Skip to content

Training Coach

The coach reads a training run and returns a diagnosis, in-whitelist proposals and an experiment plan — and every claim cites one of the cards below. This is that corpus: 22 methodology rules distilled from a real sim2real programme, and 171 episode cards behind them, each with the sentence in the war history it came from.

Doctrine

A report may cite any of these as doctrine-N.

  1. doctrine-1Contract freeze and fingerprint discipline

    The policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.

    Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).

    Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.

    contract-freeze-and-checkerlegacy-profile-pinningderived-asset-staleness-checkgain-profile-belongs-in-the-stamprecovery-two-policies-and-a-state-machinewalk-recovery-fsm-handoff

  2. doctrine-2Attribution by resolved training params - never eval-override knobs

    Capability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.

    Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).

    Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.

    kd-bandwidth-mu-law-attributiondeploy-knob-attribution-before-retrainingcycle-time-override-is-oodresolved-config-is-source-of-truth

  3. doctrine-3PASS gates become constraints; FAIL gates become objectives

    Once a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.

    Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.

    Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".

    fixed-acceptance-matrix-per-rungpreregistered-stop-criteria-per-rung

  4. doctrine-4One variable per ladder rung - counted against what the checkpoint saw

    A rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.

    Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.

    Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.

    resume-state-dr-auditorthogonal-batch-with-ablation-orderfill-the-missing-factorial-cell

  5. doctrine-5Pre-register risks, readings, and stop criteria before the ladder

    Before a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.

    Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).

    Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.

    preregister-risks-and-fork-readingspreregistered-stop-criteria-per-rungpreregistered-real-expectationsfeasibility-accounts-lock-design-point

  6. doctrine-6Plant parameters are measured, never invented

    Every plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.

    Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).

    Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.

    friction-measured-not-guessedarmature-n2-rotor-inertiatorque-limit-shape-by-measured-peakslatency-lerp-reverse-extrapolationfeasibility-accounts-lock-design-pointsim-api-friction-columnsget-up-feasibility-accounts-before-trainingsingle-support-gain-authority-probe

  7. doctrine-7Sim2sim gate before sim2real - under deployment conditions

    Every checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).

    Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).

    Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.

    sim2sim-gate-before-sim2realeval-plant-honesty-contact-paramspipeline-latency-is-plant-not-drbody-frame-velocity-api-audittorque-penalty-bought-by-leg-bracingtorque-disagreement-between-simulators-unresolved

  8. doctrine-8Observation honesty - the actor's inputs are a hardware contract

    The actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.

    Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).

    Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.

    observation-honesty-critic-onlyhistory-obs-needs-plant-variationreward-observability-limitdeploy-heading-loop-and-align-training

  9. doctrine-9Reward economics are audited in realized currency

    Reward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.

    Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).

    Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.

    realized-contribution-auditreward-cost-of-ignoring-auditgate-new-reward-terms-by-commandignore-floor-diagnosiscalibrate-threshold-between-healthy-and-sickcalibration-threshold-with-withdrawal-clauseinert-reward-term-auditseated-basin-dead-exp-kerneltail-torque-needs-hinge-on-computed-demand

  10. doctrine-10The zero-cost option must be the desired behavior

    For every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.

    Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).

    Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.

    penalize-the-slip-not-the-jointsaturation-cheating-zero-rate-costmoving-gate-42x-stand-taxtermination-closes-degenerate-basinpenalty-gate-is-an-escape-hatchsoft-limit-penalty-charges-nominal-poseunpriced-foot-attitude-is-a-free-variableenumerate-cheapest-cheats-before-trainingbinary-band-reward-fake-touchdown

  11. doctrine-11Measurement discipline: independent referees, signs, distributions

    A disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.

    Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).

    Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.

    independent-referee-for-metric-disputesbody-frame-velocity-api-auditsame-sign-response-is-yaw-biasmedian-hides-bimodal-distributionratio-metrics-need-absolute-checkheading-integral-not-body-ratesame-distribution-reward-comparisonmultiseed-sign-test-for-driftsingle-impulse-recovery-is-chaotic

  12. doctrine-12The deployment pipeline is plant

    Irreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.

    Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).

    Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.

    pipeline-latency-is-plant-not-drpower-scale-hurts-nonforward-axesdeploy-scaling-not-training-equivalentteleop-command-band-per-axislatency-dr-covers-measured-pipelinedeploy-rate-limiter-windupslew-anchor-is-an-integratorbeta-anchored-action-targetpower-derating-cuts-full-range-contract

  13. doctrine-13DR budget is finite; its distribution is the measured support

    Robustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.

    Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).

    Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.

    push-dr-conditional-budget-conservationdr-tail-plant-continuationconstant-value-dr-overfits-margintask-shaping-before-plant-hardeningcom-randomization-forces-leg-spreadcom-dr-rollback-on-symptomthin-dr-judged-by-channel-coveragefriction-priority-re-measured-after-plant-change

  14. doctrine-14Gates measure what hardware feels: posture, margins, stripped assists

    Acceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.

    Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).

    Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.

    task-metrics-vs-posture-metricsstand-gate-posture-not-survivalconstant-value-dr-overfits-marginsuspension-probe-removes-free-stabilizerhip-roll-sum-predicts-lateral-driftchirality-scored-separatelylow-speed-commands-reward-draggingend-state-confusion-matrixfrozen-acceptance-distribution-and-pinned-seedvideo-as-acceptance-recordepisode-length-bounds-what-a-gate-seesremoved-wall-returns-on-hardware

  15. doctrine-15Fork and root selection: recoverability, maturity, frozen rewards

    Choose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.

    Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).

    Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.

    fork-root-recoverable-shortfallroot-maturity-vs-product-qualityfine-tune-reward-change-falsified

  16. doctrine-16Curricula: verified engagement, lineage counters, disease-phase gating

    Automatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.

    Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).

    Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.

    auto-curriculum-engagement-checkcurriculum-counter-lineage-stepsgate-penalties-to-the-disease-phaseaggregate-metrics-mask-subgroup-failurebucket-share-is-not-a-gradient-levercurriculum-criterion-conditioned-on-lagging-categoryper-step-income-drives-speed-time-gatetime-gate-vs-wide-stance-retire-the-fix

  17. doctrine-17Probe before training: feasibility first, hypotheses in tables

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.

    Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).

    Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).

    open-loop-probe-before-reward-tuninghypothesis-table-code-auditperiod-doubling-evidence-racesuspended-test-isolates-actuator-authorityconfiguration-probe-wall-not-slopeprone-dead-end-is-foot-placementdof-vel-penalty-is-not-a-pacing-knobamplitude-cut-falsified-yaw-fix

  18. doctrine-18External advice is recomputed locally; values transfer as ratios

    Every external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.

    Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).

    Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.

    external-advice-audit-against-own-arithmetictransfer-ratios-not-absoluteslatency-dr-covers-measured-pipelinereference-structure-fk-amplitude-divisioncycle-average-tracking-for-gait-quantitiesadvisor-paraphrase-vs-paper

  19. doctrine-19Hardware sessions are scripted experiments, not tuning sessions

    Real-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.

    Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).

    Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.

    risk-ordered-real-deploymentbattery-bracketed-real-abknow-zero-command-behaviorpush-test-chirality-protocolno-field-tuning-protocolreversible-single-variable-field-experimentssim-veto-needs-real-confirmationfirst-real-get-up-violent-stage-one-policystaged-hang-mat-floor-for-get-uppower-cycle-preflighttwo-machine-config-disciplinefall-guard-becomes-a-statehardware-log-is-the-attribution-input

  20. doctrine-20Close questions in writing; restart when the debt is structural

    Audited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.

    Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).

    Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.

    frozen-verdicts-semantic-boundariesstale-verdicts-under-old-stackzero-offset-calibration-shifts-envelopeplant-swap-invariants-vs-shiftsfreeze-lineage-fix-structure-restartminimal-reward-table-with-provenancewrite-hardware-verdicts-back

  21. doctrine-21Name the quantity in the space it lives in

    A goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.

    Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).

    Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.

    joint-space-proxy-for-task-space-quantitycontact-detector-single-signal-lieszero-partial-credit-is-not-an-iteration-problemheading-integral-not-body-rate

  22. doctrine-22Continuation needs a live gradient; a release is chosen by a scan

    Continue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.

    Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).

    Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.

    converged-continuation-is-poisoncheckpoint-choice-is-a-full-gate-scanstance-decided-by-get-up-pathstop-stacking-roll-back-and-auditcurriculum-history-is-part-of-the-productcontinuation-budget-not-from-zeroroot-maturity-vs-product-quality

Experience cards

39 cards matching “independent-referee-for-metric-disputes”.

  • Never referee a suspect metric with another metric from the same code - they can share the diseaseindependent-referee-for-metric-disputes
    Mechanism understoodomniattributionmeasurementattributionsim2simprocess

    To adjudicate a disputed measurement, compute the quantity by an independent method from raw state; never accept a sibling column from the same pipeline as the tiebreaker.

    Symptom

    A triple reversal on one question: sidewalk sign diagnosis (correct) was retracted using a second metric from the same script, then the retraction itself had to be retracted when that second metric turned out to be the buggy one - two opposite-direction errors on the same problem in one day, both written into the execution sheet.

    Context

    The probe's net-displacement metric suggested the sidewalk reference sign was inverted. Worried about yaw-drift pollution of net displacement, the author checked the same table's body-frame vy_mean column (~0.003 everywhere, 20-50x smaller) and retracted the sign diagnosis. But vy_mean came from mj_objectVelocity, which was silently reporting vertical velocity due to a frame bug - the "referee" was the diseased measurement. Re-measured with a truly independent computation (xmat.T @ qvel, world trajectory), the original diagnosis was confirmed: saw -0.5 gave vy +0.058/-0.130 (76%/106%), consistent with the net-displacement values all along (yaw pollution was real but only 10-21 deg, nowhere near reversal-sized).

    Change

    Lesson written twice, verbatim, as a hard rule: when questioning a measurement, the referee must be an independent algorithm (different code path, different physical derivation), e.g. rotate qvel by the body matrix directly, or inspect the raw world trajectory.

    Outcome

    With the independent referee in place the frame bug was confirmed, fixed, and the whole C4 line re-scored - revealing sidewalk had been working (see body-frame-velocity-api-audit).

    Mechanism

    Metrics sharing a code path (or an upstream API) share failure modes; agreement between them is evidence about the code, not the world. Only a measurement with an independent derivation can break the tie, because its errors are uncorrelated with the suspect's.

    Applies when

    • two metrics of the same quantity disagree
    • about to retract a conclusion based on a second readout
    • auditing evaluation code after a surprising result
    “我用一个坏指标去质疑一个好指标,并把撤回写进了执行单。教训(写死):质疑一个测量时,不能用同一份代码里的另一个测量当裁判 —— 它们可能同源同病。裁判必须是独立算法(这次的裁判应该一开始就是 xmat.T · qvel[:3],或直接看世界轨迹)。”
    train/C_LADDER_RUN.md § 3m. 二 我今天犯了两个方向相反的错 / 3n. 五 元教训
  • The median of a bimodal metric lands in the empty gap - check the distribution, and never judge swing on 3 seedsmedian-hides-bimodal-distribution
    Replicatedomnisim-evalmeasurementgate-batteryprocess

    Before quoting a median or mean, look at the distribution; report suspected-bimodal metrics as mode share plus per-mode ranges, use small-seed smoke runs only to screen trends, and size the seed count for decisions by the share resolution you need (here: 20).

    Symptom

    Years of "high swing variance" and undecidable 3-seed swing readings turned out to be one fact: the metric was bimodal all along - "历代 swing 高方差与 3-seed 不可判由此定性:一直在测双稳态系统" - and every median reported from it (e.g. 12.1 mm) described a value no seed ever produced.

    Context

    Concrete instances: s2e_pd-1400's 20 seeds split 2.6-4.9 mm vs 19.3-24.0 mm with zero seeds between; s1e-500 read 3.9 mm on seeds 0-2 but 23.1 mm median over 20 seeds; a 3-seed reading of 19.3 was logged as "double-peak optimism, lesson recurrence #4". The selection re-audit codified the sampling rule: "3-seed 的 swing 读数不可判点, 只能筛带,选点必须 20-seed" - 3 seeds may screen a band, only 20 seeds may pick a point.

    Change

    Swing (and any suspected multi-modal metric) reported as mode shares plus per-mode ranges instead of a bare median; 3-seed smoke numbers demoted to band-screening; all shipping/selection decisions moved to 20-seed batteries.

    Outcome

    The "swing debt" bookkeeping was reinterpreted as basin probability (see swing-bistability-damping-switch), and checkpoint selection stopped being whipsawed by which basin the first three seeds happened to fall into.

    Mechanism

    Central-tendency statistics presuppose unimodality; on a bimodal distribution the median tracks the mode SHARE, not any achievable behavior, and small samples alias the share entirely. Mode-aware reporting (share + per-mode stats) is the only faithful summary, and the needed sample size is set by the share resolution required.

    Applies when

    • a quality metric shows chronic high variance across seeds
    • 3-seed smoke readings contradict 20-seed batteries
    • reporting swing height, clearance, or any basin-prone metric
    “中位数落在空档里,「swing 债 −11mm」实为「50% 概率掉进拖地吸引子」。历代 swing 高方差与 3-seed 不可判由此定性:一直在测双稳态系统。… swing 跨 seed 双峰 (500 在 seed0~2 只读 3.9mm, 20-seed 中位 23.1) —— 3-seed 的 swing 读数不可判点, 只能筛带, 选点必须 20-seed。”
    train/README.md § swing 双稳态定性 / s1e 选点重审
  • A stand gate judged by survival passes a robot that wanders a meter - judge posture insteadstand-gate-posture-not-survival
    Mechanism understoodomnireal-acceptancegate-batteryreal-acceptance

    For every gate, ask what behavior the metric is a proxy for and validate its ordering against real observations; replace metrics whose ordering disagrees with reality, and demote them explicitly rather than silently.

    Symptom

    Real robot "standing" drifted 0.5-1.4 m across the floor while the sim stand gate scored a clean 20/20 - because the gate's metric was episode survival, which wandering does not violate.

    Context

    C2's stand condition was originally written as "survival regression <=2/20". A real-robot counter-example on 2026-08-08 forced the re-judgment: wandering robots survive. Cross-checking candidate sim metrics against real-robot feel showed max-tilt median ordering agreed with hands-on ranking, while displacement ordering was actually OPPOSITE to real impressions - so displacement was demoted to a reference quantity, not a gate.

    Change

    Stand PASS criterion rewritten from survival to posture: "stand tilt max median <= root baseline +1.5 deg"; displacement kept only as reference. Applied to all subsequent rungs (C4 and redo levels inherit it).

    Outcome

    Later rungs gated stand on tilt (e.g. C4 product: 7.7 deg vs parent 7.1 deg, +0.6 deg PASS); the wandering failure mode became visible to the battery instead of hidden by survival.

    Mechanism

    A gate metric is a proxy for an intended behavior; survival is a proxy for "did not fall", not "stood still". Metric choice must be validated against ground truth (real-robot feel/measurement), and a proxy whose ordering disagrees with reality on real data is worse than no metric - it steers selection backwards.

    Applies when

    • writing PASS conditions for stand/idle/hold behaviors
    • a gate passes policies that visibly misbehave on hardware
    • choosing between candidate metrics for an acceptance battery
    “stand 必须用位姿判,不能用存活判(2026-08-08 真机反证改判):站着乱走 0.5~1.4 m 时存活照样 20/20 —— C2 的 stand 条件原写「存活退化 ≤2/20」,选错了指标。改为: stand 倾角 max 中位 ≤ 根基线 +1.5°(倾角与真机手感排序一致;位移排序与真机相反,降为参考量)”
    train/C_LADDER_RUN.md § 3b. PASS 条件 ⚠️ stand 必须用位姿判
  • A joint-velocity penalty meant to slow the get-up cut joint speed 16% and left the get-up time unchanged - the knob never moved the variable, so the idea it was meant to test stayed untesteddof-vel-penalty-is-not-a-pacing-knob
    Mechanism understoodrecoveryreward-shapingreward-shapingattribution

    Before reading a result as a test of an idea, check that the knob actually moved the independent variable; velocity regularizers smooth a schedule they do not set, and a schedule driven by per-step task income moves only when that income's time structure does.

    Symptom

    The get-up took 0.6-0.9 s in Isaac with large torque demand; the user proposed getting up more slowly so less torque would be needed.

    Context

    The idea had support in the accounts: the acceptance bound is an upper bound of 5 s (5-8x margin), the quasi-static squat path peaks at 25% of the limits, and rolling over needs no momentum. R3.2 raised dof_vel from -1e-3 to -5e-3 as the single variable.

    Change

    dof_vel -1e-3 -> -5e-3 (child-run from R3.1).

    Outcome

    Get-up medians moved +0.02-0.04 s (noise); raw joint velocity -16%; torque demand median got worse (hip_pitch 46-48% -> 63-67%) as the new term competed with torque_headroom on the same joints; MuJoCo 98 -> 96%. Verdict FAIL on the knob, not on the idea, and the rung was not adopted. When pace was later attacked through the income's time structure (V2.5/V2.5b), the MuJoCo get-up moved into the 3.5-4.5 s design band.

    Mechanism

    The pace was set by base_height_progress paying for every step spent high (stand earlier, earn more); a velocity regularizer only smooths motion along the same schedule and does not change when the robot stands up.

    Applies when

    • trying to make a skill slower or gentler with smoothness penalties
    • an experiment's primary metric did not move and a verdict is being written
    • two penalties act on the same joints
    “**关键判读:`dof_vel` 罚只把关节速度压了 16%,而起身用时一点没变。** 也就是说**这一级根本没有把"慢下来"这个自变量推动起来** —— 所以它**不构成对 用户假说的检验** … 起身节奏由 `base_height_progress` 的逐步计酬决定(早站起来就多 拿),速度正则只在同一条时间轨迹上把动作抹匀,不改变何时站起来。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §25 R3.2(dof_vel −1e-3→−5e-3,慢一点起身)
  • A ratio metric flipped the verdict - spectral share rose while absolute high-frequency energy fell 16%ratio-metrics-need-absolute-check
    Mechanism understoodwalksim-evalmeasurementattributionprocess

    Never compare share/percentage/centroid metrics across conditions whose totals differ; pair every ratio with its absolute numerator before issuing a verdict, and log retracted judgments so they are not re-derived.

    Symptom

    walk_v6 was provisionally judged "more jittery" than v5 because the joint-velocity spectral centroid rose 2.91 -> 3.62 Hz and the >4 Hz energy share rose 12.6% -> 17.8%.

    Context

    Absolute measures said the opposite: first differences of actions fell 1.54 -> 1.31, second differences 2.55 -> 2.19, and absolute high-frequency energy fell 16% (0.490 -> 0.413). The shares and centroid rose only because low-frequency content fell even more - the denominator shrank. The interim judgment was retracted in writing so it would not be reused.

    Change

    Metric discipline noted: "占比类指标在总量变化时不能直接比较" - share/ratio metrics are not comparable across conditions when the total changes; verdicts about smoothness must cite absolute energies or difference norms.

    Outcome

    v6 correctly classified as smoother, not jitterier; the retracted judgment logged under "被推翻的一个中间判断(记下来免得复用)".

    Mechanism

    A ratio confounds numerator and denominator; any intervention that removes low-frequency content raises every high-frequency share without adding a single joule of jitter. Only absolute quantities support cross-condition comparison when totals move.

    Applies when

    • comparing smoothness/jitter/spectral metrics across versions
    • any percentage-based metric moves after an intervention
    • writing an eval report that includes normalized quantities
    “我一度说"v6 动作更抖" … 错了: 动作一阶差 1.54→1.31、二阶差 2.55→2.19 都在降 … 谱质心升高只是因为低频成分掉得更多, 绝对高频能量实际下降 16%(0.490→0.413)。占比类指标在总量变化时不能直接比较。”
    train/WALK_DIAGNOSIS.md § 过程中被推翻的一个中间判断(记下来免得复用)
  • The +/-50 mm lateral COM randomization meant to spread the legs coincided with legs pulling IN - rolled back per its own pre-registered contractcom-dr-rollback-on-symptom
    Observed oncewalkdr-tuningdomain-randomizationattributiongate-battery

    When adopting a DR value that covers no local measurement, write its intent and rollback trigger into the config at adoption time; roll it back as the control arm the moment the symptom contradicts the intent, and promote the symptom's metric into the acceptance battery.

    Symptom

    After v7 adopted the reference developer's oversized lateral COM randomization (+/-50 mm) explicitly to force leg spread, the real robot's legs narrowed instead - lateral mean 154 mm / closest 107 mm in sim (nominal 214.5), narrower still on hardware with occasional leg contact.

    Context

    The rollback was clean because the adoption had been honest: the robot.yaml comment recorded the intent AND that the +/-50 value covered no local measurement (only a 16/7 mm measured offset existed; even the prior widening to +/-20 was subjective), plus the reference's own reported side effect (base sway) and the note "这一项要单独跑、 单独归因". When the opposite symptom appeared, v8 returned y to +/-20 mm as the control arm ("要么没起作用、要么帮了倒忙 … 按约定退回做 对照"), kept x/z untouched (a noise-level difference not worth another variable), and named the second suspect: the landing penalty itself, via the reference's own three-link chain (landing penalty -> stance narrows -> spacing penalty needed). A gate lesson was booked in the same table: v7's sim numbers had ALREADY crossed the line (154/107 vs v5's 182/147) - "这个指标本可拦下 v7" - so foot-distance became a standing acceptance row (min >120 mm, zero leg-leg contacts).

    Change

    base_com_offset_m y: 0.050 -> 0.020 (x/z kept), regenerated through the export tool rather than hand-editing derived files; foot-distance acceptance row added.

    Outcome

    A borrowed DR lever with no local measurement basis was retired the moment its symptom contradicted its purpose, at single-variable cost; the metric that would have caught it pre-hardware entered the gate.

    Mechanism

    DR ranges shape behavior through the policy's robustness strategy, which is jointly determined with every reward term; a lever that forces stance width on one robot can be dominated by a stronger narrowing pressure (landing softness) on another. Levers adopted without local measurement must carry their own rollback trigger, because there is no nominal to argue from when they misbehave.

    Conflicts

    Causality is not fully closed in the source: the narrowing may come from the landing penalty rather than the COM lever ("腿距的第二嫌疑人是 ④ 本身"); the rollback is the pre-agreed control experiment, not a verdict that the lever caused the narrowing.

    Applies when

    • importing DR ranges or behavioral-forcing randomizations from references
    • a DR lever's observed effect contradicts its documented purpose
    • a sim metric existed that would have caught a shipped regression
    “⑥ 的本意 … 是逼策略把脚分开;真机结果是脚向内收且偶发相碰——要么没起作用、要么帮了倒忙。… 注释当时就写了"这一项要单独跑、单独归因"。现在症状出现了,按约定退回做对照。”
    train/WALK_V8_SPEC.md § 3. 改动 C — 质心随机化退回(撤销 v7-⑥ 的 y 项)
  • Foot dragging is an attractor, not a low amplitude - and joint damping is the mode switch, adjustable at deploy timeswing-bistability-damping-switch
    Mechanism understoodomniattributionattributiondomain-randomizationactuator-modelingreal-acceptance

    When a quality metric is bimodal, stop treating it as an amplitude to be trained up: map the modes against initial conditions and plant parameters, find the parameter that switches basins, apply it first as a deployment lever, and only then bake it into the training distribution (as a plant-family shift, never as an execution-mapping change).

    Symptom

    s2e_pd-1400's swing height "median 12.1 mm" hid a perfect bimodal distribution: 20 seeds split into a drag mode (2.6-4.9 mm) and a step mode (19.3-24.0 mm) with NOT ONE seed in between - the median sat in the empty gap, and "swing debt -11 mm" really meant "50% probability of falling into the drag attractor".

    Context

    Two designed experiments closed the mechanism. Test A (nominal plant, 40 seeds): step 42% / drag 58% / middle 0 - at nominal gains, initial conditions alone pick the mode, both modes 100% survivable. Test B (fixed init, kp x kd grid): kd is the mode SWITCH - at kd 1.3 all surviving cells step (13-22 mm), at kd 0.7 nearly all drag (2.7-4.3), only at kd 1.0 does init get a vote; kp >= 1.2 is dangerous (5/6 falls). Global verification at kd x1.3 (20-seed, delay 2): survival 20/20 at ZERO cost, step share 42 -> 80%, swing median 12.1 -> 18.4 mm, slip record low 334, thicker tilt margin - costs: vx 85 -> 78%, saturation +5 pp. A Pareto sweep then priced the knob: step share 42/72/75/88/82/90 across kd 1.00-1.30 with a linear vx tax of -2.3 pp per 0.1 kd - the basin gain is fully collected at kd 1.20 ("1.30 是 over-damping 纯多付税"). Mechanism: low damping leaves a landing micro-oscillation / ground-slide channel the policy can exploit to drag; damping plugs the channel.

    Change

    Deployment lever adopted: kd-scale 1.20 (conservative 1.15) as the legitimate successor to the power-0.8 crutch ("前者削幅度保稳,后者堵 拖地通道换步态,且不牺牲存活"); training-side prescription: move the DR band to nominal-1.2 x (0.9,1.1) = [1.08,1.32], deleting the [0.7,1.0) drag-teaching zone - a contract-level change requiring digest re-baselining, gated on measuring the real robot's actual kd dispersion first.

    Outcome

    The kd surgery rung (s2e_kd) delivered basin 8 -> 11/20, slip 405 -> 331, vx 81 -> 85% with no out-of-band fragility (below-band check 20/20) - "拐杖烧进分布的正确姿势", explicitly contrasted with the failed s1g amplitude version: this one changes the plant family the policy has seen, that one changed the execution mapping the policy would have to relearn.

    Mechanism

    The gait's swing behavior is a bistable dynamical system whose basin boundaries are set by plant parameters; a policy trained across a kd band that includes the drag basin has learned to inhabit it. Shifting the deployed (and then trained) damping moves the system into the step basin without touching the policy - a plant-side fix for what looked like a training deficiency.

    Applies when

    • a gait quality metric splits into distinct modes across seeds
    • deciding between more training and a gain/damping change
    • converting a deployment crutch into a training-distribution change
    “20-seed 里拖地模式 2.6~4.9mm 与迈步模式 19.3~24.0mm 各半,中间一个不落 … kd 是模式开关——kd1.3 下 6/6 存活格全迈步 … kd0.7 下几乎全拖地 … swing 债的解(至少大半)在部署端阻尼档,不在训练端 … 机理:低阻尼下落脚微振荡/贴地滑给了策略顺势拖行的通道,加阻尼堵之。”
    train/README.md § swing 双稳态定性 + kd 部署杠杆 (2026-08-07, 用户设计 Test A/B)
  • A torque-tail penalty was paid for by bracing the legs against each other - the second simulator's leg-contact count caught it, and the first explanation ("the trainer can't see self-collision") was retracted from the run's own configtorque-penalty-bought-by-leg-bracing
    Observed oncerecoverysim2sim-gatesim2simreward-shapingattribution

    When a penalty lowers a demand metric, look for what the policy traded to get there - keep self-contact frames and foot spacing as standing sim2sim readouts - and check any "the trainer cannot see X" explanation against the run's resolved config before it enters the record.

    Symptom

    After R3.1's torque_headroom term collapsed the demand tail, MuJoCo success fell 100 -> 98% and leg-leg contact frames at mu 1.0 rose 750 -> 2,190 (worst rollout 177 -> 450). The one failure (prone seed 2) had the legs crossed, one foot on the other leg, trapped at 0.067 m - visible on video.

    Context

    Across the ten prone seeds, foot spacing and leg-leg contact frames were monotonically anti-correlated, and the failure was the extreme of the series. Pulling the legs toward the midline shortens the hip_roll lever arm and lowers torque demand. At the time the spec explained it as Isaac training without self-collisions ("a free lunch in a simulator without self-collision").

    Change

    Leg-leg contact frames and foot spacing were tracked in every MuJoCo gate; R3.2's candidates were "train with self-collision on" or "a minimum leg spacing term" - not stacked.

    Outcome

    The next rung's joint-velocity penalty incidentally erased the dependency (2,190 -> 86 frames). On 08-10 the runs' logged env.yaml showed enabled_self_collisions true in both r3_1 and v2_2 (inherited from walk v10): the tangle was physically learned bracing, visible to both simulators, and the Isaac/MuJoCo contact-count gap was mesh and contact fidelity. The "self-collision debt" narrative was withdrawn for the whole line.

    Mechanism

    A penalty on demand rewards any configuration that lowers demand; legs pressed together act as a mutual support that fails when contact geometry shifts slightly.

    Conflicts

    §24 attributes the dependency to self-collisions being disabled in training; §36 retracts that from the runs' env.yaml ("§24's mechanism explanation was wrong") and keeps the older sections unedited as history.

    Applies when

    • a torque, impact or energy penalty improves its metric and cross-sim success drops
    • legs or links approach each other after a regularization change
    • an explanation relies on a simulator setting nobody checked in the run config
    “prone 十个 seed 逐条看,脚距与腿-腿接触帧数单调反相关, 而唯一失败的那条正是最极端的一条 … 机制上说得通:把腿收到身体中线附近能缩短 `hip_roll` 力臂、降低力矩需求”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §24 代价:MuJoCo 成功率 100% → 98%,病因是两腿卡住
  • A walking policy's tilt cutoff is a legal state for a recovery policy - the default 45 deg fall guard had to be raised for recovery tests and is disabled once the switch owns falls, so the abort chain becomes the recovery timeout, the operator's cut, and the firmware torque limitsfall-guard-becomes-a-state
    Observed oncerecoveryreal-deployreal-acceptancehardwareprocess

    When a new skill makes a safety cutoff's trigger a legal state, replace the cutoff with a bound of the skill's own (a timeout ending in a safe stop) instead of just switching it off; keep the operator's cut and the firmware limits as independent layers, and write every flag change into the run sheet.

    Symptom

    deploy_policy's default protection stops the robot beyond 45 deg of tilt. A recovery policy starts lying at roughly 90-97 deg, so under the default it is refused on the spot - a flag the first hanging checklist forgot.

    Context

    The layers in the sources: deploy_policy's tilt cutoff (default 45 deg, a line in the safety chain); for standalone recovery tests the cutoff was raised (110 deg in the spec's A/B sheet; 181 deg, effectively off, in some runbook commands); with --recovery-policy the cutoff is disabled because a fall is now a state, not an exception, and RECOVERY lasting over 15 s ends in a safe stop (the runbook calls it the line where the spotter steps in). Independent of the policy: firmware torque limits checked at start (12/17/11 N*m, set_torque --check), the operator cutting enable at any kicking or oscillation, and in the one-leg teleop a space-bar stop that puts the foot down.

    Change

    The flag was added to the run sheets, and the FSM replaced the removed cutoff with its own bound (the timeout).

    Outcome

    The spec records the flag omission and its fix; it does not record the FSM's timeout being exercised on hardware.

    Mechanism

    A safety cutoff encodes one policy's notion of "abnormal"; a new skill whose normal operation lies beyond it either cannot run or runs with the cutoff off, and only a replacement bound keeps the chain closed.

    Applies when

    • deploying recovery, fall-damage or acrobatic skills behind existing safety checks
    • a run sheet disables a protection flag
    • listing the abort chain for a hardware session
    “--max-tilt-deg(默认 45°,安全链第 13 行写的那个)。recovery 的合法状态覆盖整个倾角域,把它抬到 181 = 实效关闭 … RECOVERY 超时 15s 会自动安全停(看护介入线)”
    RL系统/FOLLOW THIS copy 2.md § FSM 吊挂首测 ② 落地测 / #### Recovery Policy (operator runbook, undated)
  • Four consecutive fixes were each continued from the previous fix's degraded state until the user stopped the ladder - "change parameters, don't stack errors" - rolled back to the last good checkpoint and audited the target geometry firststop-stacking-roll-back-and-audit
    Observed oncerecoveryprocessprocessfork-selectionattribution

    When successive rungs each start from the previous rung's output and the target symptom does not move, stop, roll back to the last good checkpoint and re-derive the next change from an audit; keep the measurements, discard the stacked remedies.

    Symptom

    After the real-robot splits, the V2.7 ladder tried to widen the stance: a term swap (A), a new stance knife (b), more iterations (甲), a doubled weight (乙). Stance barely moved while hip yaw ratcheted 46.7 -> 49.9 -> 52.5 deg toward its 60 deg limit.

    Context

    Each rung started from the previous rung's output. The user ruled on 2026-08-11 that things had gone wrong from V2.7-A: go back to v2_6 and rethink which parameters to change instead of stacking errors.

    Change

    乙 was killed at start and not counted; the product baseline rolled back to v2_6c model_29399; every measurement and law learned on the ladder was kept ("the data is real; what stacked was the treatment"). Before any new training, a zero-training kinematic audit of the stance targets was run.

    Outcome

    The audit found the stand_pose target itself rewarding the narrow stance (pose-target-geometric-audit) and showed geometrically why the yawed stance could not be widened with flat feet - so 乙 was proven unnecessary without running it. The next in-lineage attempts still failed, which is what established that the stance is set by the get-up path.

    Mechanism

    A rung continued from a degraded state inherits its compensations, so each new fix answers the previous fix's side effects; the yaw ratchet was the visible trace of that stacking.

    Applies when

    • three or more corrective rungs in a row without progress on the target metric
    • a side-effect metric ratchets in one direction across rungs
    • a new rung is being planned from the latest (not the best) checkpoint
    “用户裁:"从 V2.7-A 开始就出问题了,应该回到 2.6 再思考如何改变参数而不是 错误叠加。"认账:A 的补丁 → b 的新刀 → 甲的加时 → 乙的加权,每级都从上级 的**退化状态**续(yaw 46.7→52.5° 的棘轮就是叠加痕迹)。 … 数据是真的,叠加的是处置。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 方法论裁定(2026-08-11,用户):V2.7 全阶梯叫停,回滚 v2_6
  • Every acceptance run records video of the very rollout that produced the numbers - the seated basin, edge-standing feet, tangled legs and the narrow stance were all seen on video before, or instead of, a metric catching themvideo-as-acceptance-record
    Replicatedrecoverysim-evalgate-batteryreal-acceptancemeasurement

    Make video a default output of every acceptance run, rendered from the same rollout the metrics come from (fixed views including the feet), and watch it - posture failures are visible before any gate row exists for them; never let video replace or override the numeric gate.

    Symptom

    Posture failures in the recovery line kept arriving as things the numbers had no row for: a kneeling W-sit, feet standing on their outer edges, crossed legs, a stance too narrow to hold on hardware.

    Context

    From 2026-08-09 (user decision) accept_recovery renders offscreen by default, following one env for the whole episode and archiving the clip. The MuJoCo gate's --video renders three views (side, front, feet) of the same rollout the metrics come from, with all plant modelling (delay, push); the older replay-based renderer produced an independent trajectory without delay and was not used for acceptance. Video never gates: if rendering breaks, --no-video keeps the numeric gate running.

    Change

    Video as a default acceptance artifact, named per policy, category and view, reviewed by the user.

    Outcome

    The R0.1 prone clip showed the same kneel-sit as R0; the R3.1 failure clip showed the crossed legs; the v2_5 feet view showed edge standing and led to V2.6; the v2_6 videos led the user to order a real-robot A/B between v2_5b and v2_6. Once, the recorder did not start (P1c final acceptance) and the visual material had to be produced separately.

    Mechanism

    Metrics exist only for failure modes someone anticipated; video shows the unanticipated ones, and rendering the metric rollout itself guarantees the picture and the numbers describe the same episode.

    Applies when

    • setting up an acceptance pipeline for posture-sensitive skills
    • numbers pass but a human reviewer is uneasy
    • sim videos are rendered by a separate replay tool
    “**验收存视频(用户定 2026-08-09)**:`accept_recovery.py` 默认开 Isaac 离屏 渲染,跟拍一个 env 的整局并归档 … 跪坐这类盆地在数字表出现前肉眼先看见,视频是验收的 定性存档,数字门不受它影响”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §5 R0 验收门(预注册)验收存视频(用户定 2026-08-09)
  • One fixed acceptance matrix for every rung - new skill must PASS while every old skill stays within a regression budgetfixed-acceptance-matrix-per-rung
    Replicatedomnigate-batterygate-batteryprocesscurriculum

    Freeze one acceptance matrix for the whole ladder; every promotion requires the new skill's PASS plus bounded regression on every prior skill, measured against the parent's baseline under the current (bug- fixed) metric code - and include command transitions, not just steady states.

    Symptom

    Sequential skill training silently trades old skills for new ones (C1 trained away the root's backward ability); without a constant measurement frame, each rung's numbers are incomparable and regressions hide.

    Context

    The C ladder ran the same 13-cell command matrix at 20 seeds per cell at every rung (stand; vx +0.15/+0.30; vx -0.10/-0.20; vy +/-0.10; wz +/-0.20; two vx&wz combos; two vx&vy combos), with promotion requiring "新技能 PASS 且旧技能不明显退化" - old-skill regression budget <=2/20 against the parent's recorded 20-seed baseline. For the transition-rich final rung, a command-switch block was added (forward->stop, stop->backward, forward->turn, left->right, turn->forward; survival + re-track within 2 s) because "每个 steady command 都会做 ≠ 命令切换不会摔" - steady-state success does not imply switch safety, and the joystick does switches. The C4 product's gate ran 260 cells (13 x 20) all 20/20.

    Change

    Battery frozen once, reused verbatim per rung; baselines re-measured per parent (and re-measured again after the metric-frame fix, since old baselines were taken with the buggy coordinate reading - "旧基线是坏坐标系的, 不可引用").

    Outcome

    Regressions were caught at the rung that caused them (C1's backward loss, C2's vx+0.30 decay), and cross-rung comparisons stayed valid for the ladder's whole life.

    Mechanism

    A constant matrix makes every rung's output a point in the same metric space, so "did we lose anything" is a table diff, not a judgment call; the per-skill regression budget converts previously earned PASSes into standing constraints on all future training.

    Applies when

    • designing gates for sequential skill addition
    • promoting a checkpoint to be the next rung's root
    • after any evaluation-code fix (old baselines must be re-measured)
    “新技能 PASS 且旧技能不明显退化才晋级。… C5 追加:命令切换验收(steady ≠ transition) forward→stop、stop→backward、forward→turn、left→right、turn→forward,各 20 seed,判存活 + 切换后 2 s 内是否重新跟上。”
    train/C_LADDER_RUN.md § 5. 固定验收矩阵(每级跑同一张,每项 20 seed)
  • A single-signal contact detector lied in both directions - foot height flagged 40% false flight on a walking gait, contact force alone flagged false flight during low-friction slip - so flight became force < 5 N AND sole height > 5 mmcontact-detector-single-signal-lies
    Replicatedrunsim-evalmeasurementgate-battery

    Define contact and flight events from two independent signals (force and geometry) in conjunction, validate the detector on a behaviour known not to contain the event before using it as a gate, and match the trainer's threshold when comparing across simulators.

    Symptom

    The run line needed a flight-fraction gate. The first MuJoCo version, "sole higher than 2 mm", measured 40% flight on a walking policy that never flies. Five weeks later the one-leg gate, using contact force alone, reported support-foot "flight" segments at mu 0.4 for a policy that was not hopping.

    Context

    During a walking step the toe lifts or the heel strikes with the foot pitched, so the ankle-roll origin rises a few mm while part of the sole still touches - height alone calls that flight. Switching to contact force < 5 N (the same threshold Isaac's contact reward uses) zeroed the false flight on walking. In the one-leg re-test every force-only "flight" segment had a measured sole height of 0.0 mm: the normal force chattered while slip corrections played out on low friction.

    Change

    Run line: flight = contact force < 5 N ("lift-off must be judged by contact force"). One-leg line (2026-09-16): flight = force < 5 N AND sole height > 5 mm, recorded as the same measurement lesson in the opposite direction; the gate's behavioural meaning was unchanged.

    Outcome

    With force-based detection the walking policy read 0.0 flight and the run policy's zero flight was confirmed by two plants; with the conjunctive definition the one-leg support-foot gate stopped reporting false hops.

    Mechanism

    Each signal has its own failure: geometry moves without leaving the ground (foot pitch), and forces drop without leaving the ground (slip chatter); requiring both removes both families of false positives.

    Applies when

    • writing a flight, lift-off, hop or slip detector for a gate
    • a gate reports an event the video does not show
    • reusing a detector on a different gait or floor friction
    “首版用**足底高度>2mm** 判离地, 在 omni_s1e 走路策略上测出 40% 假腾空 … 改用**接触力 <5N**(与 Isaac feet_contact_number 同源阈值)后 空检归零 (walk 策略 flight_frac 0.0)。**课文: 离地判定必须用接触力, 高度判 会把脚的俯仰当腾空**”
    train/README.md § run R1 立项 (2026-08-09): Mac 侧新工具 + 一次空检抓获
  • Late training leaned on sampling noise as a stability crutch - deterministic play collapsed while training metrics stayed greennoise-crutch-deterministic-collapse
    Replicatedomnitraining-runsim2simprocess

    Evaluate the deterministic policy in an external harness on a fixed cadence during training (not just at the end), select checkpoints on that curve, and treat a collapsing noise_std with rising training reward as a warning that noise is load-bearing.

    Symptom

    omni_s1's final checkpoint (model_5999) fell at 4 s even in Isaac's OWN deterministic play, while checkpoints from iter 1700-4000 were fine - and no training metric flagged anything. Policy noise_std had collapsed to 0.045 by iter ~990 (final 0.033).

    Context

    Diagnosis: the policy had learned to use its exploration noise as a dither/stabilizer - "策略把采样噪声当稳定拐杆,训练指标看不见" (the training metrics cannot see it, because training always runs with noise on). Countermeasures: entropy_coef 0.005 -> 0.01 to slow the std collapse, and - the structural fix - an in-training smoke loop (watch_ckpt.py): every 500 iters, export ONNX directly, run 3-seed MuJoCo evaluation, log CSV/TensorBoard curves plus three-view videos. The doctrine line was written in bold: "训练指标全绿不再是发育健康的 证据,冒烟曲线才是" - green training metrics are no longer evidence of healthy development; the smoke curve is. The follow-up run s1b showed the drift metric follow a U-shape (73 -> 8.6 at iter 3500 -> 76), making checkpoint selection BY the smoke curve (early stop at 3500) the shipping mechanism, with terminal re-degradation booked as known and unresolved.

    Change

    entropy floor raised; watch_ckpt smoke loop instituted as standing infrastructure; checkpoint selection moved from "last iteration" to "best point on the deterministic smoke curve".

    Outcome

    s1b shipped from iter 3500 (the U-bottom) instead of a degraded terminus; every later lineage (s1c/s1e, the C ladder's --every 100 loops) inherited the watcher as the standard guardrail.

    Mechanism

    PPO evaluates and improves the stochastic policy; if noise itself stabilizes the gait (dither smoothing a marginal limit cycle), the deterministic mean policy is a different, worse controller that training never measures. External deterministic evaluation on an independent simulator is the only readout of what will actually be deployed.

    Applies when

    • final checkpoints underperform mid-training ones
    • noise_std collapses early while training reward climbs
    • deciding which checkpoint to export and ship
    “训练后期确定性脆化——noise_std iter~990 收到 0.045(终 0.033),model_5999 连 Isaac 确定性 play 都 4 s 摔(1700~4000 正常):策略把采样噪声当稳定拐杖,训练指标看不见。对策:entropy_coef 0.005→0.01 + train/watch_ckpt.py 训练中冒烟曲线 … 训练指标全绿不再是发育健康的证据,冒烟曲线才是。”
    train/OMNI_V0_SPEC.md § 3. S1.1 修订记录 ②
  • A curriculum ramp keyed to the process step counter re-fires on every resume - and shipped policies never saw the penaltycurriculum-counter-lineage-steps
    Mechanism understoodomnicurriculumcurriculumattribution

    Key every curriculum/ramp schedule to lineage-cumulative progress, not per-process counters; and audit where your shipped checkpoints sit relative to every ramp - a penalty that no product ever experienced is not part of your training.

    Symptom

    vx+0.30 died at a fixed relative time in every resumed run: resume at 500 -> slide at 1100, zero at 1300; resume at 700 -> slide at 1300, zero at 1500 - absolute depths offset by exactly the resume offset, relative timetable identical.

    Context

    ramp_reward_weight (the saturation penalty ramp) read env.common_step_counter, which restarts at 0 for every run including --resume. So start_step=600 meant "600 iters after THIS resume", not "lineage iteration 600". The A/B arm pair was the clean proof: their env.yaml differed only in log_dir, only resume point distinguished them, and the omni CurriculumManager had exactly one active term - nothing else could produce that timetable. Second consequence: every shipped checkpoint (s1e-500 at +500, C2-700 at +200, A800 at +100) was selected before its run's +600, so the saturation penalty weight was 0.000 for every product ever shipped - explaining saturation 33% and raw |action| 1.9 against clip 1.0 (hip_roll in bang-bang), i.e. half the heat budget.

    Change

    Two independent recommendations recorded: (1) make the ramp count lineage steps (add the checkpoint's iteration offset on resume) or pin terminal weights in downstream rungs instead of ramping; (2) give the saturation penalty its own rung - never mixed into a skill-learning rung (that would be two variables again).

    Outcome

    Explained the recurring +600 death of the highest-amplitude command and the persistent actuator saturation of all shipped products with one root cause; honest caveat booked (at +800 the weight is only -0.086, small, but vx+0.30 is the command demanding the largest action amplitude, so it is squeezed first).

    Mechanism

    Resumable training splits "the lineage" from "the process"; any schedule keyed to process-local counters silently re-applies its transient to every descendant run, and any product-selection habit that picks checkpoints early systematically samples the pre-ramp regime - the curriculum exists in the config but never in any shipped policy.

    Applies when

    • resumed/forked training with any scheduled reward or DR ramp
    • a metric dies at a fixed offset after each resume
    • shipped policies show behavior a late-schedule penalty should prevent
    “ramp_reward_weight 读的是 env.common_step_counter,它每个 run 从 0 开始,--resume 也不例外。… 原始 run / 臂B | 500 | 1100 = +600 | 1300 = +800;臂A | 700 | 1300 = +600 | 1500 = +800 … 所有出品其实从没见过饱和罚。… 这解释了 sat_max_pct 33%、raw |a| 最大 1.9(clip 是 1.0)—— hip_roll 一直在 bang-bang,而罚它的那一项权重恒 0。热账的一半在这里。”
    train/C_LADDER_RUN.md § 3g. 系统性问题:saturation_ramp 每次 resume 归零
  • Three same-shaped judging errors - task metrics (survival, tracking, displacement) cannot stand in for posture metricstask-metrics-vs-posture-metrics
    Replicatedomnireal-acceptancereal-acceptancegate-batteryattribution

    Keep validated posture-class rows (tilt, per-joint L/R asymmetry, temperature) in every acceptance battery alongside task rows; when operator feel contradicts the gates, suspect the metric class before the operator - and never build a new skill on what is actually an asymmetry defect.

    Symptom

    The C2 product judged "full pass" on task metrics (A800: turn-gap 7 pp, vx+0.30 19/20) felt WORSE in the operator's hands than the half-pass 700: A800 tilted up to 12.90 deg (700: 6.64), drifted left while standing, showed larger per-joint asymmetries, and ran its hip_rolls 5 degC hotter.

    Context

    The re-judgment catalogued three same-type metric errors in one campaign: (1) stand judged by SURVIVAL - missed 0.5-1.4 m wandering; (2) stand ranked by DISPLACEMENT - ordering was opposite to real feel (tilt ordering matched); (3) chirality judged by wz-tracking GAP - measured turning symmetry while the robot's actual disease was postural left/right asymmetry, "两个不同的东西,且结论相反". Common pattern named: "我一直用「任务指标」当判据,而真机手感对应的是「姿态 指标」… 任务类指标不能替代它". The fix was already in the data: the per-joint left/right asymmetry table (printed identically by sim2sim and deploy) agreed with hardware in direction on every row - "判据可用、有预测力,我只是没把它写进 PASS 条件". Shipping decision followed the posture read: product reverted to 700 ("又一次「买到 精度、卖掉别的」"), and the C4 root moved to 700 as well, with the sharpest line of the episode: A800's left-drift "like sidewalking" is probably its frontal-plane asymmetry defect, not a capability - "在缺陷上建能力是危险的".

    Change

    Two posture quantities with demonstrated real-robot predictive power promoted into every PASS battery: tilt-max median and per-joint left/right asymmetry (both sim-computable, deploy-homologous); motor-temperature readout added to session close-out.

    Outcome

    Deployment flipped to the posture-better checkpoint; the hip_roll temperature table (43-48 degC vs 25-28) confirmed the earlier 90%-of-heat account; the run-line acceptance battery inherited the posture rows from birth ("任务类替代不了姿态类").

    Mechanism

    Task metrics measure goal attainment under the evaluator's episode definition; posture metrics measure the body state trajectory that operators, motors, and downstream skills actually experience. The two can rank candidates oppositely because task success tolerates postural pathology - so a battery without posture rows is blind to exactly what hardware feel reports first.

    Applies when

    • hardware feel disagrees with a green acceptance table
    • choosing between checkpoints that split task vs posture metrics
    • selecting the root for a skill that resembles an existing defect
    “共同模式:我一直用「任务指标」(存活 / 跟踪率 / 位移)当判据,而真机手感对应的是「姿态指标」(倾角、逐关节左右不对称)。→ 验收判据里必须有姿态类指标,任务类指标不能替代它。… A800 的「左飘像 side walk」很可能 … 是它更大的额平面不对称的表现 —— 在缺陷上建能力是危险的。”
    train/README.md § C2 选点改判 (2026-08-09): 手性判据第三次选错指标
  • Knee swing collapsed because it directly trades against the slip penalty - price the conflict explicitly and clamp what reward cannot holdknee-swing-vs-slip-pricing
    Mechanism understoodwalkreward-shapingreward-shapingattribution

    When a behavior collapses as another metric improves, look for the term pair trading them and set their price ratio deliberately (with escalation and reverse tripwires pre-registered); where the policy actively spends action budget to undo your target, stop paying more reward and clamp the target space structurally.

    Symptom

    Knee peak-to-peak swing collapsed across generations - v5 33 deg, v10 26-30, v10b 7-8, v11 6.5-8.6 - and rolling the clock back did not recover it, acquitting the clock; the collapse tracked the gated slip penalty instead: "屈膝与不打滑在当前奖励里直接对抗" - v10b's excellent 93 deg slip was purchased with knee amplitude.

    Context

    Reward-side flexion fixes had failed three times: raising reference amplitude backfired twice (v9/v11), and v11's deep-squat default was actively fought by the policy - it spent 0.68 of action budget pulling the squat straight ("被策略花 0.68 动作拉直反杀"). v12's design accepted the conflict as real and attacked on two tracks: (1) ECONOMICS - a direct knee_swing_amplitude reward (+0.3, target 0.55 rad, capped at 0.6/step = 55% of tracking), explicitly opposed to the slip penalty by design ("显式对立——这正是设计:v12 就是这场对抗的定价实验"), with an escalation ladder (K +0.3 -> +0.5, then slip -0.5 -> -0.3, one layer at a time) and a reverse tripwire (slip telemetry back at v10 levels -> slip weight to -0.8, accept ~20 deg knee compromise); (2) STRUCTURE - knee target bounds [0.2, 0.9] rad so full straightening is physically impossible (straightest 11.5 deg) and the 0.68 fighting budget is released. A bonus falsifiable prediction was attached: phase-lock strength tracks amplitude (v9_probe 48 deg locked 2.5 Hz; v11 low-amplitude 1.36 Hz unlocked), so if K works, hardware phase-lock should return - one change, two verdicts.

    Change

    knee_swing_amplitude reward + knee target clamp + pre-registered escalation/reverse levers; the failed reward-side-only approach retired.

    Outcome

    The lineage was frozen before v12 trained (strategic reset), but the diagnosis stands as the walk line's clearest example of two reward terms trading a behavior between them, with the pricing experiment and structural clamp fully designed and calibrated.

    Mechanism

    When two terms price opposite aspects of one motion (swing amplitude creates yaw momentum that becomes slip), the optimizer settles wherever the price ratio puts it - patching one side moves the equilibrium, not the conflict. Explicit pricing makes the trade a designed quantity; structural clamps remove the regions where the policy spends budget fighting the designer.

    Conflicts

    The pricing experiment (K vs slip) was designed and calibrated but never trained - the 2026-08-05 reset suspended v12; the collapse attribution table and the 0.68-action counterattack are measured, the remedy's效果 is untested.

    Applies when

    • one gait quality degrades in lockstep with another's improvement
    • the policy visibly fights a default pose or reference
    • repeated reward-side fixes for the same behavior have failed
    “膝摆塌在 v10→v10b,头号嫌疑是门控滑移罚(四代实测膝 p2p:v5 33° / v10 26~30° / v10b 7~8° / v11 6.5~8.6°;退时钟没救回 → 非时钟)——"屈膝"与"不打滑"在当前奖励里直接对抗 … 奖励侧修屈膝已三败 … v11 深蹲 default 被策略花 0.68 动作拉直反杀”
    train/WALK_V12_SPEC.md § 0. 定位 / 2. K —— 膝摆经济(与滑移罚的对偶)
  • Four in-lineage attempts to widen the standing stance failed - remove a tax, add a joint-space knife, change the target, add a task-space metric penalty - because the stance was the end state of the get-up path; trained from scratch with the right terms it grew right from day onestance-decided-by-get-up-path
    Replicatedrecoveryattributioncurriculumfork-selectionreward-shaping

    A posture a skill ends in is shaped by the path the policy takes to reach it; if several single-variable edits to the terminal-phase reward cannot move it, stop editing that phase and retrain with the terminal constraint present from the start.

    Symptom

    v2_6c stood with its feet 0.159 m apart (task-space) and its hips yawed 45-47 deg the same way, which split on the real robot. Standing-phase reward edits did not move it.

    Context

    V2.7-A removed the flat-feet tax on compensated stances (stance unchanged); V2.7b added a hip-roll lower-bound hinge (+5 deg in 3,000 iterations, yaw ratchet); V2.8 changed the stand_pose target to a wide flat stance (stance unchanged, yaw not unwound, feet nearly overlapping, mu 0.4 transfer 2%); V2.9 penalized lateral spacing in metres (the policy parked just outside the penalty's gate in a lunge, 0% success). The v2_6c get-up goes through a split and closes the feet together as it rises.

    Change

    In-lineage stance surgery was formally closed. V3.1 trained from scratch with task-space stance terms present from the first iteration (and, after P1, a positive width band instead of a penalty).

    Outcome

    V3.1 P1b: lateral stance 0.364 m, foot tilt 0.0 deg, all four categories 100%, MuJoCo mu 1.0 and 0.4 both 100% - with a symmetric toe-out the kinematic audit had not enumerated. P1c (with a yaw guard): 0.355 m, all six acceptance criteria passing, mu 1.0-0.4 all 100%; it became the product.

    Mechanism

    A converged policy does not rebuild the path that produced its terminal posture; a standing-phase gradient only finds the nearest hack around the posture the get-up delivers.

    Applies when

    • the final posture of a transition skill is wrong and resists terminal-phase shaping
    • repeated continuation rungs produce hacks instead of the intended posture
    • deciding between another in-lineage fix and a from-scratch retrain
    “窄站距 + yaw 扭是 v2_6c 起身策略(劈叉起身 → 双脚并拢收势)的**结构性 终态**,不是站立段的孤立参数 —— 站立形态由起身路径决定,在血统内只动 站立段奖励改不动它。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 结果:V2.8 判 FAIL —— 血统内站姿手术第三次证伪
  • With no term on foot attitude the robot stood on the edges of its feet (ankle roll -26 deg) from R0.5 to V2.5b; pricing it fixed that and turned stance width into the next free variable - the user saw both on video before any metric flagged themunpriced-foot-attitude-is-a-free-variable
    Replicatedrecoveryreward-shapingreward-shapingreal-acceptance

    List every posture quantity the hardware cares about (foot attitude, stance width, yaw, knee flexion) and make sure each is priced by a term with real gradient at the observed error; pricing one exposes the next, so re-inspect the feet-level video after every change.

    Symptom

    Watching the v2_5 video the user said the ankle roll after standing looked very strange - the feet were not flat. The feet view showed the right foot standing on its outer edge; median ankle roll at t = 8 s was -/+26 deg (limit +/-35), mirrored. This was the physical form of the ankle_roll saturation criterion that had failed since R0.5.

    Context

    No term priced foot attitude: feet_contact_upright counts an edge contact as contact, and stand_pose's wide exp kernel (sigma^2 = 9) gives almost no gradient at 0.45 rad. HoST carries an "ankle parallel" term (+20); this table never had one. Edge standing was made a blocking precondition for hardware (continuous ankle load, unstable contact, wear).

    Change

    V2.6 (single variable): flat_feet = (|q_l_ankle_roll| + |q_r_ankle_roll|) x upright gate x height gate, a linear hinge with a 5 deg margin, weight -2, continued from V2.5b.

    Outcome

    Ankle roll -/+26 -> 5.3/5.2 deg, ankle_roll saturation 6.6% (< 10%), feet flat on the feet-view video; success 99.6%, re-falls 0-1%. Then the user watched v2_6: hip yaw constantly tense and the legs very close together. The numbers: hip roll -/+4.9/4.8 deg against a nominal 25 - feet flat and hips open 25 deg cannot coexist without ankle compensation, the new term taxed that compensation, and nothing priced stance width. "Foot attitude as a free variable" was fixed and "stance width became the new free variable" - which the real robot then exposed as splits.

    Mechanism

    An optimizer spends every posture degree of freedom no term prices; closing one reallocates the slack to the next unpriced one.

    Applies when

    • a standing or landing posture looks wrong on video while gates pass
    • a saturation criterion keeps failing on one joint
    • a new posture term was just added
    “用户看 v2_5 视频:"起身之后 ankle_roll 非常奇怪,脚根本不是平着站立"。 … 机理:奖励表**无任何脚掌姿态项** —— feet_contact_upright 边缘接触也算触地, stand_pose 的 exp 核(σ²=9)对 26°=0.45 rad 梯度≈0。脚掌姿态是自由变量。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §41 预注册 V2.6:flat_feet —— ⑥ 老账的物理形态被用户目视锁定
  • A plateau in a training curve was a population mix, not a half-learned skill - 60% standing at 0.372 m and 40% sitting at 0.19 m - and a category at hard zero stayed at zero through 3,000 more iterationszero-partial-credit-is-not-an-iteration-problem
    Mechanism understoodrecoveryattributionmeasurementattributioncurriculum

    Before buying iterations for a plateau, split the metric by category and check whether it is bimodal; a category at hard zero with no partial credit is missing a capability or a reachable state, and more iterations under an unchanged config will only polish the categories that already work.

    Symptom

    After R0.1 the training-side base_height sat near 0.29 m and the curve was still climbing at the iteration cap, which read as "train it longer".

    Context

    Candidate A (user decision) was a child-run from R0.1's last checkpoint with zero config change - the logged env.yaml files differ only in log_dir - for 3,000 more iterations (R0.2). The per-category acceptance split was already available: prone had scored 0/156 with no partial credit.

    Change

    Continue training unchanged, then read the result by category rather than by the pooled curve.

    Outcome

    supine 91.5 -> 96.4%, side 82.2 -> 87.9%, get-up 0.96 -> 0.84 s, pose error 0.91 -> 0.54 - all improvements to categories that already stood. Prone stayed 0/153; mid 55.9 -> 41.2% was within noise (n=34). Height by category was binary - standing groups 0.372/0.373 m, seated groups 0.187/0.194 m, nothing between - so the pooled 0.29 was 0.61 x 0.372 + 0.39 x 0.19 = 0.30 (measured 0.307): six in ten standing, four in ten sitting. The rendered prone episode was still kneel-sitting at t = 8 s.

    Mechanism

    The pooled mean of a binary outcome only moves when the mix moves; PPO kept polishing the subpopulation that already succeeded while the failing one produced no advantage signal to follow.

    Conflicts

    R0.2 recorded the missing capability as "prone lacks rolling over"; R0.3's end-state confusion matrix retracted that - prone had righted its torso in 159/159 episodes and was failing to stand from the W-sit. The lesson that iterations could not fix it holds; the named cause was wrong.

    Applies when

    • a training curve plateaus while acceptance shows one category at zero
    • deciding between "train longer" and "change something"
    • pooled training metrics are read as the typical episode
    “**分类别 h 中位把"平台 = 人口混合"钉死了**:数值是**二值**的 —— 站立组 0.372/0.373,坐姿组 0.187/0.194,**中间没有过渡态**。 … **这也是本仓此后读该指标的通用告诫:全体混合的期望会把 双峰分布平均成一个不存在的中间值,必须分类别看。** … **结论:A 不能过门,原因确定为 prone 缺"翻身"这一技能,不是迭代不够。**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §13 R0.2(recovery_r0_2,child-run 续训):A 走完了 —— 推不动 prone
  • Adapting a lineage to one plant increment needs hundreds of iterations, not thousands - long runs only buy specializationcontinuation-budget-not-from-zero
    Mechanism understoodomnitraining-runcurriculumprocess

    Budget continuation rungs by increment class (hundreds of iterations for plant pins and smooth shifts, ~1000-1500 only for behavior-demanding changes like push), enforce a hard cap with frequent evaluation, and treat remaining budget as a reason to stop, not to continue.

    Symptom

    The default "6000 iterations per rung" (a from-zero-scale budget) was about to be applied to continuation rungs whose only change is one plant/DR increment - overspending compute and, worse, giving each rung thousands of iterations to specialize away retained skills.

    Context

    The 2026-08-07 budget table replaced the default with "最低适应窗口 + 每 100 iter 验收 + hard cap" scaled to the increment's difficulty: fixed-latency levels 300-500 (cap 500-800; the base has already seen in-band values, this only pins the plant); PD full-band 700 (cap 1000; kp+/-20%/kd+/-30% clearly widens the actuator family); COM +/-20 mm 500 (cap 800; a smooth dynamics shift); friction DR 700 (cap 1000; contact and actuator friction change the gait/contact solution together); push 1000 (cap 1500; a non-static plant change requiring recovery behavior - hardest). Rationale: "续训适应一个 plant 增量不需要从零量级的预算,跑长了只是给特化时间". The C ladder reused the scheme (per-rung caps 500-2000 by increment type), and the deep-training hazard got its own name when long runs sold quality ("深适应卖质量" - deep adaptation sells quality).

    Change

    Per-rung iteration budgets set by increment class with hard caps and 100-iter watch loops; checkpoint selection inside the window by the smoke curve, never "run to cap because budget remains".

    Outcome

    S2/C rungs completed in 300-1500 iterations each; the recurring late-run degradations (collapse valleys at 1500+, vx+0.30 decay) fell outside most rungs' caps instead of inside their runs.

    Mechanism

    A continuation rung's learning problem is local robustification around an existing optimum - low sample complexity; iterations past adaptation are spent sharpening onto the current distribution, which is exactly how retained skills and margins erode. Budgets sized to the increment bound both compute and the specialization damage window.

    Applies when

    • planning iteration budgets for a robustification or command ladder
    • a continuation run keeps improving its training metric late
    • retained skills decay in the back half of long continuation runs
    “「最低适应窗口 + 每 100 iter 验收(watch_ckpt --every 100)+ hard cap」—— 续训适应一个 plant 增量不需要从零量级的预算,跑长了只是给特化时间 … ⑥ push | 1000 | 1500 | 非静态 plant 变化,要学 recovery 行为,最难”
    train/OMNI_V0_SPEC.md § 4. 每级 iter 预算(2026-08-07 用户定)
  • A pull-assist curriculum keyed to a global success share was satisfied by the categories that already worked and withdrew before prone learned anything - conditioning the criterion on prone took it from 2.5% to 98.7%curriculum-criterion-conditioned-on-lagging-category
    Mechanism understoodrecoverytraining-runcurriculumgate-battery

    Measure a curriculum's advancement criterion on the population the scaffold is meant to help; a global success share is met by whatever already works, and the help is withdrawn before the lagging case learns.

    Symptom

    Under the beta-anchored action, supine and side stood reliably while prone still sat (1.3%). A pull-assist curriculum added to help it was withdrawn completely within ~790 iterations and prone moved only to 2.5% (noise).

    Context

    The pull assist follows HoST: an upward force on the base, active only when the torso is within 30 deg of vertical, scaled by body weight (HoST's 200 N on G1 = 0.583 BW -> 56 N here, steps of 5.6 N, ten levels to zero); the product must pass with no assist. It had already taught sit-to-stand in V1. In V2.1 its advancement criterion was the standing-time share over all envs (threshold raised to 0.55 because the share was already ~0.53).

    Change

    V2.2: PullAssistForce with gate_category = "prone" - only envs whose first step classifies them as prone count toward the criterion - and the threshold back at 0.35. A feasibility signal was pre-registered: if prone's share stayed near zero under the full 56 N, return to the roll-over path instead of adding force.

    Outcome

    The prone-conditioned curriculum withdrew level by level only as prone itself passed: prone 98.7%, and the four-category gate passed for the first time on the line (98.6% overall, re-falls 0%, torque gate PASS).

    Mechanism

    A pooled success share is filled by the categories that already succeed (supine/side ~53%), so the scaffold is removed on their account before the lagging category has used it.

    Applies when

    • an assist, guide force or easier setting is withdrawn by a success threshold
    • one task category lags while the pooled metric looks healthy
    • a curriculum ran to completion without changing the lagging category
    “**教训:全局站立占比阈会被存量类别(supine/side ~53%)凑够,拉力在 prone 学会前就撤光了 —— metric 设计失误,不是拉力机制失效**(它在 V1 教会过 坐→站)。 … **V2.2(已启动)**:`PullAssistForce` 加 `gate_category="prone"` —— 达标判据 只统计 prone 类 env”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §34 V2.1 判决(2026-08-10)
  • A get-up policy righted itself and sat - three terms paid the seated pose 84% of the return, and the only shaping term that could tell sitting from standing was an exp kernel outputting 5e-5seated-basin-dead-exp-kernel
    Mechanism understoodrecoveryreward-shapingreward-shapingattribution

    When a policy parks in a degenerate posture, tabulate what each reward term pays that posture against the target (watch contact terms that reward touching rather than bearing load) and evaluate every exp kernel at the error actually observed; a kernel narrower than the real error is switched off, and widening it is a one-variable repair that adds nothing new.

    Symptom

    R0 converged by iteration 700 and gained 1.7% over the next 2,300; success 0.0% in all four fall categories. Three of four success conditions passed (tilt median 1.0 deg, both feet in contact 99.6%, angular rate low); height passed 0.6% (median 0.204 m against 0.326). The robot knelt in a W-sit: hip yaw +/-47 deg, knees folded to 92% of the hard limit, shins flat, pelvis on the ground, torso vertical.

    Context

    Minimal reward table: upright (1-g_z)/2 +2.0, base_height linear progress +1.5, stand_pose exp(-||q-q_stand||^2/std^2) x upright gate +1.0 with std 1.0, still +0.5 and feet_on_ground +0.5 both x the upright gate, plus regularizers. The upright gate is a hinge that opens below 30 deg of tilt. The pre-registered fallbacks were then checked against the measured state: tightening the tilt gate was falsified (tilt was already 1.0 deg); a success bonus contradicted the spec's own no-cliff-bounty rule; narrowing the categories was useless (all four converged to the same pose); raising init noise was too weak for a basin this deep. Only a half-rise intermediate state addressed it, and a cheaper repair existed.

    Change

    R0.1 (user decision, single variable): stand_pose std 1.0 -> 3.0. Not a new term and not a bounty - repairing a declared term that was numerically dead. The two runs' logged env.yaml differ in log_dir and std only.

    Outcome

    R0.1 58.2% overall (R0 0.0%): supine 91.5%, side 82.2%, mid 55.9%, prone 0/156; knees fully straight; ||q-q_stand||^2 9.99 -> 0.91 and the stand_pose term 4.6e-5 -> 0.90; get-up ~1 s, no re-falls, the curve still rising at the 3,000-iteration cap. Prone stayed at zero and needed a different fix (see prone-dead-end-is-foot-placement).

    Mechanism

    Sitting earned upright 1.98/2.0, still 0.43/0.5 and feet_on_ground 0.46/0.5 - 3.0 of a 3.57 per-second return - because feet_on_ground asked for contact, not load. The only terms separating sitting from standing were base_height (+0.70/s for standing) and stand_pose, whose kernel at the real 9.99 rad^2 error (75% of it in the two knees) was exp(-9.99) = 4.6e-5 with a gradient near 1e-4. Standing up meant risking 3.0/s to gain 0.70/s while unfolding knees at 92% of their limit under load. With std 3 the same term is exp(-9.99/9) = 0.33 - a live gradient, three quarters of it on the folded knees.

    Applies when

    • a policy converges early to an upright but low, seated or kneeling pose
    • a posture-matching exp term reads ~0 in the training logs
    • contact-based rewards saturate while the task metric does not move
    “**关键:`feet_on_ground` 只问"触地"不问"承重", 跪坐时双脚确实贴地,照样满分。** 三项 3.0/s = 总回报 3.57/s 的 84%。 … **exp(−9.99) = 4.6e-5** —— 权重 1.0 的项实际输出 5e-5、梯度 ~1e-4, **不是"还没学会",是数值上根本不存在**。 … **R0.1 决定(用户 2026-08-09 定,单变量)**:`stand_pose` 的 `std` **1.0 → 3.0**。 不是加新奖励、不是悬崖悬赏,而是**修复一个已声明但数值失效的项**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §11 R0 首跑(recovery_r0, 2026-08-09):FAIL —— 翻正了但坐着
  • A time gate that had cured one lineage's rushing made the from-scratch lineage trade away its stance width twice (0.364 -> 0.235 m, 0.355 -> 0.251 m) - its disease was absent there, so the fix was retired and the pre-gate checkpoint shippedtime-gate-vs-wide-stance-retire-the-fix
    Replicatedrecoveryreward-shapingreward-shapingcurriculumfork-selection

    Carry a fix into a new lineage only if its disease is present there; a mechanism that cured one lineage can be net negative in another, and when doubling a term's weight recovers almost nothing, treat the two objectives as structurally in conflict and remove the one whose purpose is gone.

    Symptom

    V3.1's phase 2 (the 3 s zero gate on standing income, continued from P1b) kept 100% success on every friction level and slowed the get-up, but the lateral stance drifted 0.364 -> 0.235 m and hip yaw crept to 57 deg against its 60 deg limit. With the width band's weight doubled (P2c, after P1c) it drifted again, 0.355 -> 0.251 m, below the pre-registered 0.30 m failure line.

    Context

    The zero gate had been introduced in V2.5/V2.5b to slow the old lineage's get-up. In V3.1 the rushing was already absent: P1c got up in 0.90-1.06 s with a worst torque ratio of 73.3%, better than the stamped v2_6c, because the full beta curriculum, second-difference smoothing and pull curriculum had cured the violence inside training.

    Change

    Recorded as a candidate law with two data points - the zero gate and a wide stance are mutually exclusive here - and the zero gate was removed from the V3.1 recipe. P1c (the pre-gate checkpoint) went through the full stamp-level acceptance instead.

    Outcome

    P1c passed everything: all six criteria, lateral stance 0.355 m, foot tilt P75 2.0 deg, mu {1.0, 0.8, 0.6, 0.4} x 10 seeds all 100%. recovery_v3_1p1c.onnx was stamped and pushed to the robot channel.

    Mechanism

    The zero gate moves the income toward "stay stable until the end", and under low-friction DR a wide stance has a slip tail, so survival outbids the width band; doubling the band's price bought back only 0.016 m - an auction that does not converge signals structural conflict, not an under-priced term.

    Applies when

    • porting reward mechanisms from an old lineage into a fresh recipe
    • a width, margin or posture metric erodes during a late training phase
    • a weight increase produces a negligible change in its target
    “**定律候选(二实证):归零门 × 宽站互斥**。 … 加价翻倍只挽回 0.016,竞拍不收敛)。 … **归零门是 v2_5 血统的历史包袱,对 V3.1 配方是净负资产,P2 阶段除名 —— P1c 即终点形态**。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §49 终验(2026-08-14)
  • An exponential kernel on instantaneous velocity punishes gait oscillation - track the cycle averagecycle-average-tracking-for-gait-quantities
    Mechanism understoodomnireward-shapingreward-shapingattribution

    Reward velocity tracking on gait-cycle averages (or filtered values), not instantaneous samples, whenever the desired behavior oscillates at stride frequency; widening the kernel does not fix a variance penalty.

    Symptom

    Even while the robot genuinely sidewalked (verified after the metric fix), the Isaac-side tracking reward sat on the ignore-floor: true sidewalk scored 0.178 vs 0.189 for ignoring the command - the reward was mildly punishing the desired behavior.

    Context

    Sidewalking is inherently oscillatory: per-frame vy std was 0.177 while the tracking kernel width was sigma = 0.15, applied to the instantaneous value. A kernel-width scan showed widening sigma 0.15 -> 0.50 still loses (-0.14 -> -0.07): "指数核惩罚的是方差,而侧步天生带方差" (the exponential kernel penalizes variance, and side-stepping inherently carries variance). Modeling with measured parameters: replacing instantaneous vy with the mean over one gait cycle (0.5 s) flips the margin decisively (true sidewalk 1.888 vs ignore 1.281, +0.607), half-cycle is neutral (+0.006), two cycles adds nothing more. Explicitly flagged as extrapolation pending Isaac-side implementation. This also vindicated a previously dismissed external note (sigma too small) - right conclusion, different mechanism than claimed (variance, not gradient).

    Change

    Proposed fix recorded: change the tracked quantity from instantaneous vy to a one-gait-cycle running average; widening sigma alone rejected by the scan.

    Outcome

    Diagnosis complete and quantified; the C4 product shipped via feed-forward before the reward change was implemented, so the cycle-average fix remained a verified-by-model, not-yet-trained change.

    Mechanism

    E[exp(-(v-c)^2/sigma^2)] decreases with Var(v) even when E[v] = c exactly; a gait's phase-locked oscillation guarantees variance at the stride frequency, so instantaneous tracking rewards structurally prefer standing still at the command mean. Averaging over exactly one cycle removes stride-frequency variance while preserving command-following error.

    Conflicts

    The cycle-average fix itself is model-extrapolated ("⚠️ 这一条是外推,须在 Isaac 侧实装并复量后才能当结论") - the diagnosis is measured, the remedy untested in training at the time of writing.

    Applies when

    • tracking rewards for lateral/turn/any oscillation-carrying velocity
    • a verified behavior scores below the ignore-floor
    • choosing sigma for exp-kernel tracking terms
    “侧走时 vy 的逐帧摆幅 std = 0.177,而 track_lin_vel_y_exp 核宽 σ = 0.15,且作用在瞬时值上 … 真侧走(均值 66%,振荡 ±0.18)0.178 | 完全无视指令 0.189 … 真侧走的得分比无视指令还低。… σ 从 0.15 放到 0.50,侧走仍然吃亏 … 把跟踪目标从瞬时 vy 换成一个步态周期(0.5 s)的平均 vy:… 1.888 vs 1.281”
    train/C_LADDER_RUN.md § 3n. 二/三 Isaac 训练奖励为何一直坐在「无视底分」/ 修法不是放宽 σ
  • Under continuous 3-axis uniform sampling, pure straight-line walking is a zero-measure event the policy never trainedzero-measure-commands-need-mode-sampling
    Mechanism understoodomnicurriculumcurriculumobservation-honestygate-battery

    Enumerate the exact command points users will actually issue (straight, stop, in-place turn) and give each explicit probability mass via mode sampling with off-axes pinned to zero - never assume a continuous sampler covers its measure-zero subsets.

    Symptom

    "The robot drifts even in sim when told to walk straight" persisted across reward tunings - because with commands drawn as vx in [0.15,0.5] x vy ~ U(+/-0.2) x wz ~ U(+/-0.6), the event vy=0 AND wz=0 has probability zero: pure straight-line walking was never sampled even once.

    Context

    Restart evidence item #3: "纯直行是零测度点 … 'sim 里直行就漂'是分布的 必然,不是 reward 没调好" - the drift metric was legitimately drowned by commanded turning (v11's own comment self-documented this). The structural fix is discrete mode sampling: a custom ModeVelocityCommand that first draws a mode by share (stand/forward/back/turn/side/mixed), then draws values only on that mode's axes with all others pinned to exact zero - which is also what preserves single-variable discipline in the C ladder (native 3-axis uniform "采不出'离散模式桶' … 把 C1~C4 的单变量纪律直接毁掉"). The mixed mode later got an ellipsoid constraint rather than a cube for the same reason in reverse - corner combinations of a cube are unrepresentative extremes.

    Change

    Command generation moved from independent per-axis uniforms to mode-bucket sampling with pinned-zero off-axes (plus 20% rel_standing); acceptance likewise evaluates per mode.

    Outcome

    Straight-line behavior became a trained, testable mode instead of a measure-zero hope; the C ladder could add one mode per rung with provable isolation.

    Mechanism

    A policy optimizes expected reward under the command distribution; events of probability zero contribute nothing to the objective, so exact-zero-command behaviors (straight walk, stand, in-place turn) are only learned if the sampler gives them mass. Product-of-uniforms distributions concentrate mass on mixtures and give none to the pure behaviors users actually command.

    Applies when

    • a "simple" command (straight, stop) underperforms mixtures in sim
    • designing command distributions for velocity-tracking tasks
    • a ladder needs per-mode isolation for attribution
    “纯直行是零测度点:最终 command 为 vx∈[0.15,0.5] × vy∈U(±0.2) × wz∈U(±0.6) 连续均匀,vy=0∧wz=0 从未被专门采样 —— "sim 里直行就漂"是分布的必然,不是 reward 没调好 … Isaac 原生 UniformVelocityCommand 是三轴各自 uniform,采不出"离散模式桶"”
    train/OMNI_V0_SPEC.md § 0. 为什么从零 (3) / 三件前置 (1)
  • Add a single-point-suspension test to acceptance - the ground is a free stabilizer that hides divergencesuspension-probe-removes-free-stabilizer
    Mechanism understoodwalkgate-batterygate-batteryreal-acceptancesim2sim

    Include at least one acceptance condition that strips the environment's free stabilization (suspension, or equivalent) - the sim-passing policy that fails on hardware is often failing a condition the battery never posed.

    Symptom

    walk_v5 looked healthy in every on-ground sim test yet diverged on the real robot - the acceptance battery had never measured a condition that would have revealed it.

    Context

    The battery gained a single-point-suspension probe (robot hung, feet free): measure torso tilt while the policy runs without ground contact. v5 scored 45.9 deg mean tilt suspended - wildly unstable - which the file calls "最灵敏的失稳探针(拿掉地面这个免费稳定器)": ground reaction forces passively stabilize a marginal policy, so on-ground metrics saturate long before the policy's internal balance is actually sound. v6 halved it (23.0 deg, target <10 deg) - progress visible on a scale where on-ground numbers showed nothing.

    Change

    Suspended-tilt added as a standing acceptance row; run under the honest contact parameters battery (accept_v2 with measured condim 4 / torsional friction 0.035), under which v5 correctly FAILS in agreement with the real robot.

    Outcome

    The sim battery's verdict on v5 flipped from pass to fail, matching hardware; suspended tilt became the discriminating metric between v5 and v6 (45.9 vs 23.0 deg) when ground metrics differed little.

    Mechanism

    Contact with the ground closes a stabilizing feedback loop the policy gets for free; removing it exposes the policy's own attitude control authority. A metric measured only in the assisted condition cannot rank policies by the unassisted quantity that hardware will actually demand during perturbations and flight phases.

    Applies when

    • sim acceptance passes but hardware diverges
    • designing an acceptance battery for a legged robot
    • two candidates tie on ground metrics
    “单点吊那条是最灵敏的失稳探针(拿掉地面这个"免费稳定器"), v5 在地上一切正常却在真机发散, 就是因为验收从没测过这个工况。”
    train/WALK_V6_MINIMAL.md § 5. 验收
  • Mirror augmentation over an asymmetric default injects systematic error - symmetrize the default first and verify the transform bit-exactmirror-augmentation-needs-symmetric-default
    Mechanism understoodwalkobservation-designobservation-honestycurriculumprocess

    Before enabling any symmetry augmentation, make every constant inside the observation encoding exactly symmetric, and validate the mirror transform against forward kinematics to machine precision - an unverified augmentation is a new error source, not a regularizer.

    Symptom

    Mirror data augmentation was about to be added while both default poses (standing_pose, walk nominal_pose) were asymmetric - stale hand-tuned compensations from before a ground re-calibration, with hip_yaw differing 2.40 deg between sides and the foot soles actually tilted (pitch 2.88/1.35 deg, roll -2.47/+0.25 deg).

    Context

    The observation encodes joint_pos_rel = q - default. Under mirroring q_l -> -q_r, the relation (q-default)_l -> -(q-default)_r holds only if default_l = -default_r; with an asymmetric default, augmentation produces observation pairs that are NOT mirror images, i.e. "default 不对称时做镜像增强会引入系统性错误,比不做还糟" (worse than not doing it). The fix: adopt model geometric zero as standing default (MuJoCo FK verified: sole pitch/roll exactly 0, asymmetry 0.00 deg) and a symmetric crouch for walk (hip -0.25/knee -0.5/ankle -0.25 satisfying hip - knee + ankle = 0 to keep soles flat). The mirror transform itself was verified bit-exact before use: pseudovector vs polar-vector sign patterns (ang vel [-1,1,-1], gravity [1,-1,1], cmd [1,-1,-1]), joint swap-and-negate; FK check that left-foot pose under q equals the mirror of right-foot pose under mirror(q), measured error 0.00e+00.

    Change

    Defaults symmetrized first (with init heights recomputed by FK), stand policy retrained on the new default so both policies share one default; augmentation enabled only after the FK mirror test passed.

    Outcome

    stand_v1 achieved exact left/right pairing (l_knee -0.1013 / r_knee +0.1013), six-pair asymmetry 0.0 deg, height fluctuation 7 -> 1 mm, 33% less mean |action|.

    Mechanism

    Augmentation asserts an equivariance of the observation encoding; any asymmetric constant inside the encoding (the default) breaks the asserted symmetry, so the augmented data teaches a false invariance. Verifying the transform against FK geometry tests the assertion end to end, independent of the training stack.

    Applies when

    • adding mirror/symmetry augmentation to locomotion training
    • defaults or trims were hand-tuned per side at any point
    • observations are expressed relative to a default pose
    “观测里 joint_pos_rel = q − default。镜像下 q_l → −q_r,要让 (q−default)_l → −(q−default)_r 成立,必须 default_l = −default_r。default 不对称时做镜像增强会引入系统性错误,比不做还糟。… 位置误差与姿态矩阵误差实测均为 0.00e+00。”
    train/RETRAIN_v2.md § 2. 前提:default 姿态必须先对称化(不是可选项) / 3. 镜像变换
  • Every gate on a penalty is an exit - to stop paying a stance tax the policy parked 2 deg outside a 30 deg uprightness gate (lunging), and, from scratch, just under a height gate (crouching); a positive band and an always-on guard fixed bothpenalty-gate-is-an-escape-hatch
    Replicatedrecoveryreward-shapingreward-shapinggate-battery

    Never gate a penalty on a state the policy can leave by getting worse; use always-on guards for what must never happen and positive, gated bands for what you want, and count every gate on a penalty as one more escape route to check in the logs.

    Symptom

    V2.9 added stance_width_task = relu(0.34 m - foot spacing) x standing gates, weight -10. Three checkpoints scored 0% on acceptance: standing height reached, feet on the ground, angular rate low, but the torso leaned 32.5/32.3/31.9 deg in a fore-aft lunge. The training dashboard read "tax paid off, base_height at full value".

    Context

    The penalty was gated by uprightness (tilt < 30 deg) and standing height; its tax had no time gate (500 steps x -1.65) while the standing income sat behind a 3 s zero gate (~325 steps), so leaning just past 30 deg lost a little gated income and saved the whole tax. base_height has no upright gate, so the lunge still collected it. The first metric also measured full horizontal spacing, so a staggered lunge counted as "wide".

    Change

    Two laws written down: a penalty may only carry gates the policy cannot escape by getting worse (make it an always-on guard) or it becomes a positive band ("not earned" is not "escaped"); and width is measured laterally in the base yaw frame. From scratch (V3.1 P1) with the lateral metric but the same gates, the policy parked just under the height gate instead (base_height 1.176/1.5, h ~ 0.30 m against a 0.3264 gate), the beta curriculum never advanced in 1,700 iterations, and the run was stopped early. P1b flipped the penalty into a positive band +2.0 x clamp(lateral / 0.34) x standing gates; P1c added yaw_guard = -5 x relu(|hip_yaw| - 30 deg), always on, no gate, no exemption.

    Outcome

    P1b: lateral stance 0.364 m, all four categories 100%, MuJoCo mu 1.0/0.4 both 100%. P1c: all six acceptance criteria passed for the first time on the line (hip-yaw saturation 1.2%), the guard's tax converging to -0.016 (almost never touched).

    Mechanism

    A gated tax that is not paid is saved, so the policy moves to the cheapest state just outside the gate; a gated income that is not earned is simply lost, so a positive band has no exit. HoST's style penalties are ungated or binary - the same law seen from the other side.

    Applies when

    • adding a penalty multiplied by an uprightness, height, phase or contact gate
    • a policy settles just beyond a gate threshold
    • training metrics look paid-up while acceptance collapses
    “**罚项的门 = 策略的逃生门**。带直立门的负项可以靠"变得更差"(倾出门外) 全时免税;HoST 的 style 罚全部无门控/二值恰是同一律的反面实证。修律: 负项只许挂"变差逃不掉"的门(上限护栏 always-on),或改正向 band (收不到 ≠ 逃掉)。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §46 定案(血统内第四败 + 两条新律)
  • The action-delay was implemented as lerp - beyond one step it extrapolated BACKWARD, so a whole lineage trained on a fictitious actuatorlatency-lerp-reverse-extrapolation
    Mechanism understoodinfraactuator-modelingactuator-modelingplant-calibrationsim2sim

    Unit-test plant-model code (delays, filters, randomizers) against hand-computed truth across its FULL configured range, not just the nominal case - a delay must be a queue, and any interpolation used outside [0,1] is a silent plant corruption that training will faithfully absorb.

    Symptom

    Every walk/stand model up to v11 had been trained on a silently wrong plant: the action-latency implementation lerp(cur, prev, lag) is only an interpolation for lag <= 1 - at lag 3 it computes 3*prev - 2*cur, a REVERSE extrapolation. With the configured (0, 0.06) s at 50 Hz (lag in [0,3]), about 2/3 of environments were adapting to actuator dynamics that do not exist.

    Context

    Listed as evidence item #1 for the full restart: "全部旧模型训在错误 plant 上" and the head suspect for the real robot's wild kicking. The fix replaced it with a true FIFO delay line (commit 7f13793) plus its own regression test (tests/test_action_latency.py) - but every exported ONNX predated the fix, which is part of why the lineage was frozen rather than patched.

    Change

    Delay implementation rewritten as an honest FIFO with unit tests; the restart baseline trained on the corrected plant from day one.

    Outcome

    A generation-scale training investment was revealed to have a corrupt plant underneath; the class of bug (plausible-looking math that silently changes meaning outside its valid range) got a permanent test.

    Mechanism

    lerp(a, b, w) leaves the segment for w > 1; used as a delay it fabricates high-gain inverted dynamics precisely in the largest-delay draws, so the policy learns compensation for an actuator that cannot exist - and DR then trains robustness to the artifact rather than to reality. No training metric can catch this: the sim is self-consistent, just wrong.

    Applies when

    • implementing or auditing action delay / filtering in a trainer
    • a lineage behaves as if compensating dynamics nobody modeled
    • deciding whether old checkpoints are salvageable after a plant bug
    “动作延迟旧实现 lerp(cur, prev, lag) 在 lag>1 时是反向外插(w=3 → 3·prev−2·cur),配置 [0,0.06]s@50Hz 即 lag∈[0,3],约 2/3 env 在适应不存在的执行器动态。7f13793 已换真 FIFO,但所有 ONNX 均训于修复之前 —— 真机"乱踢"的头号嫌疑。”
    train/OMNI_V0_SPEC.md § 0. 为什么从零 (1)
  • Anchoring the action on the measured joint angle (target = q + beta*a), with a beta curriculum down to tau_limit/kp, bounded torque by construction, removed the re-falls and later stood the robot up on hardwarebeta-anchored-action-target
    Mechanism understoodrecoverytraining-runactuator-modelingcontract-freezecurriculum

    For large-motion skills on position-controlled actuators, bound the action relative to the measured joint angle with a per-joint authority of tau_limit/kp, curriculum the authority down from full range, keep the curriculum state out of the observation and pin acceptance at the deployed authority - and make the deployment code refuse to run the anchored contract without a measured q.

    Symptom

    The V0 full-range absolute action produced violent targets; the V1 command-anchored rate limit made standing oscillate. Both failure modes came from how the action becomes a target.

    Context

    V2.0 (user approved, from scratch): BetaAnchorJointPositionAction, target = q_measured + beta_j(m)*a, memoryless per step. beta_j(m) = floor + m*(beta0 - floor), beta0 = the contract half-range (m = 1 reproduces V0 authority), floor = min(tau_limit/kp, beta0): hip_pitch 1.309 -> 0.40, knee 1.047 -> 0.40, hip_yaw -> 0.917, the other joints unchanged - the tightening lands exactly on the joints the V0 torque account convicted. m drops 0.1 per step when a standing-share EMA exceeds 0.35. beta is NOT in the observation, so the 45-dim contract is untouched; acceptance is pinned at m = 0 because the Python curriculum state is not saved in the checkpoint. The deployment chain got a new profile (recovery_v2: action_anchor current_q, explicit per-joint beta written into the contract, independent of the gain profile), and policy_io raises if q is missing rather than silently falling back to the absolute contract; the old profile's check reproduced its pre-change deviation bit for bit.

    Change

    New action term and beta curriculum; later the RS06 floor was lowered 0.40 -> 0.30 -> 0.25 (kp*beta 7.5 N*m) and the stamped deployment profile was synced to 0.25.

    Outcome

    First acceptance at m = 0 (v2_0b): re-falls 0% in every category, the torque gate passed for the first time on the line (worst 69.9%), knee jitter 0.004; supine 98.8 / side 88.8% with prone and mid still failing (fixed by the conditional pull curriculum). MuJoCo showed demand at or under the limits (hip_pitch 11.7/12 against V0's 26.8). Lowering beta cut impact (hip_pitch demand 9.7 -> 8.5 N*m) but barely slowed the get-up - it had become coordination-limited. Enabling the policy moves the target only +/-beta around the current pose, so there is no homing fling; the 08-11 real get-up and the later v3_1p1c both run on this contract.

    Mechanism

    kp*beta caps the proportional torque in a single step with no build-up delay and no memory, giving both a hard impact bound and full balance bandwidth.

    Applies when

    • a skill needs full joint range but hardware torque limits are low
    • absolute position targets cause impacts or saturation
    • changing the action semantics of a contract that deployed policies share
    “**动作项** `BetaAnchorJointPositionAction`:`target = q_实测 + β_j(m)·a`, 逐步无记忆 … **Play/验收钉 m=0(= floor = 部署档)**:python 课程状态不进 checkpoint, Play cfg 显式 `beta_m_start=0` … 判读:**结构赌注兑现** —— 站姿零再摔 + 力矩账首过(kp·β 封顶按构造)”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §33 V2.0 预注册(2026-08-10,用户点头开工):β 锚定动作空间,从零训
  • Three hardware accounts locked the run design point - and the knee's real speed ceiling is tau_limit/kd, not the firmware limitfeasibility-accounts-lock-design-point
    Mechanism understoodrunplant-calibrationplant-calibrationactuator-modelinghardwareprocess

    Before opening a dynamic-gait training line, compute the full account set - tau_limit/kd effective speed ceilings, joint ROM under the intended reference geometry, and thermal RMS at the duty cycle - and let the accounts lock the design point; move only to pre-registered in-table alternates, re-running the accounts first.

    Symptom

    The run line was believed to require a firmware raise of the RS06 speed limit (10 rad/s) as a hard precondition, and the feasibility script's motor-envelope scan had marked 80/100 mm foot-lift cells "physically feasible".

    Context

    Three added accounts re-decided everything. (1) Damping tax: in MIT mode tau = kp*(q_des-q) - kd*qd, so sustained rotation is capped at tau_limit/kd = 12/1.5 = 8 rad/s - below the firmware's 10; at peak speeds 6.7-7.9 rad/s the damping term alone eats 10.1-11.9 N*m (84-99% of the torque limit). "提固件 limit_spd 越不过这道税 —— 它是 kd 与限扭的比,不是固件旋钮." (2) Joint ROM: the feasibility script had checked motor envelopes but NOT joint range - the ankle-pitch ROM caps 1:2:1 leg-shortening lift at 62 mm (soft) / 77 mm (hard), so the 80/100 mm "feasible" cells were voided; also firmware-independent. (3) Ankle thermal: duty 0.40 puts ankle RMS at 87% of continuous rating (0.35 -> 93%); long-period big-stride cells hit both ankle torque peak and heat. Verdict: firmware raise DEQUEUED (50 mm design point needs knee 6.7-7.3 < the 8 rad/s effective ceiling < firmware 10); vel_limit stays 10 so sim == robot. The three accounts uniquely lock the design point - 50 mm lift / T 0.60 s / duty 0.40 - "三笔账 唯一锁定,不是调参空间", with pre-registered alternates allowed only inside the table and only after re-running the accounts.

    Change

    Design point frozen from accounts; hardware precondition reversed by arithmetic rather than by test; reference amplitude (0.84 rad = FK inverse of 50 mm) derived, per-joint action scales sized to the required travel (knee 0.9, hip_pitch 0.6, ankle deliberately NOT amplified - hard limit is adjacent).

    Outcome

    A firmware work item left the critical path; an infeasible region of the design space was closed before any training; the remaining risk (knee tracking lag from the damping tax) was pre-registered with its own criterion and in-table fallback (duty 0.35) - "这不是'奖励没调好', 是 plant 账".

    Mechanism

    PD actuators in MIT mode pay kd*velocity out of the same torque budget that tracks position, so the effective speed ceiling is a ratio of configuration constants, invisible to firmware settings; and feasibility is the intersection of ALL constraint families (torque envelope, joint ROM, thermal RMS) - a scan that omits one family certifies impossible cells.

    Applies when

    • planning running/jumping or any high-rate gait on PD actuators
    • a firmware or hardware upgrade is assumed as a training precondition
    • a feasibility scan covers motor limits but not ROM or heat
    “膝的有效速度顶 = τ_limit/kd = 12/1.5 = 8 rad/s,不是固件的 10。… 提固件 limit_spd 越不过这道税 —— 它是 kd 与限扭的比,不是固件旋钮。… 可行性脚本只查了电机包络没查关节 ROM —— 其 80/100mm 的"物理可行"格作废。… 判决:RS06 提固件对 run v0 不是前置,出队”
    train/RUN_V0_SPEC.md § 1. 硬件账判决 / 2. 步态设计点
  • The 1.5 Hz step-frequency gate was retracted - the author had misread his own actuator data, and the limit fought pendulum dynamicsgate-threshold-retracted-frequency
    Mechanism understoodwalkgate-batterygate-batteryactuator-modelingprocess

    Every gate threshold must cite its measurement and survive a first-principles sanity check; when a gate keeps failing otherwise healthy behavior, re-derive the threshold from the raw data before enforcing it again - and retract wrong gates in writing.

    Symptom

    An acceptance criterion "step frequency <= 1.5 Hz" kept failing healthy policies (walk_v4 at 2.33 Hz), and an earlier attempt to force slower stepping (v3) had killed stepping altogether.

    Context

    The threshold had been derived from the author's own actuator frequency-response measurements - but re-reading the raw table showed the misread: amplitude ratio at 2.0 Hz is 0.88 (knee) / 0.83 (ankle), acceptable; the genuinely bad point was walk_v2's 3.8 Hz at 0.58. Mechanism check agreed: the leg as a compound pendulum (L ~ 0.30 m) has a natural frequency ~1.1 Hz, swing half-period 0.45 s - the observed 2.1-2.4 Hz sits near where the leg wants to swing, and forcing 1.5 Hz "是跟摆动动力学对着干" (fights the swing dynamics). The frequency definition itself was pinned by two independent methods (contact-event counting 2.39 Hz vs FFT 2.33 Hz, agreeing): reported numbers are cycle frequency = steps per leg per second.

    Change

    Gate retracted in writing: "步频 ≤1.5 Hz 应删除或放宽到 ≤2.5 Hz"; frequency definition standardized before entering any config.

    Outcome

    walk_v4/v6's 2.1-2.4 Hz reclassified from disease to normal; the v3 failure got its probable explanation (suppressing stepping to meet a wrong gate).

    Mechanism

    A gate is only as good as the measurement and the reading behind it; thresholds inherited from a misread plot become invisible design constraints that later training obeys at real cost. Cross-checking a threshold against first-principles dynamics (pendulum frequency) is a cheap way to catch such misreads.

    Applies when

    • an acceptance threshold repeatedly fails policies that look healthy
    • thresholds were set from a single person's reading of raw data
    • a forced compliance with a gate degrades the behavior it guards
    “我当初依据自己测的执行器频响定的,但看错了区间。… 2.0~2.4 Hz 的幅值比 0.83~0.88 是可接受的;真正不行的是 walk_v2 的 3.8 Hz。… 腿按复摆算(L≈0.30 m)自然频率约 1.1 Hz … 把它压到 1.5 Hz 是跟摆动动力学对着干(walk_v3 把迈步压没了,可能正是这个原因)。”
    train/WALK_DIAGNOSIS.md § ② 撤回"步频 ≤1.5 Hz"这条验收标准 —— 是我定错了
  • The run policy never left the ground and fell in the second simulator from the frontal plane - its DR (gains and latency only) covered the actuator axis, not the frontal-plane contact and inertia disturbances the doubled stride amplified; "is DR on" is the wrong questionthin-dr-judged-by-channel-coverage
    Replicatedrundr-tuningdomain-randomizationsim2simattribution

    Judge a DR recipe by whether its randomized terms cover the channel where the skill can lose stability, not by whether DR is enabled; when a new skill lengthens single support or enlarges motion in one plane, add disturbances in the plane it destabilizes before training.

    Symptom

    run R1 (6,000 iterations, 78 min): no flight phase ever appeared, and every one of 13 checkpoints failed the eight-gate MuJoCo smoke. In Isaac: zero terminations in 6,000 iterations, 4.2 deg tilt. In MuJoCo at delay 2: 1/6 survived, falls within 1.9-6.2 s at 50.8-58.7 deg, the most saturated joints all roll joints.

    Context

    The run contract doubled sagittal travel (knee action scale 0.9, knee swing peak 1.14 rad) with a 0.60 s period and 0.40 duty - long single support - while roll/yaw scales were deliberately left at 0.5. DR copied the s1e recipe: kp/kd (0.9, 1.1) and latency on; mass, COM, joint friction and push all off; ground friction pinned at (1.0, 1.0). Flight was read two independent ways: Isaac's per-foot contact reward stayed 0.845-0.857, never above 0.87 - the arithmetic ceiling of a gait with zero flight - and 30 of 36 MuJoCo seeds had flight fraction exactly 0 (the nonzero six were all tumbling falls). Foot lift itself worked (46-59 mm against a 50 mm design point): the walk-era "not enough travel" failure did not recur.

    Change

    Verdict FAIL, with the pre-registered first knob (exploration noise 1.0 -> 1.2) explicitly rejected as aimed at a different axis. The lesson was generalized and applied at the next line's design review: the one-leg spec made push, body mass, base COM and friction DR mandatory for its permanent single support and banned the thin recipe.

    Outcome

    The run line did not continue past R1 in the sources. The one-leg V0 with the wider DR passed its friction-variant gate (mu 0.4 and 1.2) inside a 40/40 acceptance.

    Mechanism

    Randomizing gains and latency covers the actuator's axis; a skill whose failure lives in frontal-plane contact and inertia needs randomization on that channel (push, mass, COM, friction), or the trainer's exact plant becomes the only one the policy can stand on - the omni_s1 transfer trap a second time, this time with DR switched on.

    Applies when

    • a policy is flawless in the trainer and falls immediately in a second simulator
    • reusing a DR recipe from a skill with a different support pattern
    • failures concentrate on one axis (roll, yaw) the DR does not touch
    “**机理**: 矢状面行程翻倍 (膝摆动峰 1.14 rad) + T 0.60 + duty 0.40 的长单支撑, 把额状面扰动放大了一个量级; 而 roll/yaw 通道按 §3 **刻意没有放大** (仍 0.5), DR 又是 s1e 复刻的薄配方 (mass/COM/关节摩擦/push **四关全关**, 地面摩擦钉死 (1.0, 1.0))。 … 说明**薄 DR 的判据不能只看"有没有开 DR"**, 要看**开的那几项 是否覆盖失稳所在的通道** —— kp/kd 与延迟是执行器轴向的, 对额状面接触/惯性 扰动零覆盖。 … 0.87 正是「零腾空的走路步态」的天花板算术”
    git:Lucen V2@origin/run-line:train/README.md § run R1 FAIL (2026-08-09, run 21-30-30_run_r1): 腾空零, 但病根在额状面不在探索
  • Raising a command bucket's share does not strengthen its per-state gradient - it only starves the other modesbucket-share-is-not-a-gradient-lever
    Mechanism understoodomnicurriculumcurriculumreward-shapingdomain-randomization

    When a skill is not learning, first prove its per-state signal is nonzero (ignore-floor and probe checks); only rebalance sampling shares to fix genuine sample starvation, and account the regression risk to the diluted modes before doing it.

    Symptom

    Sidewalk was not learning, and the reflex proposal was to give the side bucket a larger share of sampled commands.

    Context

    The C4-redo3 rung explicitly kept the 20/40/20/20 bucket (stand/forward/turn/side) with the reasoning written out: PPO computes advantages per state, so bucket proportion does not change the per-state gradient of side states; at 4096 envs x 20% x 24 steps the rollout already contained ~19.7k sidewalk states - sample count was not the bottleneck. And the cost side was already measured: cutting forward from 60% to 40% had made vx+0.30 die at +400 in an earlier run - more cuts would only collapse it sooner.

    Change

    Bucket proportions held constant across the entire C4 redo series; the actual bottlenecks (metric frame bug, reward variance penalty, exploration form) were pursued instead.

    Outcome

    Sidewalk was eventually fixed with zero bucket changes (feed-forward delivery, +100 iters); forward/turn skills never suffered starvation-induced regressions during the redo series.

    Mechanism

    Policy-gradient credit is assigned per visited state; oversampling a mode multiplies its states in the batch but not the informativeness of each, so if the per-state gradient is ~0 (behavior unreachable or reward indifferent), N times zero is still zero - while the displaced modes genuinely lose data and regress.

    Applies when

    • proposing to oversample a failing task/command mode
    • a majority mode regresses after share rebalancing
    • budgeting env count vs mode share for a multi-skill policy
    “比例不动:PPO 逐状态算优势,桶占比不改变单状态梯度;4096 env × 20% × 24 = 每 rollout 已有 1.97 万个侧走状态,样本数不是瓶颈;而 forward 60%→40% 已实测让 f30 在 +400 处死掉,再加码只会更早塌。”
    train/C_LADDER_RUN.md § 3i. 桶 20/40/20/20 不动(比例不动)
  • The walk phase machine structurally cannot express flight - rebuild the representation for duty < 0.5, teach flight with a mask tax, never a cliff bountyphase-machine-structural-limits
    Mechanism understoodrunreward-shapingreward-shapingcurriculumgate-battery

    When a new gait changes the contact pattern's structure, audit whether the phase/mask representation can express it and rebuild the representation if not; teach the new contact pattern with graded mask-mismatch pressure and count it in acceptance with artifact-proof definitions (minimum segment length), never with cliff bounties.

    Symptom

    Running requires both feet airborne, but the walk-era phase machine switches legs by the sign of sin(phase) - with duty < 0.5 the two swing windows must OVERLAP during flight, which a sign-switching representation cannot express at all.

    Context

    The run phase machine was re-architected rather than patched: per-leg phases (left = phi, right = phi+0.5 mod 1) with leg_phase < duty defining stance, aligned to the walk sin convention at duty=0.5 so the machines agree where their domains overlap. Flight is taught by the SAME mechanism that once cured foot-dragging, direction reversed: in the two planned flight windows the contact mask is (0,0) and feet_contact_number_duty charges -0.3 per foot still on the ground - a mild ~0.16/step tax, deliberately NOT a cliff: "悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文". The reference shape (half-sine bump over swing progress) is zero at window boundaries by construction, eliminating the clearing-window that the C4 probe measured to cost 13-19% on non-sinusoidal references. The shape self-check ("零代价选项是什么") was run on three behaviors: standing pays ref everywhere (known cmd=0 stepping risk, booked), walking pays only the flight-window tax, proper running collects full marks.

    Change

    New rewards.py run section (leg_phase_duty / stance_mask_duty / ref_run / clearance_run / contact_duty) with the walk versions untouched byte-for-byte; flight acceptance metric defined with a segment-length floor (>=40 ms to count) so numeric contact flicker cannot fake flight.

    Outcome

    Flight became expressible and taught by a calibrated mild pressure; the walk lineage's phase code stayed frozen as its own contract.

    Mechanism

    A phase representation defines which contact patterns exist in the reward's vocabulary; duty cycling below 0.5 introduces states (double-flight) outside a half-period sign convention's language, so no weight tuning can teach them. And rare desirable events taught by cliff-shaped bounties invite hacks (jumping in place); the graded mask tax prices the planned pattern without creating a jackpot.

    Applies when

    • extending a walking stack to running/jumping (duty < 0.5)
    • a desired contact pattern never appears despite reward increases
    • defining flight/contact acceptance metrics
    “duty<0.5 时摆动窗 (1−duty)T > T/2,两腿摆动窗在腾空段重叠 —— sin 符号切腿的机制结构上表达不了"双脚同时在空中"。… feet_contact_number_duty 对"还踩着地"持续 −0.3/脚 —— 与 walk 治拖地同一机制,方向相反。不腾空的税 ~0.16/步 … 梯度温和不构成悬崖(悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文)。”
    train/RUN_V0_SPEC.md § 5. 相位机设计
  • When the eval proxy has a known systematic bias, gate on within-proxy differences, not absolutesrelative-metrics-survive-proxy-bias
    Mechanism understoodwalkgate-batterygate-batterysim2simattribution

    Where the evaluator is known-biased, design gates as within-evaluator contrasts (left vs right, A vs B, pre vs post) that cancel common-mode error; reserve absolute thresholds for quantities whose proxy calibration has been checked.

    Symptom

    MuJoCo systematically overestimated yaw turn gain (1.58/2.45 where Isaac measured 1.04/1.02), and friction alignment recovered only part of the gap (mu 1.0 -> 0.6 pulled it to 1.43/2.27) - so any absolute turn-gain acceptance threshold in MuJoCo would be judging the proxy's bias, not the policy.

    Context

    The acceptance criterion was rewritten to use only the left/right difference of the turn gain (standard: <20%): both directions pass through the same biased proxy, so the bias largely cancels in the difference while the policy's chirality - the thing being gated - survives. The absolute-gain row was dropped: "转向只判左右差,不判绝对值". A parallel task was still opened to align MuJoCo contact parameters to Isaac's (mu, restitution), since drift proved highly friction-sensitive (straight-line drift -65 deg at mu 1.0 vs +1 deg at mu 0.4) - bias reduction and bias-robust metrics proceeded together.

    Change

    Gate metric changed from absolute turn gain to left-right gain difference; proxy-alignment work scheduled separately rather than blocking acceptance.

    Outcome

    Turn acceptance became meaningful across proxy versions (v5 70% -> v6 19% difference measured the real improvement) while the absolute bias question was pursued without holding the ladder hostage.

    Mechanism

    A systematic multiplicative or additive proxy bias applies to both arms of a mirrored measurement; differencing (or ratioing) mirrored conditions cancels the common-mode bias to first order, leaving the asymmetry signal. Metrics built this way remain valid while the proxy is imperfect - which it always is somewhere.

    Applies when

    • a sim proxy disagrees with the trainer or hardware in absolute terms
    • writing acceptance thresholds for direction-paired skills
    • proxy calibration work would otherwise block a ladder
    “转向只判左右差,不判绝对值 —— MuJoCo 的偏航增益系统性高估(实测 1.58/2.45 vs Isaac 1.04/1.02),摩擦只能解释一部分(μ 1.0→0.6 仅拉到 1.43/2.27)。差值是相对量。”
    train/WALK_V5_SPEC.md § 6. 验收
  • mj_objectVelocity returns inertial-principal-axis frame - one API assumption poisoned eval and observations for a whole linebody-frame-velocity-api-audit
    Mechanism understoodomnisim2sim-gatemeasurementsim2simobservation-honestyattribution

    Verify every frame-sensitive API against a hand-computed truth (rotate raw qvel yourself, or command a known world velocity and check where it lands) before trusting any evaluation or observation built on it - especially when a model's inertial frame is rotated from its body frame.

    Symptom

    Sidewalk vy read ~0 under every condition; separately, whole-policy performance was mysteriously mediocre in sim2sim while training-side numbers looked fine. Four training rungs were declared FAIL partly on these readings.

    Context

    base_link's URDF inertial frame is rotated 90 deg about x relative to the body frame (iquat = [0.7071, 0.7071, 0, 0]). mj_objectVelocity(flg_local=1) rotates into ximat - the inertial principal-axis frame - not the body frame, and returns center-of-mass point velocity, not body-origin velocity. Consequences measured: the "vy" column was actually vertical velocity vz (walking at cmd 0.25: old reading +0.0093 vs true -0.0424); the angular velocity fed to the policy in sim2sim was [wx, wz, -wy] - a different quantity than Isaac and the real IMU provide. RMS check over 8 s of walking: y/z axes swapped between v6[:3] and the qvel truth.

    Change

    Fixed sim2sim and both probes to compute ang_b = qvel[3:6] and lin_b = xmat.T @ qvel[0:3] (identical quantity to Isaac's root_ang_vel_b / root_lin_vel_b), with a standalone reproduction script (frame_bug_repro_0809.py).

    Outcome

    Re-scoring the "failed" C4 lineage under correct coordinates reversed the verdicts: c4r4 checkpoints showed vy 80-126% tracking (old reading: +/-2%) and vx+0.30 at 91-95% where the old metric said 0/5 - the bad frame both mis-measured vy and, via corrupted policy observations, systematically depressed all measured performance. Final product passed 260/260 cells.

    Mechanism

    A simulator API's frame convention is part of the observation contract; when the model's inertial frame is rotated relative to the body frame, frame-agnostic use of a "local" velocity silently permutes axes. Feeding a policy an axis-permuted angular velocity is an observation corruption that degrades behavior everywhere, not just on the axis being studied.

    Applies when

    • building or auditing a cross-simulator evaluation harness
    • one measured axis reads near-zero under all conditions
    • sim2sim scores are inexplicably worse than training-side metrics
    • URDF/MJCF inertial frames are rotated relative to body frames
    “base_link 的 iquat = [0.7071, 0.7071, 0, 0] … mj_objectVelocity 用的是这个 … 喂给策略的 base_ang_vel 是 [wx, wz, −wy] —— MuJoCo 侧观测与 Isaac / 真机 IMU 不是同一个量;vy_mean 报的是竖直速度 vz —— 前进 cmd 0.25 时旧读数 +0.0093,真值 −0.0424。”
    train/C_LADDER_RUN.md § 3l. ⚠️ mj_objectVelocity 读的是惯性主轴系 / 3m. 一 bug 坐实
  • When training fails repeatedly, inject the target behavior open-loop - stop tuning rewards for an unverified behavioropen-loop-probe-before-reward-tuning
    Mechanism understoodomniattributionattributionprocesscurriculum

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop on the real plant/sim first, and only resume training once you hold a measured, safe, sign-verified target trajectory.

    Symptom

    Three sidewalk training rounds failed; hypotheses multiplied (exploration failure / wrong reference waveform / insufficient authority) with no way to pick between them by running more training.

    Context

    Instead of a fourth reward guess, the team wrote probe_side_ref.py: the candidate reference is injected open-loop on top of a frozen policy's output (bypassing PPO entirely), directly measuring "what happens if the robot literally does this waveform" - separating all three hypotheses in one experiment (5 seeds x 8 s per condition, several waveform families and gains). The probe immediately eliminated the authority hypothesis (full-amplitude execution, 5/5 survival) and localized the problem to the waveform/measurement side. The closing principle was written down after the saga: without a verified target behavior, tuning rewards is "在黑暗里试钥匙" (trying keys in the dark).

    Change

    Standing method: before opening another training rung on a failing skill, build an open-loop (or task-space) generator of the intended behavior, measure whether the physical system can express it and what it looks like - then train toward a verified, quantified target.

    Outcome

    The probe chain produced the verified waveform (reversed-sign triangle, half gain), quantified safe amplitude (tilt 8.2 deg at band top, foot distance clear of the wall), exposed the metric bug when probe and training disagreed, and ultimately supplied the feed-forward that made C4 pass in +100 iters.

    Mechanism

    Training couples exploration, reward design, and feasibility into one opaque outcome; open-loop injection cuts the loop and tests feasibility and waveform alone. A behavior demonstrated open-loop converts the remaining failure into a pure credit-assignment/reward question - and its measured trajectory becomes the reference itself.

    Applies when

    • repeated training failures on one skill with multiple live hypotheses
    • uncertainty whether the platform can physically express the behavior
    • a reference trajectory's shape/sign/amplitude is guessed, not measured
    “三轮 FAIL 之后不再猜,写 train/probe_side_ref.py 把参考开环注入到策略输出之上(绕过 PPO),直接量「照这个波形做会怎样」,一次分开三个假说:甲 探索 / 乙 波形 / 丙 权限。”
    train/C_LADDER_RUN.md § 3f. C4 真因定谳(开环探针) / 3k. 建议的下一步