Training Coach
Doctrine
A report may cite any of these as doctrine-N.
doctrine-1Contract freeze and fingerprint disciplineThe policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.
Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).
Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.
doctrine-2Attribution by resolved training params - never eval-override knobsCapability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.
Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).
Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.
doctrine-3PASS gates become constraints; FAIL gates become objectivesOnce a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.
Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.
Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".
doctrine-4One variable per ladder rung - counted against what the checkpoint sawA rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.
Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.
Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.
doctrine-5Pre-register risks, readings, and stop criteria before the ladderBefore a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.
Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).
Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.
doctrine-6Plant parameters are measured, never inventedEvery plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.
Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).
Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.
doctrine-7Sim2sim gate before sim2real - under deployment conditionsEvery checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).
Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).
Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.
doctrine-8Observation honesty - the actor's inputs are a hardware contractThe actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.
Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).
Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.
doctrine-9Reward economics are audited in realized currencyReward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.
Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).
Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.
doctrine-10The zero-cost option must be the desired behaviorFor every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.
Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).
Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.
doctrine-11Measurement discipline: independent referees, signs, distributionsA disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.
Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).
Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.
doctrine-12The deployment pipeline is plantIrreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.
Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).
Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.
doctrine-13DR budget is finite; its distribution is the measured supportRobustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.
Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).
Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.
doctrine-14Gates measure what hardware feels: posture, margins, stripped assistsAcceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.
Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).
Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.
doctrine-15Fork and root selection: recoverability, maturity, frozen rewardsChoose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.
Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).
Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.
doctrine-16Curricula: verified engagement, lineage counters, disease-phase gatingAutomatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.
Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).
Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.
doctrine-17Probe before training: feasibility first, hypotheses in tablesAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.
Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).
Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).
doctrine-18External advice is recomputed locally; values transfer as ratiosEvery external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.
Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).
Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.
doctrine-19Hardware sessions are scripted experiments, not tuning sessionsReal-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.
Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).
Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.
doctrine-20Close questions in writing; restart when the debt is structuralAudited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.
Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).
Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.
doctrine-21Name the quantity in the space it lives inA goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.
Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).
Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.
doctrine-22Continuation needs a live gradient; a release is chosen by a scanContinue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.
Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).
Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.
Experience cards
161 cards matching “removed-wall-returns-on-hardware”.
Knee swing collapsed because it directly trades against the slip penalty - price the conflict explicitly and clamp what reward cannot hold
knee-swing-vs-slip-pricingWhen a behavior collapses as another metric improves, look for the term pair trading them and set their price ratio deliberately (with escalation and reverse tripwires pre-registered); where the policy actively spends action budget to undo your target, stop paying more reward and clamp the target space structurally.
Symptom
Knee peak-to-peak swing collapsed across generations - v5 33 deg, v10 26-30, v10b 7-8, v11 6.5-8.6 - and rolling the clock back did not recover it, acquitting the clock; the collapse tracked the gated slip penalty instead: "屈膝与不打滑在当前奖励里直接对抗" - v10b's excellent 93 deg slip was purchased with knee amplitude.
Context
Reward-side flexion fixes had failed three times: raising reference amplitude backfired twice (v9/v11), and v11's deep-squat default was actively fought by the policy - it spent 0.68 of action budget pulling the squat straight ("被策略花 0.68 动作拉直反杀"). v12's design accepted the conflict as real and attacked on two tracks: (1) ECONOMICS - a direct knee_swing_amplitude reward (+0.3, target 0.55 rad, capped at 0.6/step = 55% of tracking), explicitly opposed to the slip penalty by design ("显式对立——这正是设计:v12 就是这场对抗的定价实验"), with an escalation ladder (K +0.3 -> +0.5, then slip -0.5 -> -0.3, one layer at a time) and a reverse tripwire (slip telemetry back at v10 levels -> slip weight to -0.8, accept ~20 deg knee compromise); (2) STRUCTURE - knee target bounds [0.2, 0.9] rad so full straightening is physically impossible (straightest 11.5 deg) and the 0.68 fighting budget is released. A bonus falsifiable prediction was attached: phase-lock strength tracks amplitude (v9_probe 48 deg locked 2.5 Hz; v11 low-amplitude 1.36 Hz unlocked), so if K works, hardware phase-lock should return - one change, two verdicts.
Change
knee_swing_amplitude reward + knee target clamp + pre-registered escalation/reverse levers; the failed reward-side-only approach retired.
Outcome
The lineage was frozen before v12 trained (strategic reset), but the diagnosis stands as the walk line's clearest example of two reward terms trading a behavior between them, with the pricing experiment and structural clamp fully designed and calibrated.
Mechanism
When two terms price opposite aspects of one motion (swing amplitude creates yaw momentum that becomes slip), the optimizer settles wherever the price ratio puts it - patching one side moves the equilibrium, not the conflict. Explicit pricing makes the trade a designed quantity; structural clamps remove the regions where the policy spends budget fighting the designer.
Conflicts
The pricing experiment (K vs slip) was designed and calibrated but never trained - the 2026-08-05 reset suspended v12; the collapse attribution table and the 0.68-action counterattack are measured, the remedy's效果 is untested.
Applies when
- one gait quality degrades in lockstep with another's improvement
- the policy visibly fights a default pose or reference
- repeated reward-side fixes for the same behavior have failed
“膝摆塌在 v10→v10b,头号嫌疑是门控滑移罚(四代实测膝 p2p:v5 33° / v10 26~30° / v10b 7~8° / v11 6.5~8.6°;退时钟没救回 → 非时钟)——"屈膝"与"不打滑"在当前奖励里直接对抗 … 奖励侧修屈膝已三败 … v11 深蹲 default 被策略花 0.68 动作拉直反杀”
train/WALK_V12_SPEC.md § 0. 定位 / 2. K —— 膝摆经济(与滑移罚的对偶) Friction DR was demoted after a measurement (94% success at mu 0.4 with no friction randomization) and promoted again when the action contract changed and mu 0.4 fell to 76% - DR priorities belong to a plant and contract, not to a task
friction-priority-re-measured-after-plant-changeRe-measure transfer along the friction axis for every new action contract or plant, not once per task; a DR priority settled under one action parameterization does not carry to the next.
Symptom
Getting up is all scraping and pushing against the ground, and training pinned friction at 1.0, so friction looked like the first thing to randomize.
Context
The MuJoCo gate on R0.5 (5 categories x 10 seeds x 4 friction levels) measured 100/100/98/94% at mu 1.0/0.8/0.6/0.4: degradation showed first as time (prone 3.2 -> 5.3 s), not failure, so friction DR was demoted and the DR budget earmarked for mass/COM. After the switch to the beta-anchored action space, V2.2 read 90/94/90/76%: mu 0.4 was now the weak row.
Change
V2.3 (single variable): friction DR static (1.0, 1.0) -> (0.2, 2.0), dynamic (0.15, 1.6), the HiFAR range keeping the base dynamic/static ratio; restitution untouched. Continued from v2_2.
Outcome
Isaac nominal 99.8% (DR did not hurt the nominal plant); MuJoCo 98/98/96/92% - mu 0.4 76 -> 92%, mu 1.0 back to R3.1's 98% with bounded torque.
Mechanism
How much a policy leans on friction depends on how it moves; the spec records that the sensitivity rose after the action contract changed but does not establish why.
Applies when
- changing the action space, gains or authority of an existing skill
- deciding which DR axis to spend the next rung on
- an earlier sweep justified leaving an axis unrandomized
“**μ 砍到 0.4(训练值的 40%)仍有 94%**,退化先体现在**用时**(prone 3.2→5.3 s) 而不是成败。μ≥0.8 完全无损。→ **§17 曾把"摩擦随机化提到 R4 第一项"当作优先 事项,这条实测把它降级了**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §21 MuJoCo 复核门 ② 摩擦依赖 The best checkpoint to SHIP is not the best checkpoint to CONTINUE FROM - maturity is capital against adaptation shock
root-maturity-vs-product-qualityDecide shipping points and fork roots separately: gates rank products, but a root candidate must prove itself by surviving a continuation under the next rung's shift (dual-arm if in doubt) - and prefer the more-trained point as root when product metrics conflict with maturity.
Symptom
A band re-audit found s1e-300 beat the incumbent root s1e-500 on nearly every quality gate (stepping 19/20 vs 13/20 with historically-best 26.9 mm swing, speed gate 14/20 vs 2/20, heading 26 vs 54 deg/20 s) - suggesting the root had been mis-picked and the younger point should take over.
Context
The dual-arm control settled it the other way: continuing the S2 PD rung from s1e-500 adapted smoothly (3/3 smoke throughout), while the b300 control arm (same config, from s1e-300) fell into a survival valley under the PD shock (+100 iters: 1/3 -> 0/3), never climbed out within budget, and its 800-iter product scored 13/20 survival - eliminated. Verdict: "幼年点自身指标再好也扛不住新 DR 适应冲击, 成熟度是本钱,s1e-500 根被数据背书" - a young point's own metrics, however good, do not survive new-DR adaptation shock; maturity is capital. The audit still yielded value: the band scan (200-1000, per-100) mapped the lineage's arc (200 dragging -> 300 peak -> 400+ decay -> 900+ drift blowout), and 300 remains the better PRODUCT answer for shipping-as-is questions.
Change
Selection doctrine split into two questions with different answers: best-product point (quality gates at the point itself) vs best-root point (survives adaptation shocks; more training age = more capital), each decided by its own evidence - and root claims settled by a dual-arm continuation test, not by point metrics.
Outcome
s1e-500 kept the root role with data behind it; the S2e ladder built on it passed rung after rung, while the b300 line was closed at the cost of one control arm.
Mechanism
Early checkpoints sit near sharp optima with less accumulated robustness structure; their headline metrics reflect the narrow training distribution, not resilience to distribution shifts. A continuation rung is itself a distribution shift, so the root property being selected for is shock tolerance - observable only by actually continuing, never by static gates.
Applies when
- a younger checkpoint outscores the current root on quality gates
- choosing the base for a robustification or command ladder
- a continuation run stalls in an early survival valley
“b300 对照臂 … PD 冲击下存活谷(+100 起 1/3→0/3),预算尽未爬出,800 档 20-seed 存活 13/20 出局——幼年点自身指标再好也扛不住新 DR 适应冲击,成熟度是本钱,s1e-500 根被数据背书”
train/README.md § omni_s2e_pd (b300 对照臂) / s1e 选点重审 An edge-triggered landing penalty missed the tail and fired after the harm - penalize overspeed continuously inside the contact window
penalize-tail-before-touchdownPenalties aimed at impact/violation events must (a) price the excess over a threshold, not the mean, and (b) be active on the approach (state-gated window), not triggered by the event - check your control rate can even see the event you are penalizing.
Symptom
The v7 landing penalty (vz^2 on the contact-force rising edge, weight -10) did not bite: landing-velocity 95th percentile stayed at 2.61 m/s against a 0.3 target.
Context
Two structural faults were identified: (1) it penalized the MEAN over sparse events - many soft landings dilute the occasional violent slam, while the damage (GRF peaks, motor peak load) lives in the tail; (2) it fired AFTER touchdown - at 50 Hz evaluation the rising edge is aliased by physics decimation, so the read vz is often the already-decelerated post-impact value: underestimated, and with no shaping gradient before contact. Replacement: continuous penalty while the sole is inside a height gate (h < 0.03 m): relu(-vz - 0.30) - only the excess over an allowed approach speed is penalized (tail only), and gradient exists for several frames BEFORE touchdown. The sole-height computation again subtracts the 0.0585 m link offset ("WALK_DIAGNOSIS 坑#1, 别再踩"); the edge-triggered version was kept as a diagnostic only.
Change
feet_landing_vel reformulated: edge-event vz^2 -> in-window relu(-vz - v_ok) with v_ok 0.30 (conservative vs the sqrt(L)-scaled human value ~0.19, to be tightened after passing), h_gate 0.03, weight unchanged -10.
Outcome
The failure analysis of the first form was written before the second was trained; the v_ok escalation path (0.30 -> 0.45 if the robot becomes afraid to land) was pre-registered in the risk table.
Mechanism
Sparse-event mean penalties optimize the average case while the constraint is a quantile; and any penalty evaluated only at/after a discrete event gives the optimizer no gradient along the approach trajectory that determines the event. A state-gated continuous excess penalty fixes both: it prices only violations and shapes the approach.
Applies when
- impact/landing penalties fail to move tail percentiles
- a penalty is triggered by contact edges at a coarse control rate
- designing constraint-style penalties for rare violent events
“罚的是均值路径:上升沿是稀疏事件 … 大量软着陆稀释偶发猛砸;而伤害在尾部 … 罚在触地后:50 Hz 评一次,上升沿被物理 decimation 混叠,读到的 vz 常是撞完已减速的值——既低估,又没有触地前的塑形梯度。”
train/WALK_V8_SPEC.md § 2. 改动 B — 落地惩罚改罚尾部、罚在触地前 Low-friction robustness traced to kd DR bandwidth, not friction training - by digging resolved params across 8 lineages, 3840 cells
kd-bandwidth-mu-law-attributionAttribute capability differences by tabulating every lineage's resolved training params and eliminating zero-variance and non-aligned columns first; never let an eval-side override knob serve as the explanation axis, and never write a mechanism into a law before it survives a targeted test.
Symptom
Lineages differed wildly in low-ground-friction survival, and the intuitive explanation - "some trained ground friction, some didn't" - was about to steer the ladder toward a ground-mu training rung.
Context
The attribution ran as a full parameter-vs-result cross: 8 lineages x 4 eval kd levels x 6 mu levels x 20 seeds = 3840 cells, with each lineage's RESOLVED training params dug out and compared item by item. First kill: all 8 lineages had ground mu pinned at (1.0,1.0) - zero variance - so low-mu differences cannot come from friction training at all. The only training parameter aligned with the mu score was kd DR bandwidth: narrow (<=0.24) lineages scored 19.9/19.5/19.5, wide (>=0.40) scored 17.1/15.2/14.6/14.2/12.8 - the two groups completely non-overlapping. Every rival was excluded item by item (kd center no; kp band no; COM small-beneficial non-driving; friction rung a clean double null 19.5->19.5 and 15.2->14.6; iteration count non-monotonic), and the one clean single-variable causal link confirmed it: the s2e-3 kd surgery (0.7,1.3)->(1.08,1.32) moved the score 17.1->19.5. Counter-proof against "each best at its own operating point": the narrow-band lineage evaluated OUT of band (18.2) still beat the wide-band lineage at its own band center (9.2). Two axes were ordered never to be conflated (the first attribution's own error): training kd bandwidth is a parameter axis / lineage property; the eval-side --kd-scale knob is a plant axis (more damping physically helps on slippery floors for ALL policies) - "plant 轴只能当部署缓解,不能当 归因". A tempting mechanism story ("drag vs step attractor") was tested and falsified, and explicitly kept OUT of the law: "机制未定, 不入定律".
Change
The planned ground-mu training rung was recommended closed ("建议 不开") in favor of a kd band-narrowing rung (0.8,1.2)->(0.9,1.1) centered on the deployed value - with a pre-registered risk that the law demands "bandwidth = measured dispersion" and the real robot's kd dispersion was not yet measured; if it exceeds +/-10%, narrowing sacrifices real coverage and the rung must yield.
Outcome
A whole training rung was deleted from the ladder by attribution alone (the second S2 pass dropped mu and push, 5 rungs -> 3); floor material became a deployment-selection input (mu <~0.6 -> deploy the kd1.2 gain profile) rather than a training target.
Mechanism
Cross-lineage performance differences must be attributed over the actual training-parameter table, not over eval knobs or plausible stories: eval knobs act on the plant for every policy (a physical effect), while lineage properties come only from training-time parameters. Zero-variance columns are free eliminations, and one clean single-variable rung is worth more than any correlation.
Applies when
- explaining why lineages differ on a robustness axis
- an eval-side knob (gain scale, power) changes results and invites misattribution
- deciding whether to open a DR rung for an axis never actually varied in training
“8 血统地面 μ 训练带全部钉 (1.0,1.0) 零方差,低 μ 差异与「训没训地面摩擦」无关,是 kd DR 带宽的副产物 … 宽 ≤0.24 → 19.9/19.5/19.5;宽 ≥0.40 → 17.1/15.2/14.6/14.2/12.8, 两组完全不重叠。… 训练 kd 带宽 = 参数轴/血统属性;评测部署 --kd-scale = plant 轴 … plant 轴只能当部署缓解, 不能当归因。… 机制未定, 不入定律。”
train/OMNI_V0_SPEC.md § 4. 地面 μ 鲁棒性 = kd DR 带宽的副产物 (2026-08-08) Every rung gets written stop criteria - hit any one, stop; tightened when priors say results should come fast
preregistered-stop-criteria-per-rungFreeze per-rung stop criteria (old-skill floors, new-skill deadline, oscillation signature, known pathology signatures) before training, stop on first hit - and shorten the deadline in proportion to how fast your mechanism says results should appear.
Symptom
Continued training past the point of degradation had previously destroyed capabilities (C1 trained away the root's backward skill); without hard stop rules, sunk-cost reasoning keeps runs alive too long.
Context
Standard rung protocol - eval_c_matrix every 100 iters at 5 seeds with a fixed criteria list, frozen before training ("开训前写死,事后不许改"): (1) stand or vx+0.15 at <=3/5 for two consecutive points; (2) vx+0.15 tracking <70% for two consecutive points (vs recorded root baseline); (3) oscillation signature - survival repeatedly crossing zero between adjacent checkpoints, stop on first occurrence, do not wait for confirmation; (4) new-skill-no-progress deadline (e.g. wz tracking <30% or still same-signed after iter 400); (5) a known lineage pathology signature (s1c arm-B: scatter -> half-recover -> collapse).
Change
Stop budgets scale with prior knowledge: when C4-redo4's feed-forward was already proven open-loop, the no-progress deadline was tightened from +500 to +100 iters ("前馈是开环就给 71~96% 的 … 若 +100 还没有 … 不值得再烧").
Outcome
Multiple rungs were stopped exactly on criterion (C4-redo criterion 4 at iter 1200; redo2/redo3 at absolute 900), converting each into a clean hypothesis test instead of a drifting run; no rung burned its full cap on a dead hypothesis.
Mechanism
Degradation of retained skills is a lineage-level injury that later training often cannot undo, so detection latency is capital loss; and a stop rule written before the run cannot be bent by hope. Tightening deadlines when priors predict immediate results converts "no result yet" into evidence against the mechanism.
Applies when
- starting any resumed/curriculum training rung
- deciding whether to keep training a run that shows early regression
- a mechanism-backed change should produce results immediately
“每 100 iter eval_c_matrix.py --seeds 5,命中任一即停 … 存活在相邻 checkpoint 间反复跨越 0(出现即停)… 收紧版:绝对 iter 900(= +200)时 vy±0.10 跟踪仍 <30% 即停”
train/C_LADDER_RUN.md § 3h. 预注册停梯判据(开训前写死,事后不许改) The median of a bimodal metric lands in the empty gap - check the distribution, and never judge swing on 3 seeds
median-hides-bimodal-distributionBefore quoting a median or mean, look at the distribution; report suspected-bimodal metrics as mode share plus per-mode ranges, use small-seed smoke runs only to screen trends, and size the seed count for decisions by the share resolution you need (here: 20).
Symptom
Years of "high swing variance" and undecidable 3-seed swing readings turned out to be one fact: the metric was bimodal all along - "历代 swing 高方差与 3-seed 不可判由此定性:一直在测双稳态系统" - and every median reported from it (e.g. 12.1 mm) described a value no seed ever produced.
Context
Concrete instances: s2e_pd-1400's 20 seeds split 2.6-4.9 mm vs 19.3-24.0 mm with zero seeds between; s1e-500 read 3.9 mm on seeds 0-2 but 23.1 mm median over 20 seeds; a 3-seed reading of 19.3 was logged as "double-peak optimism, lesson recurrence #4". The selection re-audit codified the sampling rule: "3-seed 的 swing 读数不可判点, 只能筛带,选点必须 20-seed" - 3 seeds may screen a band, only 20 seeds may pick a point.
Change
Swing (and any suspected multi-modal metric) reported as mode shares plus per-mode ranges instead of a bare median; 3-seed smoke numbers demoted to band-screening; all shipping/selection decisions moved to 20-seed batteries.
Outcome
The "swing debt" bookkeeping was reinterpreted as basin probability (see swing-bistability-damping-switch), and checkpoint selection stopped being whipsawed by which basin the first three seeds happened to fall into.
Mechanism
Central-tendency statistics presuppose unimodality; on a bimodal distribution the median tracks the mode SHARE, not any achievable behavior, and small samples alias the share entirely. Mode-aware reporting (share + per-mode stats) is the only faithful summary, and the needed sample size is set by the share resolution required.
Applies when
- a quality metric shows chronic high variance across seeds
- 3-seed smoke readings contradict 20-seed batteries
- reporting swing height, clearance, or any basin-prone metric
“中位数落在空档里,「swing 债 −11mm」实为「50% 概率掉进拖地吸引子」。历代 swing 高方差与 3-seed 不可判由此定性:一直在测双稳态系统。… swing 跨 seed 双峰 (500 在 seed0~2 只读 3.9mm, 20-seed 中位 23.1) —— 3-seed 的 swing 读数不可判点, 只能筛带, 选点必须 20-seed。”
train/README.md § swing 双稳态定性 / s1e 选点重审 A policy's gain profile is part of its contract - the one-leg policy needs per-joint gains the default profile lacks, and the manifest refused an evaluation under the default once; the recovery contract's beta was never stamped, a known gap not to repeat
gain-profile-belongs-in-the-stampStamp everything that defines the closed loop a policy was trained in - gains included - into its manifest, and make every consumer refuse a mismatch; a profile field that is not in the stamp is a silent misconfiguration waiting for an operator to forget a flag.
Symptom
A policy trained with hip_roll kp 80 and ankle_roll kp 60 behaves differently, or falls, under the default kp 20/12 profile - and the gain profile is a command-line flag an operator can forget.
Context
The one-leg line added a gain_profile field to the contract so the stamped manifest carries it; the spec's deployment note says the manifest guard blocks rl_default and that it had already bitten once in simulation (an evaluation run without the one-leg profile). The same spec states the general rule - any new profile field must be synced into the manifest builder - and names the counter-example: the recovery line's beta was never put into the manifest. The recovery line itself had decided that its anchored authority is computed from the base rl gains and written into the contract so it cannot drift with the gain flag, and that the older kp x 0.9 profile chosen in the V0 era does not match the beta contract and must not be used.
Change
Gain profile as a contract field checked at load; per-contract gain choices written into the run sheets.
Outcome
Evaluations and hardware runs of the one-leg policy run under rl_oneleg or are refused; the recovery beta gap stayed recorded as known.
Mechanism
A policy is trained against a closed loop whose gains are part of the plant; running it under other gains is an out-of-distribution plant, exactly like a wrong observation scale.
Applies when
- a skill introduces per-joint or skill-specific gains
- deployment gains are chosen by a command-line flag
- adding any new field to a policy profile
“增益档 `--profile rl_oneleg` 必须给 —— manifest 防线会拦 `rl_default`(sim 已咬合一次) … (recovery 的 β 未进 manifest 是已知缺口,不再复制)”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §9b AGX 真机手顺 要点 / §3 契约 A joint-velocity penalty meant to slow the get-up cut joint speed 16% and left the get-up time unchanged - the knob never moved the variable, so the idea it was meant to test stayed untested
dof-vel-penalty-is-not-a-pacing-knobBefore reading a result as a test of an idea, check that the knob actually moved the independent variable; velocity regularizers smooth a schedule they do not set, and a schedule driven by per-step task income moves only when that income's time structure does.
Symptom
The get-up took 0.6-0.9 s in Isaac with large torque demand; the user proposed getting up more slowly so less torque would be needed.
Context
The idea had support in the accounts: the acceptance bound is an upper bound of 5 s (5-8x margin), the quasi-static squat path peaks at 25% of the limits, and rolling over needs no momentum. R3.2 raised dof_vel from -1e-3 to -5e-3 as the single variable.
Change
dof_vel -1e-3 -> -5e-3 (child-run from R3.1).
Outcome
Get-up medians moved +0.02-0.04 s (noise); raw joint velocity -16%; torque demand median got worse (hip_pitch 46-48% -> 63-67%) as the new term competed with torque_headroom on the same joints; MuJoCo 98 -> 96%. Verdict FAIL on the knob, not on the idea, and the rung was not adopted. When pace was later attacked through the income's time structure (V2.5/V2.5b), the MuJoCo get-up moved into the 3.5-4.5 s design band.
Mechanism
The pace was set by base_height_progress paying for every step spent high (stand earlier, earn more); a velocity regularizer only smooths motion along the same schedule and does not change when the robot stands up.
Applies when
- trying to make a skill slower or gentler with smoothness penalties
- an experiment's primary metric did not move and a verdict is being written
- two penalties act on the same joints
“**关键判读:`dof_vel` 罚只把关节速度压了 16%,而起身用时一点没变。** 也就是说**这一级根本没有把"慢下来"这个自变量推动起来** —— 所以它**不构成对 用户假说的检验** … 起身节奏由 `base_height_progress` 的逐步计酬决定(早站起来就多 拿),速度正则只在同一条时间轨迹上把动作抹匀,不改变何时站起来。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §25 R3.2(dof_vel −1e-3→−5e-3,慢一点起身) FK-verify a borrowed reference's structure, then size its amplitude by the reference's job - it pins phase, the policy adds lift
reference-structure-fk-amplitude-divisionWhen borrowing a reference trajectory: verify its structural claim against your own kinematics (an invariant like flat-foot), assign it the phase-pinning job, and size amplitude low enough that the policy contributes the lift - moving toward a proven foreign value in halves, not jumps.
Symptom
walk_v4 had big knee swing (40-46 deg) but only 18-24 mm foot lift - amplitude without hip/knee/ankle phase coordination; later, walk_v5's real-robot swing ballooned to 73.6 deg (sim 55.7) with violent footfalls - amplitude over-driven by the reference.
Context
Structure first: Humanoid-Gym's 1:2:1 hip:knee:ankle reference was verified on the local model before adoption - the ratio exactly satisfies the locally derived flat-foot constraint hip - knee + ankle = 0, FK-tested at multiple amplitudes with sole pitch 0.00 deg throughout. Amplitude second, and here the first reasoning failed honestly: FK said shorter legs need LARGER reference scale (0.30 for 30 mm lift), and the FK was correct - but the premise was wrong ("FK 没错, 但前提错了"): it assumed foot lift must come from the reference. HighTorque Pi, same scale, uses 0.08 with a 0.02 m foot-height target - proof that lift is added by the policy ON TOP of the reference, whose actual job is pinning the phase relationship. Scale 0.30 made the reference the entire gait: over-constrained and over-driven. The correction went to 0.15, deliberately not Pi's 0.08: "一次只走一半, 留退路" (walk half the distance, keep a retreat).
Change
target_joint_pos_scale 0.30 -> 0.15 as one of v6-minimal's three changes, treating both the footfall force and the lateral kicking (yaw momentum scales with leg swing amplitude).
Outcome
v6 improved landing force 1.72x -> 1.55x, suspended tilt 45.9 -> 23.0 deg, turn-gain asymmetry 70% -> 19%; the later v6-halved-shaping experiment (35 mm -> 4 mm collapse) confirmed the reference still carries the gait's existence on this machine - the division of labor is real but machine-specific.
Mechanism
A joint-space reference plays two separable roles: encoding structure (phase relations that keep the foot flat) and injecting amplitude (energy). Structure transfers across robots and is checkable by FK against an invariant; amplitude is a negotiation with the policy, and over-assigning it to the reference removes the policy's freedom to modulate lift with state.
Applies when
- importing a reference gait / imitation target from another codebase
- reference amplitude reasoning based on leg length alone
- real swing amplitude far exceeds sim's under a strong reference
“FK 没错, 但前提错了。我默认抬脚必须由参考轨迹产生。HighTorque Pi 同尺度机器人 … 用 0.08, 而它 target_feet_height = 0.02 m —— 说明抬脚是策略在参考之上加出来的, 参考只负责钉住髋/膝/踝的相位配合。我们取 0.30 等于让参考本身就是整个步态, 过约束 + 过驱动”
train/WALK_V6_MINIMAL.md § ① target_joint_pos_scale 0.30 → 0.15 Every new penalty ships with a pre-registered withdrawal clause - if healthy gait must pay above the cap, the term stands down
calibration-threshold-with-withdrawal-clauseIntroduce every new penalty with: the zero-cost-option audit, a replay-calibrated weight formula (healthy pays a fixed small fraction of tracking), and a pre-registered withdrawal condition - and let the clause fire without argument when the calibration says the term cannot be afforded.
Symptom
Three same-shaped crashes had established a failure archetype: v4's clearance, v8a's landing window (weight off by 58x uncalibrated), and v6a's bare hip_yaw suppression all combined a zero-cost "don't move" option with a fee on any motion - a reverse barrier that pushes policies toward standing still.
Context
The v11 landing-window penalty was therefore introduced under a calibration-threshold protocol: (1) shape chosen with the window tightened (h_gate 0.03 -> 0.02, because 0.03 equaled the clearance target and priced the entire descent); (2) weight from a FORMULA, not judgment: measure the term's raw value on healthy replays (v5/v10b), set w = -(0.10-0.15 x tracking reward) / raw_healthy; (3) withdrawal clause pre-registered: if healthy gait must pay >15% of tracking no matter the tuning, the term is withdrawn to the next version rather than forced in - "不硬上". The companion hip_yaw quieting term ran the same protocol (calibrate on replays, healthy pays <=5%) and was later retired entirely when a structural fix (zero action scale) made its shaping tax unnecessary.
Change
Penalty introduction protocol: shape audit (what is the zero-cost option?), replay-based weight formula, healthy-pay cap with a written stand-down condition - all before training.
Outcome
The landing term was in fact withdrawn under its clause (v12 records "P5 落地窗口罚 已撤 … 维持撤下"), demonstrating the protocol firing as designed instead of the fourth same-type crash.
Mechanism
A penalty's damage mode is mispricing healthy behavior; since the healthy price is measurable in advance on replays, both the weight and the go/no-go decision can be computed rather than discovered by a ruined training run. The withdrawal clause converts "make it work" pressure into a clean deferral.
Applies when
- adding any motion-taxing penalty to a working gait
- a proposed term's weight has no measurement behind it
- a previous same-shaped term crashed training
“权重公式而非拍脑袋:先在 v5/v10b 回放上量 h_gate=0.02 的原始值,w = −(0.10~0.15 × 跟踪奖励) / raw_健康;标定门槛:若健康步态无论如何要付 >15% 跟踪,本项撤下留 v12,不硬上 (v4 clearance/v8a-B/v6a 三次同型翻车的教训:代价为零的"不动"选项 + 一动就收费 = 反向壁垒)。”
train/WALK_V11_SPEC.md § 6. P5 —— 落地窗口罚(三代欠账,标定门槛制) Exponential tracking kernels go flat exactly when the error is largest - pair them with an L2 term for the far field
exp-kernel-needs-l2-far-fieldNever let an exp/Gaussian kernel be the only tracking pressure on a quantity that can drift far from target: pair it with an unbounded (L2) term sized as the "don't diverge" floor, and check which frame the kernel reads.
Symptom
With only an exp-type yaw tracking term (exp(-err/std^2), std 0.25), a robot whose heading had drifted badly received almost no corrective gradient: at error 0.6 rad/s the term evaluates to exp(-0.36/0.0625) = 0.003 - near zero AND flat.
Context
The exp kernel is excellent for fine tracking near zero error but its gradient vanishes at large error - precisely when correction matters most. Fix: add track_ang_vel_z_err_l2 (-0.5), a plain quadratic on the same quantity: "exp 管精细跟踪、L2 管'别发散', 互补". Both terms deliberately read WORLD-frame wz (matching the exp term's source), because this torso sways enough that body-frame wz means are systematically off (measured -0.039 while actually turning +0.152). The same far-field-gradient argument reappears in the v8 risk list: frozen joints could not climb back because their huge error put them on the exp plateau ("远端梯度消失是冻结自锁的帮凶").
Change
Added the L2 companion term at -0.5 alongside the existing exp term (a term that had been in an earlier draft and was lost in a rewrite - itself worth noticing).
Outcome
Corrective pressure restored across the whole error range; the exp+L2 pairing became the house pattern for tracking terms.
Mechanism
d/de[exp(-e^2/s^2)] -> 0 as e grows: the kernel saturates and cannot distinguish bad from terrible. A quadratic's gradient grows with error, covering the far field; summing the two yields monotone corrective pressure with fine shaping near the target.
Applies when
- tracking rewards use exp/Gaussian kernels alone
- a drifted or frozen state fails to recover during training
- designing tracking terms for quantities with large transient errors
“exp 在误差大时梯度趋零, 恰好在最需要纠正的时候失灵。… 误差 0.6 → exp(-0.36/0.0625) = 0.003, 接近零且平坦。… exp 管精细跟踪、L2 管"别发散", 互补。”
train/WALK_V7_SPEC.md § ② track_ang_vel_z_err_l2 −0.5 —— 补 exp 的梯度洞 One fixed acceptance matrix for every rung - new skill must PASS while every old skill stays within a regression budget
fixed-acceptance-matrix-per-rungFreeze one acceptance matrix for the whole ladder; every promotion requires the new skill's PASS plus bounded regression on every prior skill, measured against the parent's baseline under the current (bug- fixed) metric code - and include command transitions, not just steady states.
Symptom
Sequential skill training silently trades old skills for new ones (C1 trained away the root's backward ability); without a constant measurement frame, each rung's numbers are incomparable and regressions hide.
Context
The C ladder ran the same 13-cell command matrix at 20 seeds per cell at every rung (stand; vx +0.15/+0.30; vx -0.10/-0.20; vy +/-0.10; wz +/-0.20; two vx&wz combos; two vx&vy combos), with promotion requiring "新技能 PASS 且旧技能不明显退化" - old-skill regression budget <=2/20 against the parent's recorded 20-seed baseline. For the transition-rich final rung, a command-switch block was added (forward->stop, stop->backward, forward->turn, left->right, turn->forward; survival + re-track within 2 s) because "每个 steady command 都会做 ≠ 命令切换不会摔" - steady-state success does not imply switch safety, and the joystick does switches. The C4 product's gate ran 260 cells (13 x 20) all 20/20.
Change
Battery frozen once, reused verbatim per rung; baselines re-measured per parent (and re-measured again after the metric-frame fix, since old baselines were taken with the buggy coordinate reading - "旧基线是坏坐标系的, 不可引用").
Outcome
Regressions were caught at the rung that caused them (C1's backward loss, C2's vx+0.30 decay), and cross-rung comparisons stayed valid for the ladder's whole life.
Mechanism
A constant matrix makes every rung's output a point in the same metric space, so "did we lose anything" is a table diff, not a judgment call; the per-skill regression budget converts previously earned PASSes into standing constraints on all future training.
Applies when
- designing gates for sequential skill addition
- promoting a checkpoint to be the next rung's root
- after any evaluation-code fix (old baselines must be re-measured)
“新技能 PASS 且旧技能不明显退化才晋级。… C5 追加:命令切换验收(steady ≠ transition) forward→stop、stop→backward、forward→turn、left→right、turn→forward,各 20 seed,判存活 + 切换后 2 s 内是否重新跟上。”
train/C_LADDER_RUN.md § 5. 固定验收矩阵(每级跑同一张,每项 20 seed) Curriculum-gate a penalty to the phase where its disease occurs - early on it only taxes exploration
gate-penalties-to-the-disease-phaseFor penalties aimed at late-stage pathologies (freezing, saturation, degenerate attractors), ramp the weight in only after exploration noise has decayed; anchor the terminal weight to measured healthy-vs-sick raw values, and shift all related tripwires to after the ramp completes.
Symptom
The action_saturation penalty, applied from iteration 0 in v8a, taxed exploration itself: with init_noise_std 1.2 the sampled actions paid ~-2.45/step before any policy had formed - while the disease it targets (clamp freezing) is a LATE pathology (v9 froze at iteration ~2624).
Context
v10 re-introduced the same penalty behind a curriculum gate: weight 0 until iter 1000, ramping linearly to -1.0 by iter 2000 - present only when the disease can occur, absent while exploration noise dominates. The trust argument was evidence, not hope: in v8a the term, while active, had pulled joint_pos_ref from 0.041 up to 0.155 and climbing - proof it can extract a policy from the frozen pit. Weight magnitudes were anchored to measured raw values (healthy v5 0.310 / v6 0.106 vs frozen v7 1.145 / v9 1.22 per step: at -1.0 healthy pays 6-18% of tracking, frozen pays 60%+, standing ~0). v10c then isolated the gated term as THE anti-freeze mechanism by single variable, upgraded to untouchable status in v11: "S 的门控机制(v10c 单变量铁案:任何情况下 不许撤,只许调终值)" - and v11 dared to relax other penalties only because S stood guard.
Change
action_saturation gated 0 -> -1.0 over iters 1000-2000 (later terminal value tuned -1.0 -> -0.5 with the gate mechanism itself frozen); tripwires adjusted to respect the gate's timing (freeze check moved to iter 2500-3000 to give the ramped term its effect window).
Outcome
Freezing stopped recurring while early training kept full exploration; the mechanism graduated from experiment to invariant within two versions.
Mechanism
A penalty's incidence depends on who occupies its support: early in training that is exploration noise (whose suppression starves learning), late it is the converged pathology. Time-gating aligns the penalty's presence with its target's presence, buying the constraint without the exploration tax - and tripwire timing must then be computed from the gate schedule, not from ungated precedents.
Applies when
- a structural penalty punishes exploration in early training
- a late-onset pathology (freeze/saturation) needs a standing guard
- deciding when a curriculum ramp should engage
“v8a 实锤它的病根是"罚在采样动作上"——init_noise_std 1.2 的早期等于罚探索(~−2.45/步);而冻结是晚期病(v9 速率 2624 才死平)… 门控让它只在病发期在场。… v8a 里它在场时 joint_pos_ref 从 0.041 爬到 0.155 且仍在升——有从低谷爬出的实证。”
train/WALK_V10_SPEC.md § 2. S 保险 —— action_saturation 课程门控 Four consecutive fixes were each continued from the previous fix's degraded state until the user stopped the ladder - "change parameters, don't stack errors" - rolled back to the last good checkpoint and audited the target geometry first
stop-stacking-roll-back-and-auditWhen successive rungs each start from the previous rung's output and the target symptom does not move, stop, roll back to the last good checkpoint and re-derive the next change from an audit; keep the measurements, discard the stacked remedies.
Symptom
After the real-robot splits, the V2.7 ladder tried to widen the stance: a term swap (A), a new stance knife (b), more iterations (甲), a doubled weight (乙). Stance barely moved while hip yaw ratcheted 46.7 -> 49.9 -> 52.5 deg toward its 60 deg limit.
Context
Each rung started from the previous rung's output. The user ruled on 2026-08-11 that things had gone wrong from V2.7-A: go back to v2_6 and rethink which parameters to change instead of stacking errors.
Change
乙 was killed at start and not counted; the product baseline rolled back to v2_6c model_29399; every measurement and law learned on the ladder was kept ("the data is real; what stacked was the treatment"). Before any new training, a zero-training kinematic audit of the stance targets was run.
Outcome
The audit found the stand_pose target itself rewarding the narrow stance (pose-target-geometric-audit) and showed geometrically why the yawed stance could not be widened with flat feet - so 乙 was proven unnecessary without running it. The next in-lineage attempts still failed, which is what established that the stance is set by the get-up path.
Mechanism
A rung continued from a degraded state inherits its compensations, so each new fix answers the previous fix's side effects; the yaw ratchet was the visible trace of that stacking.
Applies when
- three or more corrective rungs in a row without progress on the target metric
- a side-effect metric ratchets in one direction across rungs
- a new rung is being planned from the latest (not the best) checkpoint
“用户裁:"从 V2.7-A 开始就出问题了,应该回到 2.6 再思考如何改变参数而不是 错误叠加。"认账:A 的补丁 → b 的新刀 → 甲的加时 → 乙的加权,每级都从上级 的**退化状态**续(yaw 46.7→52.5° 的棘轮就是叠加痕迹)。 … 数据是真的,叠加的是处置。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 方法论裁定(2026-08-11,用户):V2.7 全阶梯叫停,回滚 v2_6 A curriculum ramp keyed to the process step counter re-fires on every resume - and shipped policies never saw the penalty
curriculum-counter-lineage-stepsKey every curriculum/ramp schedule to lineage-cumulative progress, not per-process counters; and audit where your shipped checkpoints sit relative to every ramp - a penalty that no product ever experienced is not part of your training.
Symptom
vx+0.30 died at a fixed relative time in every resumed run: resume at 500 -> slide at 1100, zero at 1300; resume at 700 -> slide at 1300, zero at 1500 - absolute depths offset by exactly the resume offset, relative timetable identical.
Context
ramp_reward_weight (the saturation penalty ramp) read env.common_step_counter, which restarts at 0 for every run including --resume. So start_step=600 meant "600 iters after THIS resume", not "lineage iteration 600". The A/B arm pair was the clean proof: their env.yaml differed only in log_dir, only resume point distinguished them, and the omni CurriculumManager had exactly one active term - nothing else could produce that timetable. Second consequence: every shipped checkpoint (s1e-500 at +500, C2-700 at +200, A800 at +100) was selected before its run's +600, so the saturation penalty weight was 0.000 for every product ever shipped - explaining saturation 33% and raw |action| 1.9 against clip 1.0 (hip_roll in bang-bang), i.e. half the heat budget.
Change
Two independent recommendations recorded: (1) make the ramp count lineage steps (add the checkpoint's iteration offset on resume) or pin terminal weights in downstream rungs instead of ramping; (2) give the saturation penalty its own rung - never mixed into a skill-learning rung (that would be two variables again).
Outcome
Explained the recurring +600 death of the highest-amplitude command and the persistent actuator saturation of all shipped products with one root cause; honest caveat booked (at +800 the weight is only -0.086, small, but vx+0.30 is the command demanding the largest action amplitude, so it is squeezed first).
Mechanism
Resumable training splits "the lineage" from "the process"; any schedule keyed to process-local counters silently re-applies its transient to every descendant run, and any product-selection habit that picks checkpoints early systematically samples the pre-ramp regime - the curriculum exists in the config but never in any shipped policy.
Applies when
- resumed/forked training with any scheduled reward or DR ramp
- a metric dies at a fixed offset after each resume
- shipped policies show behavior a late-schedule penalty should prevent
“ramp_reward_weight 读的是 env.common_step_counter,它每个 run 从 0 开始,--resume 也不例外。… 原始 run / 臂B | 500 | 1100 = +600 | 1300 = +800;臂A | 700 | 1300 = +600 | 1500 = +800 … 所有出品其实从没见过饱和罚。… 这解释了 sat_max_pct 33%、raw |a| 最大 1.9(clip 是 1.0)—— hip_roll 一直在 bang-bang,而罚它的那一项权重恒 0。热账的一半在这里。”
train/C_LADDER_RUN.md § 3g. 系统性问题:saturation_ramp 每次 resume 归零 Fall recovery was defined as the whole chain - any fallen pose, a stable stand, a clean hand-back to walking - and built as a second policy behind a deploy-side switch, not folded into the walking PPO
recovery-two-policies-and-a-state-machineDefine a recovery skill by the whole chain it must complete, including the hand-back to the next controller; if it is built as a separate policy, make the switching logic and its handoff contract a deliverable of their own, and keep the recovery observation contract a subset of the locomotion one so a unified policy stays possible later.
Symptom
A walking robot that falls needs a human to stand it back up. The design question on 2026-08-09 was whether to teach getting up inside the existing omni walking policy or beside it.
Context
The user set the goal as "any fallen pose -> stand up alone -> stand stably", and the spec named the real difficulty as the full chain fall -> recovery -> stable stand -> correctly initialised walking history and clock -> walking, making the deploy state machine a first-class deliverable. A unified single policy had a real-robot precedent (arXiv:2605.18611, a state-dependent gate near 37 deg tilt) but was deferred until a recovery policy and an omni policy were each reliable. The line ran on its own branch and worktree with every walk/stand/omni/run config path untouched. The development path copied the G1 learned get-up logic (arXiv:2502.12152): first find any feasible get-up (ugly accepted), then add smoothing, torque and real-robot constraints. The recovery contract kept the base 45-dim observation (command slice held at 0, no gait phase, no frame history - their reasons do not apply to a skill without a clock or a velocity task), so it stays a prefix of the 215-dim omni contract and a later merge is not foreclosed.
Change
Two policies and a deploy-side switch instead of one retrained walking policy; recovery got its own minimal contract (45 dims, full-range action, later the beta-anchored profile) and its own acceptance battery.
Outcome
The split held for the whole line: on 08-14 deploy_policy gained a second (PolicyIO, ONNX) pair behind --recovery-policy, each loaded under its own manifest contract, and the runbook runs stand_v1b or omni_c4_ff800 as the locomotion side with recovery_v3_1p1c. The literature scan of 08-10 found that every verified get-up implementation deploys one end-to-end policy (or softly gated experts) and stages only on the training side - so the runtime state machine here is the walk/recovery switch, not a staged get-up.
Mechanism
A separate policy keeps each reward table single-purpose and lets a proven walking lineage stay byte-frozen; the cost moves to the handoff, where every piece of state one policy leaves behind (history, clock, last action, command) must be reset for the other.
Applies when
- adding fall recovery or get-up to a robot that already walks
- choosing between one unified policy and a switched pair of policies
- designing the observation/action contract of a secondary skill
“先做 recovery policy + omni policy 两个策略,部署侧状态机切换;不把 recovery 硬塞进现有 omni PPO。 … 任务定义:**任意跌倒姿态 → 自己站起来 → 稳定站立**。真正的难点不只是"起身", … omni walk**(§6 部署状态机是本 spec 的一等公民,不是附录)”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §0 目标口径与架构判决(用户 2026-08-09 定) Diagnose a behavior failure by enumerating hypotheses and auditing each against the actual config, cheapest first
hypothesis-table-code-auditBefore changing anything, write the full hypothesis list for the symptom and audit each against the resolved config and measured magnitudes, cheapest check first; train only on the survivors.
Symptom
Real robot leaned forward "wanting to walk" but dragged its feet instead of lifting them - a symptom with many plausible causes and no obvious single fix.
Context
Seven hypotheses were listed and each checked against the actual training config files (velocity_env_cfg.py, isaac_values.py), ordered by check cost: missing foot clearance term (CONFIRMED, primary - feet_air_time existed but no swing-height term at all); energy penalties dominating (REJECTED - energy terms total -0.19 vs tracking +1.2, 16%); command range too narrow (CONFIRMED - (0.15,0.35)); nominal pose too crouched / action scale too small (HALF - knee 0.5 rad = 28.6 deg deep, scale fine); mixed PD across motor types (REJECTED - already grouped); missing base-height reward (REJECTED - present at -5.0); height-drop termination (REJECTED - none exists, which itself became finding #4 of the fix list).
Change
The audit produced a ranked fix list (add clearance penalty; widen speed range; reduce nominal crouch) with each rejected hypothesis documented so it would not be re-litigated.
Outcome
Three confirmed causes fixed over v5/v6: swing height went 22-23 mm -> 34 mm, tracking 81% -> 87%; the rejected hypotheses stayed rejected (no wasted rungs on energy weights or PD grouping).
Mechanism
Multi-cause symptoms invite guess-and-train loops; a written hypothesis table forces each candidate to be confirmed or rejected against actual values (not impressions), and cost-ordering the checks means most hypotheses die for the price of reading a config.
Applies when
- a real or sim behavior failure has multiple plausible causes
- the team is about to "try a fix" without an audit
- post-mortems keep re-proposing already-rejected causes
“真机现象:躯干前倾像要走,脚抬不起来(拖着蹭)。按成本从低到高逐条核查 … | 1 | 缺 foot clearance | ✅ 成立,首要 | 有 feet_air_time,无任何摆动足高度项 | | 2 | 能量惩罚压过跟踪 | ❌ 不成立 | 能量类合计 −0.19,跟踪 +1.2,只占 16% |”
train/WALK_DIAGNOSIS.md § walk 拖地问题 — 七条假设的代码核查结果 When the training reset distribution changes, freeze the acceptance distribution separately and pin its seed - the same checkpoint measured twice differed by 3.8 and 10.2 points
frozen-acceptance-distribution-and-pinned-seedAcceptance distributions are frozen artifacts, decoupled from whatever the training distribution becomes and evaluated with a pinned seed, and every gate row carries its sample size so that a small-n row is never read as a regression.
Symptom
R0.3 added easier roll-arc start states to the training resets, and the acceptance script shared the training category table; separately, the same checkpoint scored side 3.8 points and mid 10.2 points differently on two runs of the same script.
Context
Acceptance ran 512 parallel envs drawn from the fall categories. Had it kept following the training table, 15% of acceptance samples would have landed on states easier than prone - inflated scores and generations that could not be compared. The two same-checkpoint runs had identical per-item height medians, so the ruler had not changed; the spread was reset resampling noise (side n ~ 169, sigma 2.3%; mid n ~ 35, sigma 8.3%). The spec's "20 seeds" had always meant controlled seeds.
Change
accept_recovery.ACCEPT_CATEGORIES pinned to the four R0-R0.2 categories and decoupled from the training FALL_CATEGORIES; --seed 20260809 pinned, after which two consecutive runs were bit-identical. The mid row (n ~ 35) was labelled the bluntest gate.
Outcome
Every later generation (R0.3 through V3.1) was scored on the frozen distribution and seed, which is what let R0.3's intermediate state be read as "no measurable gain" (62.7 -> 62.1%) and R0.2's mid drop be booked as noise rather than a regression.
Mechanism
An acceptance set that follows the training distribution measures a moving target, and an unpinned reset draw adds sampling noise that small-n rows cannot absorb.
Applies when
- the training reset or command distribution changes between generations
- repeated evaluations of one checkpoint disagree
- a small category drives a pass/fail decision
“**分布冻结**:`accept_recovery.ACCEPT_CATEGORIES` 钉死 §5 四类 … 与训练侧 `FALL_CATEGORIES` **解耦**。 … **side 差 3.8 点、mid 差 10.2 点**(h 中位逐项一致,证明不是尺子变了)—— 纯 reset 重采样噪声 … 钉死后两次连跑逐位相同。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §14 验收尺子的两处加固(R0.3 起生效,向后兼容) Record where every failed episode ends - an end-state confusion matrix showed all failures finishing seated and overturned a "cannot roll over" diagnosis that per-category success rates hide by construction
end-state-confusion-matrixFor any multi-category acceptance, report where each failed episode ends, not only which category it started in; it costs a few lines and no extra simulation, and it separates "cannot reach the goal" from "reaches the wrong basin".
Symptom
Prone scored 0% for three generations; the working diagnosis was "prone lacks the roll-over skill", and R0.3 spent a run adding prone-to-side roll-arc start states. It bought nothing: the 45-deg roll band itself only moved from 24.2% to 26.6% after 3,000 iterations.
Context
Acceptance reported success per starting category. A final-state table (lying prone / on the side / supine / seated / standing for every failed episode) was added to accept_recovery.py at R0.3.
Change
The confusion matrix became a permanent part of the acceptance output, and the prone diagnosis was rewritten from it.
Outcome
The prone, side and supine columns were all zero - every failure ended seated - and prone had righted its torso in 159/159 episodes (tilt under 30 deg in 100%). The missing ability was standing up from one specific seated configuration, not rolling over, which redirected the next rungs to foot placement and to a configuration probe.
Mechanism
Per-category success rates collapse "reached the wrong basin" and "never reached anything" into the same zero; the end state separates them.
Applies when
- a category sits at 0% and the diagnosis rests on its label
- recovery, manipulation or navigation tasks with distinct terminal states
- an intervention aimed at the presumed cause shows no effect
“**① 末态混淆矩阵 —— 固化(已在 `accept_recovery.py`)。** 它给出的 "趴/侧躺/仰躺三列全 0、所有失败都终于坐姿"是本线最改变决策的一个事实, 而**逐类成功率按构造看不见它**。 … 成本十来行、零额外仿真。 … **② prone 病因更正(旧诊断作废)。** 旧:"缺翻身"。新:**prone 159/159 全部 把躯干翻正**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §15 R0.3 判决 + 三件事的判断 Design the next run to complete the 2x2 - either outcome then convicts or acquits a factor cleanly
fill-the-missing-factorial-cellWhen two config factors are jointly suspected, lay out the factorial of existing evidence, spend one run on the missing cell with both interpretations and an early-abort tripwire written in advance - and treat either outcome as a verdict, not a disappointment.
Symptom
Hip joints froze at the action clamp in v7, but the history could not say whether the culprit was the raised action_rate (-0.2) or the halved reference amplitude (scale 0.15): existing versions covered only three corners of the (rate x scale) space - v5 (-0.03, 0.30) healthy, v6 (-0.03, 0.15) healthy, v7 (-0.2, 0.15) frozen.
Context
v9 was designed explicitly as the missing cell (-0.2, 0.30), with the readings pre-registered: v9 not frozen -> the real anti-freeze force was always the reference amplitude and -0.2 may stay; v9 frozen -> -0.2 is convicted beyond appeal (freezes at both amplitudes) and the next version goes straight to a structural fix. "两个结局都是干净的信息" - both endings are clean information.
Change
One training run allocated purely to complete the factorial, with freeze tripwires (joint_pos_ref telemetry <0.1 at iter 1000-1500 -> abort, do not run to 6000) so a conviction costs the minimum compute.
Outcome
v9 froze - the rate weight was convicted at both amplitudes ("−0.2 铁案定罪"), and v10 moved to the structural saturation fix with the weight question closed instead of re-litigated.
Mechanism
Three corners of a 2x2 leave the two factors confounded in the failure corner; the fourth observation makes each factor's marginal effect identifiable. Pre-registering both readings turns the run into a guaranteed-informative experiment regardless of outcome.
Applies when
- two config changes are confounded in a failure
- version history already covers some corners of a factor grid
- deciding what single experiment buys the most attribution
“这恰好补齐一个 2×2 实验矩阵的缺格 … v9 不冻 → 真正的抗冻结主力一直是参考摆幅,−0.2 可以留;v9 仍冻 → −0.2 铁案定罪(两种摆幅下都冻),v10 直接上结构修复 … 两个结局都是干净的信息。”
train/WALK_V9_SPEC.md § 0. 设计原则 (2×2 实验矩阵) After seven patch-generations, freeze the lineage as a regression baseline, fix the structural debts, and retrain from zero
freeze-lineage-fix-structure-restartWhen successive rungs keep trading one symptom for another, ask whether the remaining problems are structural (contracts, latency, sampling, reward-table architecture); if so, freeze the lineage as regression baselines, pay the structural debts, and restart minimal - carrying forward laws and instruments, not weights and weights' patches.
Symptom
The v5-v11 walk lineage had accumulated interacting patches (reward terms, gates, clamps, per-joint scales) faster than it converged on the user's goal; v12's spec itself was superseded before training by an external review's verdict that the remaining problems were structural, not parametric.
Context
The 2026-08-05 status banner records the pivot: the walk profile was rolled back wholesale to v10b parameters, the v5-v11 lineage frozen "只作回归对照" (kept only as regression baselines), and four structural debts were named as prerequisites for a from-zero straight-walk baseline: the action-latency FIFO (fixed with its own test), the ONNX manifest contract, discrete command sampling, and a minimal reward table. The v12 spec - fully designed, partially implemented - was suspended: "本规格挂起,不再按此开训".
Change
Strategy switched from "one more patch generation" to freeze-fix-restart: lineage checkpoints retained as comparison anchors, infrastructure hardened first, then a clean retrain with a minimal reward table (this restart produced the s* generation that later became the real-robot SOTA line).
Outcome
A designed-and-ready training generation was deliberately not run - the review's structural findings outranked sunk design cost; the restart line inherited seven generations of laws (calibrations, gate batteries, falsified fixes) without inheriting their entangled reward table.
Mechanism
Patch lineages accumulate coupled terms whose interactions eventually cost more to reason about than a restart costs to train; the knowledge worth keeping is the laws and instruments (measured plant values, calibrated gates, falsified directions), not the entangled weights. A restart on hardened structure converts the lineage's lessons into a clean initial design instead of another delta.
Applies when
- repeated rungs shuffle symptoms without net progress
- an external review flags infrastructure/contract debts
- deciding between another patch generation and a clean retrain
“同日外部评审定调换路线:冻结 v5~v11 血统(只作回归对照),修结构性问题(latency FIFO 已修 tests/test_action_latency.py、ONNX manifest 契约、离散命令采样、最小奖励表)后从零训直行基线。本规格挂起,不再按此开训。”
train/WALK_V12_SPEC.md § ⚠️ 状态 (2026-08-05) Fix a too-deep nominal pose before adding any penalties - the default stance defines the basin training starts in
nominal-posture-before-penaltiesBefore tuning penalties on a degenerate gait, audit the nominal pose and height targets against morphology and published ratios; if the default stance encodes the degenerate behavior, fix it first - and recompute dependent quantities (init height) by FK, not by hand.
Symptom
Policy lived in a crouched shuffle; nominal knee angle was 0.5 rad (28.6 deg) - deeper than published configs (Unitree G1 0.3 rad / 17.2 deg, Booster T1 0.4 rad) - so the policy's starting point and its action-space center both sat inside the crouch basin.
Context
Initially ranked "secondary" in the local diagnosis, this was promoted to co-first priority by the cross-check against published reward tables, which states that with nominal knee flexion above ~0.4 rad, fixing the posture must precede adding any penalties ("改这个之前别加任何 惩罚都是白费"). Companion base-height items: walk profile had weakened base_height_l2 to -5.0 (base class -10, field standard -10 to -20, "the second most common cause of death"), and the height target must be the STANDING height (0.384), not the crouch height.
Change
Nominal knee 0.5 -> 0.3 rad with init_base_height recomputed by MuJoCo FK (0.3739 -> 0.3802); base_height_l2 restored to -10 with standing height target; both bundled as first-priority alongside the clearance term.
Outcome
Part of the v5/v6 package that lifted swing height to 34 mm and tracking to 87%; the crouch basin stopped being the default answer.
Mechanism
The nominal pose is the fixed point every regularizer pulls toward and the point where action=0 lands; if that point is itself the degenerate posture, every penalty fights the geometry. Correcting the attractor is prior to shaping the gradient field around it.
Applies when
- policy converges to a crouched or collapsed posture
- nominal joint angles were chosen for stability rather than gait
- base-height reward targets or weights were locally weakened
“研究明确说"nominal 膝屈超过 ~0.4 rad 必须先改,改这个之前别加任何惩罚"。我们是 0.50,超标。… base_height_l2 在 walk profile 里被减到 −5.0(基类是 −10)。研究说这是"第二常见死因"且应 −10 ~ −20。改回 −10。目标高度用站立高 0.384 是对的(研究要求 target 必须是*站立*高度而非蹲姿)。”
train/WALK_DIAGNOSIS.md § 修正 ①(升级优先级) / 修正 ④ The smoke watcher panicked at the wrong operating point and its score goes blind when gates saturate - treat it as a survival sentinel, not a judge
smoke-watcher-operating-pointConfigure every automated evaluator at the lineage's declared deployment operating point, give it graded metrics that cannot saturate, and until then scope its authority to catastrophe-detection - never let a mis-configured watcher stop or rank a rung on its own.
Symptom
Two watcher misfires in one ladder: (1) during the friction rung the watcher (evaluating at kd 1.0) reported panic-level 1/3 survival from iter 3300 - falsified by the official kd 1.2 scan, because the lineage's design operating point was kd 1.2 and the watcher lacked the --kd-scale passthrough; (2) during the PD rung the watcher's early-stop score froze at iter 1050 despite ongoing drift improvements, because with all eight gates passing (constant 0/3 failures) the score has no gradient left - "八门全过恒 0/3 时 score 对漂移改善盲, s1f 课文三现".
Context
Both are the same category: the in-training smoke loop is an instrument with its own configuration (operating point, score design), and its verdicts are only as aligned as that configuration. The booked doctrine: "冒烟只当存活哨兵" - until the watcher evaluates at the deployment operating point with graded metrics, its role is detecting catastrophes, not ranking checkpoints; ranking belongs to the full battery at the design operating point (and the watcher's scoring was separately patched to weight survival 3x so recoveries during hard phases are not early-stopped away).
Change
Watcher debt booked (--kd-scale passthrough); score saturation acknowledged with graded columns planned; selection authority kept with 20-seed batteries at the declared operating point.
Outcome
A false panic did not abort a rung that was actually passing at its design point; a frozen score did not hide real drift gains; the instrument's authority was scoped to what its configuration can actually see.
Mechanism
An evaluator is itself configured (gain profile, delay, metrics); evaluating a policy away from its design operating point measures a counterfactual robot, and bounded scores saturate once binary gates pass, losing all sensitivity. Instruments need the same operating-point discipline as deployments and graded outputs to retain gradient.
Applies when
- an automated smoke loop contradicts the official battery
- early-stop scores freeze while graded metrics still improve
- lineages with non-default deployment gain/delay profiles
“watcher (kd1.0 口径) 3300 起 1/3 恐慌被 kd1.2 正式扫描证伪为考纲外假象 —— 工作点评测口径教训: watch_ckpt 缺 --kd-scale 透传 (待补), 冒烟只当存活哨兵。… watcher score 饱和误停 @1050(八门全过恒 0/3 时 score 对漂移改善盲, s1f 课文三现)”
train/README.md § omni_s2e_fric (watcher 恐慌被证伪) / omni_s2e_pd (500 臂) PPO's Gaussian noise cannot compose phase-locked oscillations - deliver them as feed-forward and let the policy learn the residual
feedforward-for-phase-locked-skillsIf a skill needs a temporally coherent (phase-locked) action component, do not expect step-wise exploration to find it: inject a verified feed-forward and train the policy as a residual stabilizer, keeping the feed-forward inside the deployment contract.
Symptom
Four different reward arrangements (no reference / wrong-sign reference / correct-sign reference / cage released) all failed to elicit sidewalk, while open-loop probes proved the behavior existed and was safe on the same platform with the same policy as base.
Context
Producing lateral velocity requires a phase-locked hip_roll oscillation synchronized to the gait clock. PPO's exploration is per-step, zero-mean, uncorrelated Gaussian noise - it can never compose a sustained phase-locked component, so the behavior is unreachable by exploration regardless of how it is rewarded. The fix changed the delivery channel: target = default + scale*action + lat_ff(cmd_vy, phi). The policy's action becomes a residual on top of the feed-forward, retaining full balance authority (it can even cancel the feed-forward); the feed-forward supplies exactly the component exploration cannot. This mirrors why the sagittal joint_pos_ref worked (it also delivered phase structure), just via a different channel.
Change
Contract-level change, done cleanly: new profile omni_ff (= omni + lat_ff_gain -0.5), existing omni profile bit-identical; feed-forward applied after the action delay stage; missing cmd/phase raises instead of silently dropping; deployment must use the same phi as build_obs (recomputing gives a one-tick phase misalignment).
Outcome
From C2-700, +100 iterations sufficed: product omni_c4_ff800 scored vy +120%/+125% (from +4%/-1%), 260/260 cells at 20/20 survival, zero old-skill regression, left/right gap 5 pp - the entire C4 saga resolved by changing the delivery mechanism, not the reward.
Mechanism
Exploration noise spans only the subspace its correlation structure can express; skills requiring coherent oscillation lie outside the span of i.i.d. per-step noise. Feed-forward moves the required structure into the action pipeline where it needs zero probability mass to appear, reducing the learning problem to stabilizing around a demonstrated behavior - which PPO does well.
Applies when
- a periodic/oscillatory skill trains flat under every reward variant
- open-loop injection of the behavior already works
- considering GRU/curriculum/exploration tricks for a rhythmic skill
“病因不在奖励,在探索形式:产生侧向速度需要相位锁定的 hip_roll 振荡,PPO 的逐步高斯噪声零均值无相关,合不出相位锁定分量。… target = default + scale·a + lat_ff(cmd_vy, φ)。策略动作因此是前馈之上的残差,保留全部平衡权限”
train/C_LADDER_RUN.md § 3j. C4-redo4:唯一变量 = 侧步参考改为前馈注入(契约级) The get-up kept getting faster because standing earlier paid more every step - lowering torque authority barely slowed it, and only zeroing the standing income for the first 3 s moved the pace into the design band
per-step-income-drives-speed-time-gateWhen a skill is too fast, find the term that pays for finishing early and gate that income by time; keep the "get into position" term ungated so the policy does not learn to wait, use a ramp instead of a cliff, and confirm with a paired same-level experiment that the drift is motivational before changing it.
Symptom
The user judged the get-up too fast (Isaac medians about 0.7-1.6 s) and suspected path dependence: the policy seemed to get faster the longer it trained.
Context
Lowering the beta authority 0.40 -> 0.30 cut impact but moved supine only 1.70 -> 2.00 s: coordination-limited, not torque-limited. A paired experiment inside one beta level (checkpoint 15,600 vs 18,499, +2,900 iterations, same ruler) measured the drift: get-up medians -7 to -10%. The spec concluded the motive, not the path, was the cause - per-step standing income pays for every early step, and any lineage (even one from scratch) races toward the fastest solution inside its constraints.
Change
V2.5: the standing income (base_height, stand_pose, still, feet_on_ground) multiplied by w(t) = clamp(t/3 s, 0, 1); upright deliberately NOT gated, so righting and sitting up early still pay and the policy is not taught to lie flat and wait; a ramp, not a step. V2.5b: zero before t0 = 3 s, then a 1 s ramp.
Outcome
V2.5: Isaac 100%, get-up +17-43% slower, MuJoCo 100/100/98/100% (the best cross-simulator reading yet), still short of the 3.5-4.5 s design band - a linear ramp only discounts early income. V2.5b: MuJoCo supine 2.04 -> 4.10 s and prone 3.18 -> 4.04 s, inside the band; the Isaac pace barely moved (a lineage habit on a gradient-free plateau). Later the zero gate proved harmful when trained from scratch (curriculum-history-is-part-of-the-product) and in the V3.1 lineage (time-gate-vs-wide-stance-retire-the-fix).
Mechanism
Constraints on authority or velocity change how the fastest solution looks; the time structure of the task income decides how fast the fastest solution is.
Applies when
- a policy is faster or more aggressive than wanted and constraints do not slow it
- progress-style rewards pay every step spent at the goal
- performance drifts faster with more training at fixed settings
“**V2.5 机制(唯一)**:站立收入(base_height/stand_pose/still/feet_on_ground) 乘时间斜坡 w(t)=clamp(t/T_gate,0,1),T_gate=3.0 s;**upright 刻意不门控** (翻正/坐直早期照常拿钱,防"躺平等门开" … **用户假设量化 证实:逐步计酬动机在 β 包络内持续压缩时间,约束挡不住动机 —— V2.5 动机层 修法为正解。**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §39 V2.5 预注册 / §39 补 配对实验 Armature must be N^2 x rotor inertia, never 0 - measure it no-load
armature-n2-rotor-inertiaEvery geared actuator carries N^2 * I_rotor of reflected inertia at the joint; set armature from a no-load measurement, never leave it 0 and never guess it.
Symptom
Sim joints accelerate more easily than real joints; old MJCF had armature = 0 (rotor reflected inertia entirely unmodeled), a systematic sim2real gap on every joint.
Context
Original hand-written MJCF plant used armature 0. The team derived and then measured the correct value: torque needed at the rotor is I_rotor * N * alpha; after the N:1 gearbox the output shaft "feels" an extra N^2 * I_rotor of inertia. With a 9:1 reduction that is an 81x amplification of the rotor inertia, far too large to ignore.
Change
Set per-motor armature from no-load (motor out of the robot) measurement instead of 0: RS06 = 0.0070 kg*m^2, RS02 = 0.0032 kg*m^2, RS00 = 0.0015 kg*m^2. Landed together with measured friction as the first fully-measured plant parameter set.
Outcome
Plant parameters "第一次全部来自实测" (first time all from measurement); became the frozen plant baseline for all subsequent training generations.
Mechanism
Reflected inertia scales with the square of the gear ratio: the rotor spins N times faster than the joint, so its kinetic energy (and the torque needed to accelerate it) appears N^2 larger at the output. Omitting it makes simulated joints unrealistically fast/light, so policies learn action rates the real actuator cannot deliver.
Applies when
- building or auditing a simulation plant model for a geared/QDD actuator
- sim policy moves joints faster or snappier than the real robot can
- MJCF/URDF review shows armature or rotor inertia set to 0 or a default
“armature 转子反射惯量有问题 在sim里面一定要处理 不能是0,空机测试。转子处需要的力矩 = I_rotor × N × α 经减速箱放大 N 倍后 = N² × I_rotor × α 所以输出轴"感觉到"多了一个 N² × I_rotor 的惯量。 这就是 armature。… 关键是那个平方。减速比 9:1 就放大 81 倍。”
Experience.md § # armature 转子反射惯量有问题 (line 9) Prove a new penalty actually fires - two ways a clearance term silently did nothing
inert-reward-term-auditBefore training with a new reward term, log its realized per-step value under the current policy and confirm it is nonzero where intended - check coordinate zero-points against FK and check who occupies the term's gate; and never weaken the term that creates the states your new term needs.
Symptom
A newly designed swing-height clearance penalty could have trained as a no-op twice over, and the companion advice to lower feet_air_time actively backfired when tried.
Context
Instance 1 (zero-point offset): the proposed code used body_pos_w of the foot link, but that is the ankle_roll_link frame origin, which sits 0.0585 m above the ground even with the foot flat on it - so (0.03 - 0.0585) is always negative and the penalty is永远 0; the 0.0585 offset must be subtracted (verified identical in MuJoCo FK and Isaac). Instance 2 (gate occupancy): the clearance penalty fires only in swing phase; a dragging policy keeps both feet in contact, so the penalty is constantly 0 for exactly the policy it was meant to fix - and worse, any slight lift immediately incurs it, a reverse threshold. Lowering feet_air_time to 0.5 on that advice measurably collapsed air time to 0.0002 (below v3). Corrected understanding: "clearance 是把已有的摆动相抬高, 造出摆动相仍要靠 air_time" - air_time creates the swing phase, clearance raises it.
Change
Fixed the height zero-point; kept feet_air_time as the swing-phase creator with clearance layered on top; both errors documented as corrections to the team's own earlier advice.
Outcome
With both fixed, swing height rose from 22-23 mm (v2) to 29 mm (v5) to 34 mm (v6); the inert-term failure class entered the standing checklist.
Mechanism
A penalty's gradient exists only where its gate is occupied and its argument crosses its threshold; frame offsets shift the threshold out of reach, and phase gates can have zero occupancy under exactly the policy being treated. Terms interact as an ecology - one term must create the states in which another can act.
Applies when
- adding any gated or thresholded penalty (clearance, impact, slip)
- a new term produces no behavioral change at any weight
- body-frame positions are used in reward code
“body_pos_w 是 ankle_roll_link 坐标系原点,平放触地时仍高出地面 0.0585 m。… (0.03 − 0.0585) 恒为负 → 惩罚永远是 0 … clearance 惩罚只在摆动相生效,拖地时两脚始终触地 → 惩罚恒 0;而一旦轻微抬脚就立刻扣分,对正在拖地的策略是反向门槛。… 正确认识:clearance 是"把已有的摆动相抬高",造出摆动相仍要靠 air_time。”
train/WALK_DIAGNOSIS.md § walk_v4 独立验收 — 本文档给的两处代码/建议是错的 Audit rewards by realized contribution (weight x achieved value) - a weight of 2.0 was really paying 0.04
realized-contribution-auditEvaluate a reward table by each term's realized per-step contribution under the current policy, never by its weight column; if a term's realized value is ~0, escalating its weight is a no-op - change the term's structure instead.
Symptom
Foot dragging persisted through repeated weight escalation: feet_air_time had been raised 0.25 -> 1.0 -> 2.0 across versions with no behavioral change, and the training-side comment even recorded the fact ("几乎没有 单支撑相, 是拖着脚蹭") without the fix changing form.
Context
Computing realized per-term contributions in the trained state exposed the economy: feet_air_time contributed weight 2.0 x achieved 0.019 = 0.038 per step, against tracking's +1.20 - lifting the leg earned 3% of what tracking earned, so dragging was the rational optimum no matter the weight escalation. The same table acquitted the energy penalties (total -0.30 negative vs +1.74 positive) that a naive read of weights (-5.0 orientation!) would have blamed.
Change
Fix redirected from "raise the weight again" to "add a term whose realized contribution changes the optimum": a clearance penalty sized so its realized magnitude (~0.018/foot when dragging) is comparable to feet_air_time's, enough to flip the optimum without drowning tracking.
Outcome
With the term economy corrected (plus posture/range fixes), swing height reached 34 mm and tracking 87% by v6; weight escalation of the old term was abandoned.
Mechanism
A reward weight is only a multiplier on whatever the policy currently achieves on that term; when the achieved value is near zero (behavior absent), escalating the weight multiplies near-zero. Optimizer behavior is governed by realized per-step magnitudes, so audits must be conducted in that currency.
Applies when
- a behavior persists despite repeated weight increases
- auditing whether penalties are "too strong" or rewards "too weak"
- sizing a new reward term against existing ones
“把 feet_air_time 权重从 0.25 → 1.0 → 2.0 一路加,但没有加高度项。量级算下来:feet_air_time 权重 2.0 × 实得 0.019 = 0.038,而跟踪奖励是 1.2。抬腿的边际收益只有跟踪的 3%,拖地当然是最优解。… 正项 +1.74,负项 −0.30。能量惩罚不是瓶颈,抬腿没收益才是。”
train/WALK_DIAGNOSIS.md § 决定性证据(#1) / 各项奖励的实际量级 The walk phase machine structurally cannot express flight - rebuild the representation for duty < 0.5, teach flight with a mask tax, never a cliff bounty
phase-machine-structural-limitsWhen a new gait changes the contact pattern's structure, audit whether the phase/mask representation can express it and rebuild the representation if not; teach the new contact pattern with graded mask-mismatch pressure and count it in acceptance with artifact-proof definitions (minimum segment length), never with cliff bounties.
Symptom
Running requires both feet airborne, but the walk-era phase machine switches legs by the sign of sin(phase) - with duty < 0.5 the two swing windows must OVERLAP during flight, which a sign-switching representation cannot express at all.
Context
The run phase machine was re-architected rather than patched: per-leg phases (left = phi, right = phi+0.5 mod 1) with leg_phase < duty defining stance, aligned to the walk sin convention at duty=0.5 so the machines agree where their domains overlap. Flight is taught by the SAME mechanism that once cured foot-dragging, direction reversed: in the two planned flight windows the contact mask is (0,0) and feet_contact_number_duty charges -0.3 per foot still on the ground - a mild ~0.16/step tax, deliberately NOT a cliff: "悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文". The reference shape (half-sine bump over swing progress) is zero at window boundaries by construction, eliminating the clearing-window that the C4 probe measured to cost 13-19% on non-sinusoidal references. The shape self-check ("零代价选项是什么") was run on three behaviors: standing pays ref everywhere (known cmd=0 stepping risk, booked), walking pays only the flight-window tax, proper running collects full marks.
Change
New rewards.py run section (leg_phase_duty / stance_mask_duty / ref_run / clearance_run / contact_duty) with the walk versions untouched byte-for-byte; flight acceptance metric defined with a segment-length floor (>=40 ms to count) so numeric contact flicker cannot fake flight.
Outcome
Flight became expressible and taught by a calibrated mild pressure; the walk lineage's phase code stayed frozen as its own contract.
Mechanism
A phase representation defines which contact patterns exist in the reward's vocabulary; duty cycling below 0.5 introduces states (double-flight) outside a half-period sign convention's language, so no weight tuning can teach them. And rare desirable events taught by cliff-shaped bounties invite hacks (jumping in place); the graded mask tax prices the planned pattern without creating a jackpot.
Applies when
- extending a walking stack to running/jumping (duty < 0.5)
- a desired contact pattern never appears despite reward increases
- defining flight/contact acceptance metrics
“duty<0.5 时摆动窗 (1−duty)T > T/2,两腿摆动窗在腾空段重叠 —— sin 符号切腿的机制结构上表达不了"双脚同时在空中"。… feet_contact_number_duty 对"还踩着地"持续 −0.3/脚 —— 与 walk 治拖地同一机制,方向相反。不腾空的税 ~0.16/步 … 梯度温和不构成悬崖(悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文)。”
train/RUN_V0_SPEC.md § 5. 相位机设计 Audit which joints your imitation term constrains - a task that needs deviation is fighting the reference
imitation-term-scope-auditList which joints your imitation/deviation terms actually constrain and check the new skill's required motion against that list; for balance-coupled joints deliver references as feed-forward residuals, not absolute-position targets - and never assume "reference = 0" is neutral.
Symptom
Sidewalk would not learn despite a dedicated tracking reward; meanwhile the gait-shaping imitation term (joint_pos_ref) computed its error norm over ALL 12 joints while its reference covered only the 6 sagittal joints - roll/yaw reference was constantly 0.
Context
Two prior generations had shown the forward gait itself was taught by joint_pos_ref, not discovered by PPO (v6 halved the shaping and swing height collapsed 35 mm -> 4 mm). So the reference is load-bearing - but sidewalk requires hip_roll to deviate from nominal, and the all-joints norm punished exactly that deviation: "一边悬赏一边罚过程" (posting a bounty while punishing the process). A follow-up experiment (C4-redo3, free_roll=True releasing the 4 roll joints from the norm) raised the regularization headroom 6x -> 27x yet sidewalk stayed flat and released hip_roll wandered, killing other skills - net negative, withdrawn. A --roll-absolute probe showed the converse failure: pinning roll to a clock-driven absolute trajectory drove tilt 6.9 -> 13.7 deg. Conclusion recorded: absolute-position imitation cannot teach actions that must be superimposed on state feedback.
Change
The audit reframed the problem: neither punishing roll deviation nor freeing roll nor absolute roll tracking works; the reference for a balance-coupled joint must be delivered as feed-forward under the policy's residual control (see feedforward-for-phase-locked-skills).
Outcome
free_roll rung: joint_pos_ref term rose 0.887 -> 1.104 (release confirmed effective) but vy stayed flat; regularization hypothesis eliminated by experiment.
Mechanism
An imitation error norm defines a cage: joints inside it are pulled to the reference in absolute position, so any skill requiring systematic deviation is taxed per step; but joints carrying active balance cannot follow absolute references either, since their correct position depends on state. The scope and the delivery mechanism of the reference are therefore design decisions per joint, not defaults.
Applies when
- adding a skill that moves joints your reference sets to zero/nominal
- an imitation or deviation penalty coexists with a new tracking reward
- considering releasing joints from a shaping term mid-lineage
“前进步态也不是 PPO 自己发现的,是 joint_pos_ref 教出来的(v6 砍半塑形 → 抬脚 35 mm 塌到 4 mm…)。而 ref_joint_offset 原本只写 6 个矢状面关节,roll/yaw 参考恒 0 —— 侧走既没被教,roll 一偏离 nominal 反被 joint_pos_ref 扣分。一边悬赏一边罚过程。”
train/C_LADDER_RUN.md § 3e. 为什么首战 FAIL / 3i. 解锁笼子 Before resuming a checkpoint, diff the current cfg against what the checkpoint was trained with
resume-state-dr-audit"One variable per rung" counts variables against what the checkpoint actually experienced: audit the checkpoint's logged training config and align every unintended difference before resuming.
Symptom
Two consecutive rungs (C1 back-mode, C2' forward-turn) failed from the same root with the same full-regression signature despite adding different new modes - so the mode was not the cause.
Context
Both runs resumed s1e-500 with the then-current cfg, which carried PD band (0.8,1.2) plus three DR events (base_com, joint_friction, push_robot) accumulated by later lineages. Verified on the training machine from the source of truth (the run's logged params/env.yaml): s1e-500's actual training state was PD +/-10% (0.9,1.1) and all three DR events None. Resuming it under the new cfg meant eating 4 new plant variables plus a new mode at once - the intended "1 variable" was actually 5. A worse variant (c1_redo from s2e_pd-1400) added push +/-0.3 to a root that had never seen it: near-total collapse within +100 iters.
Change
C2 aligned the cfg to the checkpoint's training state before resuming (PD back to (0.9,1.1), three DR events off) - making the new mode the only true variable. Permanent rule recorded: compare the checkpoint's training-time DR with the current cfg before any resume.
Outcome
C2 trained successfully from the same root that had "failed" twice (wz 20/20 with genuine sign-antisymmetric response by iter 700-800); the A/B falsification ("两个不同模式同签名崩") plus the env.yaml verification closed the attribution.
Mechanism
A resumed policy is instantly evaluated (and its value function trained) under whatever plant distribution the cfg specifies; every DR term the checkpoint never adapted to is a distribution shift applied on day one, compounding with the intended change. Single-variable discipline is therefore a property of (cfg diff) x (checkpoint history), not of the cfg diff alone.
Applies when
- resuming or forking any checkpoint under an evolved config
- a resumed run degrades broadly within the first few hundred iterations
- two different changes from the same root fail with the same signature
“A/B 定谳:两个不同模式同签名崩 → 病因不是模式,是「从 s1e-500 续训」。… s1e-500 训练态 = kp/kd ±10% (0.9,1.1),base_com / joint_friction / push_robot 全 None;而 cfg 里带着 (0.8,1.2) + … 三个 DR —— 从它续训等于一次吃 4 个新 plant 变量 + 新模式 … 永久教训:续训前必须比对 checkpoint 的训练态 DR 与现行 cfg。单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」。”
train/C_LADDER_RUN.md § 3b. 这不是重复实验 —— 前两次的病根已定位并修掉 Under continuous 3-axis uniform sampling, pure straight-line walking is a zero-measure event the policy never trained
zero-measure-commands-need-mode-samplingEnumerate the exact command points users will actually issue (straight, stop, in-place turn) and give each explicit probability mass via mode sampling with off-axes pinned to zero - never assume a continuous sampler covers its measure-zero subsets.
Symptom
"The robot drifts even in sim when told to walk straight" persisted across reward tunings - because with commands drawn as vx in [0.15,0.5] x vy ~ U(+/-0.2) x wz ~ U(+/-0.6), the event vy=0 AND wz=0 has probability zero: pure straight-line walking was never sampled even once.
Context
Restart evidence item #3: "纯直行是零测度点 … 'sim 里直行就漂'是分布的 必然,不是 reward 没调好" - the drift metric was legitimately drowned by commanded turning (v11's own comment self-documented this). The structural fix is discrete mode sampling: a custom ModeVelocityCommand that first draws a mode by share (stand/forward/back/turn/side/mixed), then draws values only on that mode's axes with all others pinned to exact zero - which is also what preserves single-variable discipline in the C ladder (native 3-axis uniform "采不出'离散模式桶' … 把 C1~C4 的单变量纪律直接毁掉"). The mixed mode later got an ellipsoid constraint rather than a cube for the same reason in reverse - corner combinations of a cube are unrepresentative extremes.
Change
Command generation moved from independent per-axis uniforms to mode-bucket sampling with pinned-zero off-axes (plus 20% rel_standing); acceptance likewise evaluates per mode.
Outcome
Straight-line behavior became a trained, testable mode instead of a measure-zero hope; the C ladder could add one mode per rung with provable isolation.
Mechanism
A policy optimizes expected reward under the command distribution; events of probability zero contribute nothing to the objective, so exact-zero-command behaviors (straight walk, stand, in-place turn) are only learned if the sampler gives them mass. Product-of-uniforms distributions concentrate mass on mixtures and give none to the pure behaviors users actually command.
Applies when
- a "simple" command (straight, stop) underperforms mixtures in sim
- designing command distributions for velocity-tracking tasks
- a ladder needs per-mode isolation for attribution
“纯直行是零测度点:最终 command 为 vx∈[0.15,0.5] × vy∈U(±0.2) × wz∈U(±0.6) 连续均匀,vy=0∧wz=0 从未被专门采样 —— "sim 里直行就漂"是分布的必然,不是 reward 没调好 … Isaac 原生 UniformVelocityCommand 是三轴各自 uniform,采不出"离散模式桶"”
train/OMNI_V0_SPEC.md § 0. 为什么从零 (3) / 三件前置 (1) Oversized lateral COM randomization (+/-5 cm) deliberately forces leg spread
com-randomization-forces-leg-spreadDR ranges can be behavior-shaping tools, not just robustness padding: oversize a randomization axis to force a strategy the reward struggles to express - and expect a compensating behavior to appear as the cost.
Symptom
Feet drift toward the centerline and even collide; policy has no incentive to keep a lateral support base.
Context
COM randomization ranges were chosen asymmetrically by axis: lateral +/-5 cm ("比常规大,故意的" - larger than usual, on purpose), fore-aft +/-2 cm, vertical +/-2 cm. The oversized lateral range is not robustness padding but a behavioral forcing function. Lucen logged it as directly relevant to its own roll-channel / sideways leg-kick symptom.
Change
Set COM randomization to lateral +/-5 cm, fore-aft +/-2 cm, vertical +/-2 cm, with the lateral band intentionally oversized to make narrow stances fail during training.
Outcome
Effective at separating the feet on the reference robot; side effect - the base began swaying left-right, which then required a foot-centerline distance penalty (see reward-chain-foot-height-landing-spacing).
Mechanism
Randomizing COM laterally makes narrow-stance policies fall for some draws, so PPO discovers wide stances as the only strategy robust across the band - DR used as an implicit reward. The sway side effect appears because the policy hedges against unknown COM by active lateral correction.
Applies when
- feet too close / self-collision in a learned gait
- roll-axis instability suspected to come from narrow stance
- choosing COM or mass-offset DR ranges
“两脚太近甚至互撞 → 先试质心横向随机化 ±5 cm,逼迫策略把脚分开;有效但引发新问题——基座开始左右摇摆 … 横向 ±5 cm(比常规大,故意的,用来逼出分腿)/ 前后 ±2 cm / 垂直 ±2 cm”
Experience.md § 质心随机化范围 (lines 75, 84-86) Choose the fork root by which candidate's shortfalls are recoverable, not by headline score
fork-root-recoverable-shortfallWhen picking a checkpoint to fork from, rank candidates by whether their weaknesses are trainable-back, not by current headline metrics; prefer the candidate whose deficits the upcoming training directly pays for.
Symptom
Multiple candidate checkpoints for the omni-command ladder root, each best at something different: fric-3000 had the best tracking precision (vx 88-91%) and hardened plant robustness; s1e-500 had lower precision (vx 82%) but was the only candidate that could still walk backward.
Context
Root selection ran as a data probe, not a preference vote: 6 candidates x 8 out-of-distribution omni commands x 20 seeds = 960 cells (probe_omni_0808.json). s1e-500 @pw1.0 survived 20/20 in all eight conditions including backward at 67% tracking; the deep-trained fric lineage scored backward 0-3/20 despite better forward precision.
Change
Decision criterion made explicit: list what each candidate exclusively wins at, then ask which of those wins the loser could train back. fric-3000's exclusive wins (precision, plant robustness) are both retrainable - precision is directly optimized by the reward, plant hardening is a planned later pass. s1e-500's exclusive wins (backward plasticity 20/20 vs 2/20, disturbance margin 159/160 vs 125/160, push chirality symmetry 40/40 vs 17/40) had all been shown unrecoverable - push-level rungs failed twice, chirality never recovered even with mirror augmentation on. Root = s1e-500.
Outcome
s1e-500 carried the whole C ladder; its backward skill was preserved through C2/C4 gates (regress budget <=2/20 enforced), and the final C4 product passed a 260-cell battery at 20/20 everywhere.
Mechanism
Training can re-earn anything the objective directly pays for, but capabilities that earlier training destroyed and never restored (plasticity, symmetry, robustness margins) are empirically one-way doors. The information-bearing comparison is therefore recoverability of each candidate's deficit, which the team stated as "独占项的可恢复性正好相反 —— 这就是判据" (the exclusive items' recoverability is exactly opposite - that is the criterion).
Applies when
- selecting a resume/fork root among several checkpoints
- one candidate is more precise but another retains a skill the rest lost
- planning a task-extension ladder from an existing lineage
“fric-3000 赢在精度(vx 88~91%…)与 plant 鲁棒性 → 两样都训得回来…;s1e-500 赢在可塑性(C1 20/20 vs 2/20)、抗扰余量(159/160 vs 125/160)、手性对称(推 ±6 N·s 40/40 vs 17/40)→ 三样都训不回来 … 独占项的可恢复性正好相反 —— 这就是判据。”
train/C_LADDER_RUN.md § 0. 为什么根是 s1e-500(数据,不是偏好) Mirror augmentation over an asymmetric default injects systematic error - symmetrize the default first and verify the transform bit-exact
mirror-augmentation-needs-symmetric-defaultBefore enabling any symmetry augmentation, make every constant inside the observation encoding exactly symmetric, and validate the mirror transform against forward kinematics to machine precision - an unverified augmentation is a new error source, not a regularizer.
Symptom
Mirror data augmentation was about to be added while both default poses (standing_pose, walk nominal_pose) were asymmetric - stale hand-tuned compensations from before a ground re-calibration, with hip_yaw differing 2.40 deg between sides and the foot soles actually tilted (pitch 2.88/1.35 deg, roll -2.47/+0.25 deg).
Context
The observation encodes joint_pos_rel = q - default. Under mirroring q_l -> -q_r, the relation (q-default)_l -> -(q-default)_r holds only if default_l = -default_r; with an asymmetric default, augmentation produces observation pairs that are NOT mirror images, i.e. "default 不对称时做镜像增强会引入系统性错误,比不做还糟" (worse than not doing it). The fix: adopt model geometric zero as standing default (MuJoCo FK verified: sole pitch/roll exactly 0, asymmetry 0.00 deg) and a symmetric crouch for walk (hip -0.25/knee -0.5/ankle -0.25 satisfying hip - knee + ankle = 0 to keep soles flat). The mirror transform itself was verified bit-exact before use: pseudovector vs polar-vector sign patterns (ang vel [-1,1,-1], gravity [1,-1,1], cmd [1,-1,-1]), joint swap-and-negate; FK check that left-foot pose under q equals the mirror of right-foot pose under mirror(q), measured error 0.00e+00.
Change
Defaults symmetrized first (with init heights recomputed by FK), stand policy retrained on the new default so both policies share one default; augmentation enabled only after the FK mirror test passed.
Outcome
stand_v1 achieved exact left/right pairing (l_knee -0.1013 / r_knee +0.1013), six-pair asymmetry 0.0 deg, height fluctuation 7 -> 1 mm, 33% less mean |action|.
Mechanism
Augmentation asserts an equivariance of the observation encoding; any asymmetric constant inside the encoding (the default) breaks the asserted symmetry, so the augmented data teaches a false invariance. Verifying the transform against FK geometry tests the assertion end to end, independent of the training stack.
Applies when
- adding mirror/symmetry augmentation to locomotion training
- defaults or trims were hand-tuned per side at any point
- observations are expressed relative to a default pose
“观测里 joint_pos_rel = q − default。镜像下 q_l → −q_r,要让 (q−default)_l → −(q−default)_r 成立,必须 default_l = −default_r。default 不对称时做镜像增强会引入系统性错误,比不做还糟。… 位置误差与姿态矩阵误差实测均为 0.00e+00。”
train/RETRAIN_v2.md § 2. 前提:default 姿态必须先对称化(不是可选项) / 3. 镜像变换 Prone get-up sat at 0/159 until two gated hinge terms moved the seated feet - first sideways (561 -> 360 mm), then fore-aft (-168 -> +56 mm) - and success went to 158/159 with nothing else changed
prone-dead-end-is-foot-placementWhen a stuck state and a successful state differ geometrically, penalize the discriminating quantity with a gated hinge that is exactly zero in the state the policy actually reaches (measure it - not the nominal), then re-probe: flattening one axis can move the discriminant to another.
Symptom
Prone falls always righted and then sat with the feet splayed wide or tucked behind the hips, from where the policy never stood (0% for four generations).
Context
Four lines of evidence pointed at foot position: the configuration probe (ankles 215 mm apart stood 52.3%, 561 mm apart 0.0%); FK showing the action contract's nominal (a = 0) is itself a 465 mm straddle, so the action_rate and still terms were pulling toward the splits; biomechanics (feet tucked under the body cut peak hip-extension torque 148.8 -> 32.7 N*m, -78%); and HoST's foot-displacement term, which this reward table lacked. The earlier "not a reward hole" reading was corrected to "a gradient hole, not a level hole": at the dead point the heaviest term (upright) was saturated with zero gradient, still paid for not moving, and the one live gradient (base_height) pointed at the thigh-horizontal torque barrier. A prone ROM scan had already ruled out pushing up from prone.
Change
R0.4: feet_spread_excess = clamp(ankle distance - 0.215, 0, inf) x upright gate, weight -2.0, plus a height-decay factor added after measuring that the policy's real standing stance was 406 mm, not the 215 mm nominal (the plain version would have taxed every successful stand 0.38/s). R0.5: the same shape on the fore-aft axis, feet_fore_seated = |fore-aft offset - 0.05| x upright gate x height decay, target +50 mm (the measured natural offset of standing postures). One variable per rung.
Outcome
R0.4: seated ankle distance 561 -> 360 mm, supine/side exactly unchanged, prone 0 -> 1.9%, mid 45.9 -> 62.2%; a probe then showed the discriminant had moved to the fore-aft axis (standing starts +42 to +51 mm, the prone seat -168 mm). R0.5: supine 99.4, prone 99.4, side 100, mid 100%, re-falls 0%; both geometry terms collapsed to ~0 near iteration 13,100 as base_height rose, and the prone fore-aft offset went -168 -> +56 mm - the term's own target, closing the causal chain. The cost, unmeasured at the time: action jitter rose 33% (sum |da|^2 6.82 -> 9.06).
Mechanism
An upright-gated hinge is inert while the robot rolls and exactly zero in the achieved stance, so it adds gradient only inside the stuck basin; a seated robot with its feet behind or outside its COM must make a kinematically unfavourable transition to stand, and moving the feet under the body removes it.
Applies when
- a get-up or transition skill fails from one start category only
- successful and failed episodes differ in a measurable geometric quantity
- a shaping term might tax the posture successful episodes already use
“`recovery_r0_5`,唯一变量 = 追加 `feet_fore_seated`(与 R0.4 同形状,只换测量轴)。 … 对照 R0.4 的 prone(3.1%,360 mm,**−168 mm**):前后偏移从 −168 走到 +56, 正是这一项的目标量,**判别量被消掉后成功率随之到顶** —— 因果链完整。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §20 R0.5(前后向脚位置项):成功率门全过 —— prone 0/159 → 158/159 Foot dragging is an attractor, not a low amplitude - and joint damping is the mode switch, adjustable at deploy time
swing-bistability-damping-switchWhen a quality metric is bimodal, stop treating it as an amplitude to be trained up: map the modes against initial conditions and plant parameters, find the parameter that switches basins, apply it first as a deployment lever, and only then bake it into the training distribution (as a plant-family shift, never as an execution-mapping change).
Symptom
s2e_pd-1400's swing height "median 12.1 mm" hid a perfect bimodal distribution: 20 seeds split into a drag mode (2.6-4.9 mm) and a step mode (19.3-24.0 mm) with NOT ONE seed in between - the median sat in the empty gap, and "swing debt -11 mm" really meant "50% probability of falling into the drag attractor".
Context
Two designed experiments closed the mechanism. Test A (nominal plant, 40 seeds): step 42% / drag 58% / middle 0 - at nominal gains, initial conditions alone pick the mode, both modes 100% survivable. Test B (fixed init, kp x kd grid): kd is the mode SWITCH - at kd 1.3 all surviving cells step (13-22 mm), at kd 0.7 nearly all drag (2.7-4.3), only at kd 1.0 does init get a vote; kp >= 1.2 is dangerous (5/6 falls). Global verification at kd x1.3 (20-seed, delay 2): survival 20/20 at ZERO cost, step share 42 -> 80%, swing median 12.1 -> 18.4 mm, slip record low 334, thicker tilt margin - costs: vx 85 -> 78%, saturation +5 pp. A Pareto sweep then priced the knob: step share 42/72/75/88/82/90 across kd 1.00-1.30 with a linear vx tax of -2.3 pp per 0.1 kd - the basin gain is fully collected at kd 1.20 ("1.30 是 over-damping 纯多付税"). Mechanism: low damping leaves a landing micro-oscillation / ground-slide channel the policy can exploit to drag; damping plugs the channel.
Change
Deployment lever adopted: kd-scale 1.20 (conservative 1.15) as the legitimate successor to the power-0.8 crutch ("前者削幅度保稳,后者堵 拖地通道换步态,且不牺牲存活"); training-side prescription: move the DR band to nominal-1.2 x (0.9,1.1) = [1.08,1.32], deleting the [0.7,1.0) drag-teaching zone - a contract-level change requiring digest re-baselining, gated on measuring the real robot's actual kd dispersion first.
Outcome
The kd surgery rung (s2e_kd) delivered basin 8 -> 11/20, slip 405 -> 331, vx 81 -> 85% with no out-of-band fragility (below-band check 20/20) - "拐杖烧进分布的正确姿势", explicitly contrasted with the failed s1g amplitude version: this one changes the plant family the policy has seen, that one changed the execution mapping the policy would have to relearn.
Mechanism
The gait's swing behavior is a bistable dynamical system whose basin boundaries are set by plant parameters; a policy trained across a kd band that includes the drag basin has learned to inhabit it. Shifting the deployed (and then trained) damping moves the system into the step basin without touching the policy - a plant-side fix for what looked like a training deficiency.
Applies when
- a gait quality metric splits into distinct modes across seeds
- deciding between more training and a gain/damping change
- converting a deployment crutch into a training-distribution change
“20-seed 里拖地模式 2.6~4.9mm 与迈步模式 19.3~24.0mm 各半,中间一个不落 … kd 是模式开关——kd1.3 下 6/6 存活格全迈步 … kd0.7 下几乎全拖地 … swing 债的解(至少大半)在部署端阻尼档,不在训练端 … 机理:低阻尼下落脚微振荡/贴地滑给了策略顺势拖行的通道,加阻尼堵之。”
train/README.md § swing 双稳态定性 + kd 部署杠杆 (2026-08-07, 用户设计 Test A/B) A single run's drift direction may be a limit cycle, not a policy bias - check the sign distribution across seeds
multiseed-sign-test-for-driftDistinguish "bias" from "broken symmetry limit cycle" by the sign distribution over many seeds; report drift as (mean, sign split), and never compare single-run drift magnitudes across versions.
Symptom
Net yaw over 15 s appeared to worsen from -41 deg (v2) to -84 deg (v4), inviting the conclusion that the new version drifted more.
Context
The Isaac-side view across 32 environments told a different story: per-env yaw was mixed-sign (20 negative / 12 positive) with mean ~0 - the drift is a limit cycle whose direction depends on initial conditions, not a systematic policy bias. The single MuJoCo run had sampled one draw from that distribution, so its magnitude could not be compared across versions as if it were a property.
Change
Evaluation rule: before classifying drift as systematic, run multiple seeds and examine the sign distribution; single-trajectory drift magnitudes are samples, not properties.
Outcome
The v2-vs-v4 drift "regression" was reclassified as not-established; later drift work (hip_roll l+r bias) used cross-policy, cross-seed evidence instead.
Mechanism
Symmetric dynamical systems can settle into either of two mirrored limit cycles; the selected cycle is decided by noise and initial state. A statistic whose sign is initial-condition-dependent has no meaning as a single sample - only its distribution does.
Applies when
- comparing heading drift or lateral drift across policy versions
- a symmetric-looking behavior shows a consistent direction in one run
- deciding whether to fix "drift" in reward or calibration
“偏航反而变差(−41° → −84°):注意 Isaac 侧 32 env 的逐 env 偏航是正负混合(20/12)、均值 ≈0,说明这是极限环性质(方向随初值)而非策略偏置 —— MuJoCo 单次跑测到的是分布里的一个样本,不能当作系统性偏差。要判断需多种子统计。”
train/WALK_DIAGNOSIS.md § walk_v4 独立验收 读法 (偏航) Size each joint's action authority to its measured working range - lock a channel to zero only when its job is provably elsewhere
per-joint-action-scale-lockdownSet per-joint action scales from measured target ranges and momentum decompositions: full authority for working channels, working-range authority for balance channels, zero for channels whose contribution is measured negligible - and prefer this structural quieting over perpetual reward penalties, keeping the contract dimensions intact.
Symptom
"Walks crooked" - roll and yaw channels wandered; hip_yaw peak-to-peak reached 22.3 deg in v11 while contributing essentially nothing to locomotion; a uniform action scale of 0.5 gave every joint the same authority regardless of its actual job.
Context
The v12 design replaced the scalar action scale with per-joint scales justified by measurements: pitch-class 0.5 (the gait's entire working channel - untouched); hip_yaw 0.0 - lossless because it carries only 1.4% of yaw momentum (v6 decomposition) and straight-line targets use merely +/-0.006-0.03 ("乱动纯属浪费"); roll 0.2 - NOT zero, because lateral balance and weight transfer are roll's unique job (locking it would degenerate into foot-edge rocking, "比现在更歪"), and 0.2 covers the measured working range +/-0.06-0.19 while "0.5 的另一半全是 '歪歪扭扭'的来源". Costs were accepted consciously: turning demoted to an observation row with a fallback (yaw 0 -> 0.1 in v13). Contract preserved: the 12-dim action interface unchanged, yaw values simply neutralized. The structural lockdown also RETIRED the reward-side hip_yaw_quiet penalty - "A 的 yaw=0 结构性取代,不再付奖励塑形成本".
Change
action_scale_joint pitch 0.5 / roll 0.2 / yaw 0.0 wired through robot.yaml -> policy_io (verified: all-ones action gives hip_yaw target exactly 0) -> Isaac action term, guarded by the contract checker ("它就是抓这种双侧不一致的").
Outcome
Designed and verified on the shared side before the lineage freeze; stands as the pattern for authority sizing: structure replaces reward shaping wherever a channel should simply not act.
Mechanism
Action scale is a per-channel authority budget; uniform budgets give noise channels the same voice as working channels, and reward-side quieting then pays a permanent shaping tax for what a zero scale provides for free. But zeroing is only lossless when decomposition proves the channel's contribution negligible AND no unique function (balance) lives there.
Conflicts
Wired and verified on the config/deploy side but never trained - the 2026-08-05 reset suspended v12 before the Isaac-side run.
Applies when
- some joints wander without contributing to the task
- a quieting penalty (deviation/L1) taxes every step forever
- deciding action-space authority for a new task or robot
“yaw=0 是无损的:实测它只贡献 1.4% 偏航动量、直行目标只 ±0.006~0.03,乱动纯属浪费。… roll 不能为 0:横向平衡/重心换脚是它的独有职责,锁死会退化成脚缘摇摆(比现在更歪)。0.2 的依据:各代实测 roll 目标只用 ±0.06~0.19,0.5 的另一半全是"歪歪扭扭"的来源。”
train/WALK_V12_SPEC.md § 3. A —— 逐关节动作幅度(用户"只动 pitch"的安全版)