Training Coach
Doctrine
A report may cite any of these as doctrine-N.
doctrine-1Contract freeze and fingerprint disciplineThe policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.
Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).
Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.
doctrine-2Attribution by resolved training params - never eval-override knobsCapability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.
Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).
Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.
doctrine-3PASS gates become constraints; FAIL gates become objectivesOnce a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.
Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.
Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".
doctrine-4One variable per ladder rung - counted against what the checkpoint sawA rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.
Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.
Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.
doctrine-5Pre-register risks, readings, and stop criteria before the ladderBefore a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.
Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).
Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.
doctrine-6Plant parameters are measured, never inventedEvery plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.
Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).
Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.
doctrine-7Sim2sim gate before sim2real - under deployment conditionsEvery checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).
Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).
Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.
doctrine-8Observation honesty - the actor's inputs are a hardware contractThe actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.
Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).
Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.
doctrine-9Reward economics are audited in realized currencyReward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.
Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).
Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.
doctrine-10The zero-cost option must be the desired behaviorFor every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.
Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).
Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.
doctrine-11Measurement discipline: independent referees, signs, distributionsA disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.
Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).
Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.
doctrine-12The deployment pipeline is plantIrreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.
Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).
Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.
doctrine-13DR budget is finite; its distribution is the measured supportRobustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.
Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).
Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.
doctrine-14Gates measure what hardware feels: posture, margins, stripped assistsAcceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.
Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).
Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.
doctrine-15Fork and root selection: recoverability, maturity, frozen rewardsChoose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.
Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).
Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.
doctrine-16Curricula: verified engagement, lineage counters, disease-phase gatingAutomatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.
Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).
Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.
doctrine-17Probe before training: feasibility first, hypotheses in tablesAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.
Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).
Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).
doctrine-18External advice is recomputed locally; values transfer as ratiosEvery external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.
Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).
Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.
doctrine-19Hardware sessions are scripted experiments, not tuning sessionsReal-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.
Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).
Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.
doctrine-20Close questions in writing; restart when the debt is structuralAudited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.
Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).
Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.
doctrine-21Name the quantity in the space it lives inA goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.
Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).
Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.
doctrine-22Continuation needs a live gradient; a release is chosen by a scanContinue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.
Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).
Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.
Experience cards
102 cards matching “stand-gate-posture-not-survival”.
A nonzero response with the same sign for + and - commands is bias, not ability
same-sign-response-is-yaw-biasBefore crediting any directional skill, test both command signs: response must flip sign with the command; a same-signed pair is a bias to subtract, not an ability to report.
Symptom
Root-selection probe showed nonzero wz "tracking percentages" on turn commands, tempting the read that candidates could partially turn.
Context
During C-ladder root selection, s1e-500's measured yaw rate was +0.084 rad/s for cmd +0.3 and +0.093 rad/s for cmd -0.3 - same sign both ways. The same check on the C2 baseline gave wz+0.20 -> -0.13 and wz-0.20 -> +0.12 (again same sign), while the alternative root s2e_pd-1400 gave +0.16 / -0.16 - opposite signs, i.e. a genuine 16% command response.
Change
Reading corrected and written into the execution sheet: percentages on directional commands are meaningless unless the +cmd and -cmd responses have opposite signs; all three candidates were re-classified as "cannot turn, cannot sidewalk - C2/C3/C4 learn from zero". Acceptance criteria thereafter required "tracking >=50% AND left/right opposite-signed".
Outcome
Prevented crediting turn/sidewalk ability that did not exist; the antisymmetry clause became a standing part of every turn and sidewalk PASS condition (C2, C4, C4-redo levels all carry "且左右反号").
Mechanism
A constant yaw (or lateral) bias projects onto any command's sign convention and shows up as fake fractional tracking; only sign-antisymmetry under command reversal distinguishes a feedback response to the command from an open-loop offset.
Applies when
- evaluating turn/sidewalk/any signed-command tracking percentages
- a candidate shows partial tracking on an axis it was never trained on
- writing PASS criteria for a new directional skill
“C2/C3 那些非零的 wz 百分比不是转向能力 —— 转向+ 与 转向− 的实测同号(s1e:cmd +0.3 → +0.084,cmd −0.3 → +0.093 rad/s),那是恒定偏航偏置。… 三个候选都不会转、都不会侧走。”
train/C_LADDER_RUN.md § 0. 读数纠正(重要,别引错) A joint frozen at the action clamp pays zero action_rate forever - penalize pre-clip saturation to make the cheat cost money
saturation-cheating-zero-rate-costWhenever actions are clipped and any smoothness/rate penalty exists, add a pre-clip saturation penalty so living at the clamp costs more than oscillating - and audit for frozen-at-clamp joints (action std ~0, |a| at exactly the clip value) as a standing acceptance row.
Symptom
With action_rate_l2 raised to -0.2, walk_v7's hip_pitch actions froze at exactly +/-1.000 (the clamp), reproduced bit-for-bit on hardware (splits frozen at +/-0.35 rad); the gait-shaping term joint_pos_ref collapsed to 0.026-0.035. The repo had died in the same trap once before (walk_v0: four joints pinned at +/-1.0).
Context
Mechanism: a joint pinned at the clamp has action-rate cost exactly zero and forever zero - under a strong smoothness tax, "push to the clamp and freeze" becomes the dominant optimum. Lowering the weight (-0.2 -> -0.1) only reduces temptation; the frozen state still costs nothing, so the structural fix adds action_saturation = sum(relu( |a_raw| - 0.9)) at weight -1.0, computed on the PRE-clip network output - post-clip, |a|=1.01 and |a|=3 punish identically and the out-of-range gradient dies (v0's old disease: mean |a| 1.71 soaked in saturation). Economics: freezing at |a|=1.0 now pays 0.1/joint/step (two hips = 40% of alive) vs ~0.0004/step for the healthy reference oscillation - the cheat flips from free to ~250x negative. Honest limits were recorded: A1 does not forbid freezing at 0.89 (the anti-freeze pressure must come from the oscillation demand of joint_pos_ref), and the alternative "rate on post-clip target" was rejected as 换汤不换药 - a pinned target also has zero rate.
Change
v8-A: add action_saturation (-1.0, thresh 0.9, pre-clip) AND halve action_rate_l2 (-0.2 -> -0.1, still 3.3x the v5 value); success criterion pre-declared (joint_pos_ref telemetry returns to v6 scale).
Outcome
Booked as the structural repair of the v7 freeze; also fixed a config hygiene trap discovered on the way - action_rate was assigned twice in __post_init__ (v5 comment line then v7 line), merged to one assignment "别再留两处赋值给下次审计埋雷".
Mechanism
Clipping creates a zero-gradient, zero-cost absorbing region in action space; any penalty on action derivatives makes that region strictly optimal once entered. Only a penalty on clamp proximity itself (measured pre-clip so depth of violation is visible) restores a slope out of the absorbing region.
Applies when
- joints sit at exactly the action clip with near-zero variance
- raising a smoothness penalty degrades gait amplitude
- shaped-oscillation terms collapse after a rate-weight increase
“钉死在钳位的关节 action_rate 代价精确为零且永远为零;−0.2 之下"推到钳位冻起来"成了压倒性最优 … 本仓第二次栽在同一坑(walk_v0 死于四关节钉死 ±1.0)。回调权重(−0.2→−0.1)只降低诱惑不消除作弊 … 算在 clip 前的原始网络输出上 … 作弊收支从"白赚"变成"倒贴 ~250 倍"。”
train/WALK_V8_SPEC.md § 1. 改动 A — 治饱和作弊 Every new penalty ships with a pre-registered withdrawal clause - if healthy gait must pay above the cap, the term stands down
calibration-threshold-with-withdrawal-clauseIntroduce every new penalty with: the zero-cost-option audit, a replay-calibrated weight formula (healthy pays a fixed small fraction of tracking), and a pre-registered withdrawal condition - and let the clause fire without argument when the calibration says the term cannot be afforded.
Symptom
Three same-shaped crashes had established a failure archetype: v4's clearance, v8a's landing window (weight off by 58x uncalibrated), and v6a's bare hip_yaw suppression all combined a zero-cost "don't move" option with a fee on any motion - a reverse barrier that pushes policies toward standing still.
Context
The v11 landing-window penalty was therefore introduced under a calibration-threshold protocol: (1) shape chosen with the window tightened (h_gate 0.03 -> 0.02, because 0.03 equaled the clearance target and priced the entire descent); (2) weight from a FORMULA, not judgment: measure the term's raw value on healthy replays (v5/v10b), set w = -(0.10-0.15 x tracking reward) / raw_healthy; (3) withdrawal clause pre-registered: if healthy gait must pay >15% of tracking no matter the tuning, the term is withdrawn to the next version rather than forced in - "不硬上". The companion hip_yaw quieting term ran the same protocol (calibrate on replays, healthy pays <=5%) and was later retired entirely when a structural fix (zero action scale) made its shaping tax unnecessary.
Change
Penalty introduction protocol: shape audit (what is the zero-cost option?), replay-based weight formula, healthy-pay cap with a written stand-down condition - all before training.
Outcome
The landing term was in fact withdrawn under its clause (v12 records "P5 落地窗口罚 已撤 … 维持撤下"), demonstrating the protocol firing as designed instead of the fourth same-type crash.
Mechanism
A penalty's damage mode is mispricing healthy behavior; since the healthy price is measurable in advance on replays, both the weight and the go/no-go decision can be computed rather than discovered by a ruined training run. The withdrawal clause converts "make it work" pressure into a clean deferral.
Applies when
- adding any motion-taxing penalty to a working gait
- a proposed term's weight has no measurement behind it
- a previous same-shaped term crashed training
“权重公式而非拍脑袋:先在 v5/v10b 回放上量 h_gate=0.02 的原始值,w = −(0.10~0.15 × 跟踪奖励) / raw_健康;标定门槛:若健康步态无论如何要付 >15% 跟踪,本项撤下留 v12,不硬上 (v4 clearance/v8a-B/v6a 三次同型翻车的教训:代价为零的"不动"选项 + 一动就收费 = 反向壁垒)。”
train/WALK_V11_SPEC.md § 6. P5 —— 落地窗口罚(三代欠账,标定门槛制) A joint-velocity penalty meant to slow the get-up cut joint speed 16% and left the get-up time unchanged - the knob never moved the variable, so the idea it was meant to test stayed untested
dof-vel-penalty-is-not-a-pacing-knobBefore reading a result as a test of an idea, check that the knob actually moved the independent variable; velocity regularizers smooth a schedule they do not set, and a schedule driven by per-step task income moves only when that income's time structure does.
Symptom
The get-up took 0.6-0.9 s in Isaac with large torque demand; the user proposed getting up more slowly so less torque would be needed.
Context
The idea had support in the accounts: the acceptance bound is an upper bound of 5 s (5-8x margin), the quasi-static squat path peaks at 25% of the limits, and rolling over needs no momentum. R3.2 raised dof_vel from -1e-3 to -5e-3 as the single variable.
Change
dof_vel -1e-3 -> -5e-3 (child-run from R3.1).
Outcome
Get-up medians moved +0.02-0.04 s (noise); raw joint velocity -16%; torque demand median got worse (hip_pitch 46-48% -> 63-67%) as the new term competed with torque_headroom on the same joints; MuJoCo 98 -> 96%. Verdict FAIL on the knob, not on the idea, and the rung was not adopted. When pace was later attacked through the income's time structure (V2.5/V2.5b), the MuJoCo get-up moved into the 3.5-4.5 s design band.
Mechanism
The pace was set by base_height_progress paying for every step spent high (stand earlier, earn more); a velocity regularizer only smooths motion along the same schedule and does not change when the robot stands up.
Applies when
- trying to make a skill slower or gentler with smoothness penalties
- an experiment's primary metric did not move and a verdict is being written
- two penalties act on the same joints
“**关键判读:`dof_vel` 罚只把关节速度压了 16%,而起身用时一点没变。** 也就是说**这一级根本没有把"慢下来"这个自变量推动起来** —— 所以它**不构成对 用户假说的检验** … 起身节奏由 `base_height_progress` 的逐步计酬决定(早站起来就多 拿),速度正则只在同一条时间轨迹上把动作抹匀,不改变何时站起来。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §25 R3.2(dof_vel −1e-3→−5e-3,慢一点起身) mj_objectVelocity returns inertial-principal-axis frame - one API assumption poisoned eval and observations for a whole line
body-frame-velocity-api-auditVerify every frame-sensitive API against a hand-computed truth (rotate raw qvel yourself, or command a known world velocity and check where it lands) before trusting any evaluation or observation built on it - especially when a model's inertial frame is rotated from its body frame.
Symptom
Sidewalk vy read ~0 under every condition; separately, whole-policy performance was mysteriously mediocre in sim2sim while training-side numbers looked fine. Four training rungs were declared FAIL partly on these readings.
Context
base_link's URDF inertial frame is rotated 90 deg about x relative to the body frame (iquat = [0.7071, 0.7071, 0, 0]). mj_objectVelocity(flg_local=1) rotates into ximat - the inertial principal-axis frame - not the body frame, and returns center-of-mass point velocity, not body-origin velocity. Consequences measured: the "vy" column was actually vertical velocity vz (walking at cmd 0.25: old reading +0.0093 vs true -0.0424); the angular velocity fed to the policy in sim2sim was [wx, wz, -wy] - a different quantity than Isaac and the real IMU provide. RMS check over 8 s of walking: y/z axes swapped between v6[:3] and the qvel truth.
Change
Fixed sim2sim and both probes to compute ang_b = qvel[3:6] and lin_b = xmat.T @ qvel[0:3] (identical quantity to Isaac's root_ang_vel_b / root_lin_vel_b), with a standalone reproduction script (frame_bug_repro_0809.py).
Outcome
Re-scoring the "failed" C4 lineage under correct coordinates reversed the verdicts: c4r4 checkpoints showed vy 80-126% tracking (old reading: +/-2%) and vx+0.30 at 91-95% where the old metric said 0/5 - the bad frame both mis-measured vy and, via corrupted policy observations, systematically depressed all measured performance. Final product passed 260/260 cells.
Mechanism
A simulator API's frame convention is part of the observation contract; when the model's inertial frame is rotated relative to the body frame, frame-agnostic use of a "local" velocity silently permutes axes. Feeding a policy an axis-permuted angular velocity is an observation corruption that degrades behavior everywhere, not just on the axis being studied.
Applies when
- building or auditing a cross-simulator evaluation harness
- one measured axis reads near-zero under all conditions
- sim2sim scores are inexplicably worse than training-side metrics
- URDF/MJCF inertial frames are rotated relative to body frames
“base_link 的 iquat = [0.7071, 0.7071, 0, 0] … mj_objectVelocity 用的是这个 … 喂给策略的 base_ang_vel 是 [wx, wz, −wy] —— MuJoCo 侧观测与 Isaac / 真机 IMU 不是同一个量;vy_mean 报的是竖直速度 vz —— 前进 cmd 0.25 时旧读数 +0.0093,真值 −0.0424。”
train/C_LADDER_RUN.md § 3l. ⚠️ mj_objectVelocity 读的是惯性主轴系 / 3m. 一 bug 坐实 Tightening the bridge's rate limiter under an unchanged policy cut torque peaks 30-50% and made other things worse - the policy cannot see the limiter, keeps commanding and winds up; a deploy-side limiter is a safety net, not a cure
deploy-rate-limiter-windupA rate or torque limiter added at deployment lowers peaks but the policy still commands as if unconstrained (saturation, windup, new contacts); use it as a safety net mirrored in evaluation, and put the constraint where the policy can learn around it.
Symptom
After the violent first real-robot get-up, the cheapest candidate fix was to tighten the bridge's slew (rate) limit for the recovery policy without retraining.
Context
Probe on R3.1 in MuJoCo (5 categories x 3 seeds, mu 1.0), monkeypatching the limiter with no repository change: TIGHT = RS06 4.0 / RS02 3.0 / RS00 2.0 rad/s (about 0.08/0.06/0.04 rad per policy step) against the current vel_limit setting.
Change
The probe decided the role of the limiter rather than a deployment.
Outcome
Success 14/15 -> 12/15; get-up median 2.35 -> 3.53 s (max 9.30); torque demand peak median hip_pitch 164% -> 111%, knee 166% -> 86%; action saturation still 100%; leg-leg contact 558 -> 860 frames. The limiter was kept only as a real-robot safety net (mirrored into sim2sim evaluation); the cure moved into training - where the next lesson was that a limiter anchored on the last command is itself an integrator (slew-anchor-is-an-integrator).
Mechanism
A policy that never trained with the limiter keeps issuing the targets it learned; the limiter clips them, the target window runs ahead (windup), and the robot follows a trajectory the policy never evaluated.
Applies when
- a trained policy is too violent on hardware and a quick deploy-side fix is tempting
- adding slew, torque or velocity limits in a bridge or firmware
- evaluation and deployment use different limiter settings
“判读:**链路侧收紧立等可取地把 τ 峰值砍 30~50%,但成功率掉、饱和率仍 100%、 腿-腿接触反升** —— 策略感知不到限速器,目标窗口继续狂奔。⇒ 收紧 slew 只配当 **真机侧安全网**(必须同步进 sim2sim 口径,基础设施现成),**不配当治法; 治法必须进训练**。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 探针:收紧桥层 slew,r3_1 不重训直接测 An edge-triggered landing penalty missed the tail and fired after the harm - penalize overspeed continuously inside the contact window
penalize-tail-before-touchdownPenalties aimed at impact/violation events must (a) price the excess over a threshold, not the mean, and (b) be active on the approach (state-gated window), not triggered by the event - check your control rate can even see the event you are penalizing.
Symptom
The v7 landing penalty (vz^2 on the contact-force rising edge, weight -10) did not bite: landing-velocity 95th percentile stayed at 2.61 m/s against a 0.3 target.
Context
Two structural faults were identified: (1) it penalized the MEAN over sparse events - many soft landings dilute the occasional violent slam, while the damage (GRF peaks, motor peak load) lives in the tail; (2) it fired AFTER touchdown - at 50 Hz evaluation the rising edge is aliased by physics decimation, so the read vz is often the already-decelerated post-impact value: underestimated, and with no shaping gradient before contact. Replacement: continuous penalty while the sole is inside a height gate (h < 0.03 m): relu(-vz - 0.30) - only the excess over an allowed approach speed is penalized (tail only), and gradient exists for several frames BEFORE touchdown. The sole-height computation again subtracts the 0.0585 m link offset ("WALK_DIAGNOSIS 坑#1, 别再踩"); the edge-triggered version was kept as a diagnostic only.
Change
feet_landing_vel reformulated: edge-event vz^2 -> in-window relu(-vz - v_ok) with v_ok 0.30 (conservative vs the sqrt(L)-scaled human value ~0.19, to be tightened after passing), h_gate 0.03, weight unchanged -10.
Outcome
The failure analysis of the first form was written before the second was trained; the v_ok escalation path (0.30 -> 0.45 if the robot becomes afraid to land) was pre-registered in the risk table.
Mechanism
Sparse-event mean penalties optimize the average case while the constraint is a quantile; and any penalty evaluated only at/after a discrete event gives the optimizer no gradient along the approach trajectory that determines the event. A state-gated continuous excess penalty fixes both: it prices only violations and shapes the approach.
Applies when
- impact/landing penalties fail to move tail percentiles
- a penalty is triggered by contact edges at a coarse control rate
- designing constraint-style penalties for rare violent events
“罚的是均值路径:上升沿是稀疏事件 … 大量软着陆稀释偶发猛砸;而伤害在尾部 … 罚在触地后:50 Hz 评一次,上升沿被物理 decimation 混叠,读到的 vz 常是撞完已减速的值——既低估,又没有触地前的塑形梯度。”
train/WALK_V8_SPEC.md § 2. 改动 B — 落地惩罚改罚尾部、罚在触地前 Audit rewards by realized contribution (weight x achieved value) - a weight of 2.0 was really paying 0.04
realized-contribution-auditEvaluate a reward table by each term's realized per-step contribution under the current policy, never by its weight column; if a term's realized value is ~0, escalating its weight is a no-op - change the term's structure instead.
Symptom
Foot dragging persisted through repeated weight escalation: feet_air_time had been raised 0.25 -> 1.0 -> 2.0 across versions with no behavioral change, and the training-side comment even recorded the fact ("几乎没有 单支撑相, 是拖着脚蹭") without the fix changing form.
Context
Computing realized per-term contributions in the trained state exposed the economy: feet_air_time contributed weight 2.0 x achieved 0.019 = 0.038 per step, against tracking's +1.20 - lifting the leg earned 3% of what tracking earned, so dragging was the rational optimum no matter the weight escalation. The same table acquitted the energy penalties (total -0.30 negative vs +1.74 positive) that a naive read of weights (-5.0 orientation!) would have blamed.
Change
Fix redirected from "raise the weight again" to "add a term whose realized contribution changes the optimum": a clearance penalty sized so its realized magnitude (~0.018/foot when dragging) is comparable to feet_air_time's, enough to flip the optimum without drowning tracking.
Outcome
With the term economy corrected (plus posture/range fixes), swing height reached 34 mm and tracking 87% by v6; weight escalation of the old term was abandoned.
Mechanism
A reward weight is only a multiplier on whatever the policy currently achieves on that term; when the achieved value is near zero (behavior absent), escalating the weight multiplies near-zero. Optimizer behavior is governed by realized per-step magnitudes, so audits must be conducted in that currency.
Applies when
- a behavior persists despite repeated weight increases
- auditing whether penalties are "too strong" or rewards "too weak"
- sizing a new reward term against existing ones
“把 feet_air_time 权重从 0.25 → 1.0 → 2.0 一路加,但没有加高度项。量级算下来:feet_air_time 权重 2.0 × 实得 0.019 = 0.038,而跟踪奖励是 1.2。抬腿的边际收益只有跟踪的 3%,拖地当然是最优解。… 正项 +1.74,负项 −0.30。能量惩罚不是瓶颈,抬腿没收益才是。”
train/WALK_DIAGNOSIS.md § 决定性证据(#1) / 各项奖励的实际量级 PPO's Gaussian noise cannot compose phase-locked oscillations - deliver them as feed-forward and let the policy learn the residual
feedforward-for-phase-locked-skillsIf a skill needs a temporally coherent (phase-locked) action component, do not expect step-wise exploration to find it: inject a verified feed-forward and train the policy as a residual stabilizer, keeping the feed-forward inside the deployment contract.
Symptom
Four different reward arrangements (no reference / wrong-sign reference / correct-sign reference / cage released) all failed to elicit sidewalk, while open-loop probes proved the behavior existed and was safe on the same platform with the same policy as base.
Context
Producing lateral velocity requires a phase-locked hip_roll oscillation synchronized to the gait clock. PPO's exploration is per-step, zero-mean, uncorrelated Gaussian noise - it can never compose a sustained phase-locked component, so the behavior is unreachable by exploration regardless of how it is rewarded. The fix changed the delivery channel: target = default + scale*action + lat_ff(cmd_vy, phi). The policy's action becomes a residual on top of the feed-forward, retaining full balance authority (it can even cancel the feed-forward); the feed-forward supplies exactly the component exploration cannot. This mirrors why the sagittal joint_pos_ref worked (it also delivered phase structure), just via a different channel.
Change
Contract-level change, done cleanly: new profile omni_ff (= omni + lat_ff_gain -0.5), existing omni profile bit-identical; feed-forward applied after the action delay stage; missing cmd/phase raises instead of silently dropping; deployment must use the same phi as build_obs (recomputing gives a one-tick phase misalignment).
Outcome
From C2-700, +100 iterations sufficed: product omni_c4_ff800 scored vy +120%/+125% (from +4%/-1%), 260/260 cells at 20/20 survival, zero old-skill regression, left/right gap 5 pp - the entire C4 saga resolved by changing the delivery mechanism, not the reward.
Mechanism
Exploration noise spans only the subspace its correlation structure can express; skills requiring coherent oscillation lie outside the span of i.i.d. per-step noise. Feed-forward moves the required structure into the action pipeline where it needs zero probability mass to appear, reducing the learning problem to stabilizing around a demonstrated behavior - which PPO does well.
Applies when
- a periodic/oscillatory skill trains flat under every reward variant
- open-loop injection of the behavior already works
- considering GRU/curriculum/exploration tricks for a rhythmic skill
“病因不在奖励,在探索形式:产生侧向速度需要相位锁定的 hip_roll 振荡,PPO 的逐步高斯噪声零均值无相关,合不出相位锁定分量。… target = default + scale·a + lat_ff(cmd_vy, φ)。策略动作因此是前馈之上的残差,保留全部平衡权限”
train/C_LADDER_RUN.md § 3j. C4-redo4:唯一变量 = 侧步参考改为前馈注入(契约级) When training fails repeatedly, inject the target behavior open-loop - stop tuning rewards for an unverified behavior
open-loop-probe-before-reward-tuningAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop on the real plant/sim first, and only resume training once you hold a measured, safe, sign-verified target trajectory.
Symptom
Three sidewalk training rounds failed; hypotheses multiplied (exploration failure / wrong reference waveform / insufficient authority) with no way to pick between them by running more training.
Context
Instead of a fourth reward guess, the team wrote probe_side_ref.py: the candidate reference is injected open-loop on top of a frozen policy's output (bypassing PPO entirely), directly measuring "what happens if the robot literally does this waveform" - separating all three hypotheses in one experiment (5 seeds x 8 s per condition, several waveform families and gains). The probe immediately eliminated the authority hypothesis (full-amplitude execution, 5/5 survival) and localized the problem to the waveform/measurement side. The closing principle was written down after the saga: without a verified target behavior, tuning rewards is "在黑暗里试钥匙" (trying keys in the dark).
Change
Standing method: before opening another training rung on a failing skill, build an open-loop (or task-space) generator of the intended behavior, measure whether the physical system can express it and what it looks like - then train toward a verified, quantified target.
Outcome
The probe chain produced the verified waveform (reversed-sign triangle, half gain), quantified safe amplitude (tilt 8.2 deg at band top, foot distance clear of the wall), exposed the metric bug when probe and training disagreed, and ultimately supplied the feed-forward that made C4 pass in +100 iters.
Mechanism
Training couples exploration, reward design, and feasibility into one opaque outcome; open-loop injection cuts the loop and tests feasibility and waveform alone. A behavior demonstrated open-loop converts the remaining failure into a pure credit-assignment/reward question - and its measured trajectory becomes the reference itself.
Applies when
- repeated training failures on one skill with multiple live hypotheses
- uncertainty whether the platform can physically express the behavior
- a reference trajectory's shape/sign/amplitude is guessed, not measured
“三轮 FAIL 之后不再猜,写 train/probe_side_ref.py 把参考开环注入到策略输出之上(绕过 PPO),直接量「照这个波形做会怎样」,一次分开三个假说:甲 探索 / 乙 波形 / 丙 权限。”
train/C_LADDER_RUN.md § 3f. C4 真因定谳(开环探针) / 3k. 建议的下一步 A suspended (no-load) test acquits or convicts the actuator before you blame authority
suspended-test-isolates-actuator-authorityBefore attributing a failure to actuator authority, measure no-load tracking error and steady-state torque fraction; blame authority only if the task fails while the error grows with demanded force - and then fix gains or targets, not training.
Symptom
hip_roll sagged 0.21 rad on the ground and saturation questions loomed over the sidewalk plan - was the roll axis physically too weak (authority), or was something else limiting it?
Context
Before C4, the roll-authority question was settled by measurement triage: suspended test (--suspend, feet off ground) showed hip_roll tracking error 0.0008 rad - actuator acquitted; the entire 0.21 rad ground sag is load-induced. Steady-state torque was 25% of limit - 75% margin remains. Since sidewalk needs lateral force, not exact angles, authority was ruled "not a hard limit", with a pre-registered criterion for when it WOULD become one: sidewalk fails to track AND roll error keeps growing - then the fix is raising hip_roll kp or lowering the vy target, not more training.
Change
Hypothesis "roll authority insufficient" demoted from blocker to a monitored branch with an explicit trigger condition; C4 proceeded.
Outcome
Later open-loop probes confirmed the actuator could produce the behavior (sidewalk feed-forward ran at full amplitude, 5/5 survival), and the eventual C4 failure causes were measurement and reward, never authority.
Mechanism
Suspended vs loaded comparison separates the actuator's closed-loop competence from the load path: tiny no-load tracking error means the motor/controller is fine and any loaded deviation is statics (gravity / stiffness budget, kp trading error for force). Torque-fraction measurement then bounds how much force headroom actually remains.
Applies when
- suspecting an axis is "too weak" for a new skill
- large position sag on a loaded joint
- deciding between hardware fix, gain change, and more training
“吊挂(--suspend)实测 hip_roll 跟踪误差 0.0008 rad → 执行器无罪,地面下垂 0.21 rad 全是负载所致;稳态占限扭 25% → 仍有 75% 扭矩余量。… 判据:若 C4 出现「侧走跟不动且 roll 误差继续变大」,那才是权限账 … 解法是提 hip_roll 的 kp 或降 vy 目标,不是硬训。”
train/C_LADDER_RUN.md § 3d. roll 权限:已部分澄清,不是硬上限 Foot dragging is an attractor, not a low amplitude - and joint damping is the mode switch, adjustable at deploy time
swing-bistability-damping-switchWhen a quality metric is bimodal, stop treating it as an amplitude to be trained up: map the modes against initial conditions and plant parameters, find the parameter that switches basins, apply it first as a deployment lever, and only then bake it into the training distribution (as a plant-family shift, never as an execution-mapping change).
Symptom
s2e_pd-1400's swing height "median 12.1 mm" hid a perfect bimodal distribution: 20 seeds split into a drag mode (2.6-4.9 mm) and a step mode (19.3-24.0 mm) with NOT ONE seed in between - the median sat in the empty gap, and "swing debt -11 mm" really meant "50% probability of falling into the drag attractor".
Context
Two designed experiments closed the mechanism. Test A (nominal plant, 40 seeds): step 42% / drag 58% / middle 0 - at nominal gains, initial conditions alone pick the mode, both modes 100% survivable. Test B (fixed init, kp x kd grid): kd is the mode SWITCH - at kd 1.3 all surviving cells step (13-22 mm), at kd 0.7 nearly all drag (2.7-4.3), only at kd 1.0 does init get a vote; kp >= 1.2 is dangerous (5/6 falls). Global verification at kd x1.3 (20-seed, delay 2): survival 20/20 at ZERO cost, step share 42 -> 80%, swing median 12.1 -> 18.4 mm, slip record low 334, thicker tilt margin - costs: vx 85 -> 78%, saturation +5 pp. A Pareto sweep then priced the knob: step share 42/72/75/88/82/90 across kd 1.00-1.30 with a linear vx tax of -2.3 pp per 0.1 kd - the basin gain is fully collected at kd 1.20 ("1.30 是 over-damping 纯多付税"). Mechanism: low damping leaves a landing micro-oscillation / ground-slide channel the policy can exploit to drag; damping plugs the channel.
Change
Deployment lever adopted: kd-scale 1.20 (conservative 1.15) as the legitimate successor to the power-0.8 crutch ("前者削幅度保稳,后者堵 拖地通道换步态,且不牺牲存活"); training-side prescription: move the DR band to nominal-1.2 x (0.9,1.1) = [1.08,1.32], deleting the [0.7,1.0) drag-teaching zone - a contract-level change requiring digest re-baselining, gated on measuring the real robot's actual kd dispersion first.
Outcome
The kd surgery rung (s2e_kd) delivered basin 8 -> 11/20, slip 405 -> 331, vx 81 -> 85% with no out-of-band fragility (below-band check 20/20) - "拐杖烧进分布的正确姿势", explicitly contrasted with the failed s1g amplitude version: this one changes the plant family the policy has seen, that one changed the execution mapping the policy would have to relearn.
Mechanism
The gait's swing behavior is a bistable dynamical system whose basin boundaries are set by plant parameters; a policy trained across a kd band that includes the drag basin has learned to inhabit it. Shifting the deployed (and then trained) damping moves the system into the step basin without touching the policy - a plant-side fix for what looked like a training deficiency.
Applies when
- a gait quality metric splits into distinct modes across seeds
- deciding between more training and a gain/damping change
- converting a deployment crutch into a training-distribution change
“20-seed 里拖地模式 2.6~4.9mm 与迈步模式 19.3~24.0mm 各半,中间一个不落 … kd 是模式开关——kd1.3 下 6/6 存活格全迈步 … kd0.7 下几乎全拖地 … swing 债的解(至少大半)在部署端阻尼档,不在训练端 … 机理:低阻尼下落脚微振荡/贴地滑给了策略顺势拖行的通道,加阻尼堵之。”
train/README.md § swing 双稳态定性 + kd 部署杠杆 (2026-08-07, 用户设计 Test A/B) Symmetrizing the config made the gait MORE asymmetric - the asymmetry lived in the policy weights
asymmetry-in-weights-not-configLocalize a persistent asymmetry by intervening at the config layer first: if the symptom survives (or worsens), it is in the weights - fix it with symmetry-constrained training, not with trims or offsets.
Symptom
walk_v1 on hardware: straight-line command curved 149 deg in 15 s (9.9 deg/s) with 3.06 m lateral runout; turn gain +31% one way vs +129% the other (75% difference); knee asymmetry 4.4 deg in sim, 9.6 deg on the robot.
Context
The obvious suspect was the asymmetric default pose in the config. The decisive test: symmetrize standing_pose and run the SAME policy in sim - the asymmetry got LARGER (hip_pitch 6.8 -> 9.3 deg). Root cause therefore not in the config but baked into the policy weights: PPO without a symmetry constraint commonly converges one-sided, because splitting the work 50/50 and loading one side yield the same return, and the gradient falls randomly into one of the equivalent optima.
Change
Fix redirected from config trimming to retraining with mirror data augmentation (walk_v2 spec) - a weights-level fix for a weights-level disease.
Outcome
With augmentation (and the symmetric-default precondition), stand_v1 reached 0.0 deg asymmetry on all six joint pairs (from 4.4-7.7 deg), height fluctuation 7 mm -> 1 mm, mean |action| down 33%.
Mechanism
Reward-equivalent solution families (who carries the load) leave the symmetric solution unpreferred; SGD picks an arbitrary member and entrenches it. Config changes move the coordinate frame around the entrenched asymmetric function - they cannot move the function. The counterfactual test (change config, watch symptom) localizes the layer the disease lives in.
Applies when
- a robot veers or loads one side despite a symmetric-looking config
- deciding between config trims and retraining for an asymmetry
- mirrored-turn gains differ by tens of percent
“根因不在配置里:把 standing_pose 对称化后在 sim 里跑同一策略,不对称反而变大(hip_pitch 6.8°→9.3°)—— 说明不对称烙在策略权重里。这是无对称约束的 PPO 的常见收敛结果(左右各担一半与一边多担的回报相同,梯度会随机落进其中一个)。”
train/RETRAIN_v2.md § 1. 为什么是对称增强(证据) Raising a command bucket's share does not strengthen its per-state gradient - it only starves the other modes
bucket-share-is-not-a-gradient-leverWhen a skill is not learning, first prove its per-state signal is nonzero (ignore-floor and probe checks); only rebalance sampling shares to fix genuine sample starvation, and account the regression risk to the diluted modes before doing it.
Symptom
Sidewalk was not learning, and the reflex proposal was to give the side bucket a larger share of sampled commands.
Context
The C4-redo3 rung explicitly kept the 20/40/20/20 bucket (stand/forward/turn/side) with the reasoning written out: PPO computes advantages per state, so bucket proportion does not change the per-state gradient of side states; at 4096 envs x 20% x 24 steps the rollout already contained ~19.7k sidewalk states - sample count was not the bottleneck. And the cost side was already measured: cutting forward from 60% to 40% had made vx+0.30 die at +400 in an earlier run - more cuts would only collapse it sooner.
Change
Bucket proportions held constant across the entire C4 redo series; the actual bottlenecks (metric frame bug, reward variance penalty, exploration form) were pursued instead.
Outcome
Sidewalk was eventually fixed with zero bucket changes (feed-forward delivery, +100 iters); forward/turn skills never suffered starvation-induced regressions during the redo series.
Mechanism
Policy-gradient credit is assigned per visited state; oversampling a mode multiplies its states in the batch but not the informativeness of each, so if the per-state gradient is ~0 (behavior unreachable or reward indifferent), N times zero is still zero - while the displaced modes genuinely lose data and regress.
Applies when
- proposing to oversample a failing task/command mode
- a majority mode regresses after share rebalancing
- budgeting env count vs mode share for a multi-skill policy
“比例不动:PPO 逐状态算优势,桶占比不改变单状态梯度;4096 env × 20% × 24 = 每 rollout 已有 1.97 万个侧走状态,样本数不是瓶颈;而 forward 60%→40% 已实测让 f30 在 +400 处死掉,再加码只会更早塌。”
train/C_LADDER_RUN.md § 3i. 桶 20/40/20/20 不动(比例不动) Sort external training advice into adopt / already-have / modify / would-trap by recomputing it on your own config
external-advice-audit-against-own-arithmeticNever apply external tuning advice directly: recompute each claim on your own reward table and probe data, classify it adopt / have / modify / trap, and record why - and verify external citations actually exist.
Symptom
External AI/literature advice for the omni ladder arrived plausible-sounding but was written without knowledge of this robot's actual reward table, contract, and history; following it blindly would have broken single-variable discipline and, in one case, made sidewalk unlearnable.
Context
Before the C ladder, every external suggestion was audited: 3 adopted (ellipsoid command sampling; staged wz bands; command-switch acceptance), 3 already present (unified reward; frame history - the frozen 168-dim 5-frame window; per-100-iter acceptance), 2 modified (stand share kept at 20% to avoid a second variable; back share NOT raised because the probe showed backward works untrained 20/20@67%, so oversampling would only crowd out forward), and 1 flagged as a trap: "start vy very small (0.06-0.15)" - on THIS reward table vy was only an L2 tax, so ignoring a vy=0.06 command costs 0.4% of the vx tracking scale, 28-180x cheaper than ignoring forward, with quadratic shrinkage making small commands weaker still. A separate retrieval-reliability note: two search agents returned fabricated verbatim quotes from arXiv PDFs (2 papers, verified fake and discarded); only HTML/abstract/source-verifiable material was used.
Change
Advice classified only after recomputing each claim with local numbers; the "start small" trap was replaced by adding a gated lateral tracking term (the ladder's only true reward surgery) instead of shrinking the command.
Outcome
The adopted items (ellipsoid modes, staged wz, transition acceptance) entered the ladder; the trap was avoided; one external factual error (calling s1g the mainline start - it was falsified 0/20) was caught. Later, one initially-dismissed item (sigma=0.15 too narrow) turned out right for a different reason than claimed - see cycle-average-tracking-for-gait-quantities.
Mechanism
External advice encodes the advisor's reward table and robot, not yours; the transfer-validity test is whether the claim survives recomputation under your own arithmetic (reward margins, probe baselines, contract freeze). Items that survive become experiments; items that don't become documented traps.
Applies when
- incorporating LLM or literature advice into a training plan
- advice conflicts with locally measured baselines
- an external claim depends on reward-table details the advisor cannot know
“其建议 C4「先很小,vy = ±0.06~0.15」—— 在我们这张奖励表下会让侧走学不起来 … 忽略侧走比忽略前进便宜 28~180 倍,且指令越小激励越弱(平方缩放)—— "先很小"在稳定性上对、在梯度上正好把信号缩没了。”
train/C_LADDER_RUN.md § 1. 外部 AI 训练建议的评估(采纳 / 已有 / 要改 / 会踩坑) Prove an armless get-up exists before training it - a connected static domain, 25% torque on the cheapest path, an 8 mm hand-over gap, a static roll-over - and write down what each scan cannot represent
get-up-feasibility-accounts-before-trainingBefore training a get-up or any multi-contact skill, compute the quasi-static accounts - connectivity of the static domain, torque along the cheapest path, hand-over gaps, COM shift available for rolling - and state which configurations each scan cannot represent; when a policy gets stuck in one of those, extend the scan before blaming the reward.
Symptom
A torso-and-legs robot has no arms to push off the ground; whether it can get up from the floor at all was unknown when the line opened.
Context
recovery_feasibility.py ran three accounts before any training (the run line's "hard accounts first" discipline): a sagittal quasi-static scan (0.05 rad grid, 44,520 configurations, MuJoCo FK, flat-foot assumption). (1) The static standing domain (COM over the feet, torques in limit) has 29,586 cells, flood-fill connected with no islands, from a 0.097 m deepest squat to the 0.384 m stand. (2) The minimum-torque path peaks at 25% of the limits (knee 2.9/12, ankle 3.7/17 N*m) - a 4x margin. (3) All 508 ground-contact configurations have the contact behind the COM; the smallest gap to pure foot support is 8 mm. A roll-over account: swinging both straight legs to one side shifts the COM 96 mm against a 62 mm torso half-width - 1.6x, so rolling needs no momentum. Three conclusions were written down for later attribution: the legs are 80% of the mass (swinging them moves the COM), prone has no flat-foot hand-over face (merge into a supine/side sit first), and supine needs no sit-up (hip flexion is limited to 75 deg).
Change
The accounts gated opening the line and were cited in every later argument about what the robot can physically do.
Outcome
They held where they applied: in V1.0 every fall category was righted under a hard rate limit, which the spec records as the quasi-static roll-over account verified by training, and the 25% torque path was the basis for pursuing a slow get-up. They also misled once: account (3) is sagittal, and on 08-09 the spec corrected its scope - it cannot represent the splayed W-sit where the policy actually stalled. A follow-up prone hip-ROM scan (471,625 cells) found 3,912 two-foot-contact cells and none with both soles within 25 deg of level (best 40.2 deg): a flat-footed push-up from prone is infeasible on this robot, so the fix became where the feet go after sitting up.
Mechanism
A get-up needs a connected path through statically feasible configurations and enough torque along it; quasi-static accounts bound both cheaply, and momentum can only make the real problem easier. A reduced-dimensional scan, though, only speaks for the configurations it can express.
Conflicts
In R0.1-R0.2 the spec read account (3)'s "prone has no front hand-over" as "prone lacks the roll-over skill"; R0.3's confusion matrix showed prone had righted its torso 159/159, and the spec then restricted account (3) to the sagittal configurations it models.
Applies when
- opening a get-up, recovery or climbing skill on a new robot
- a robot lacks arms or other obvious contact options
- a policy stalls in a configuration a feasibility scan never modelled
“本机 **torso + legs、无手臂可撑地**,开训前先证明存在不依赖手臂的物理解 … 双腿同侧直腿摆最大横移 **96 mm = 1.6×** … **静态摆腿即可翻身,无需动量**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §1 机械可行性判决(2026-08-09,recovery_feasibility.py 三笔账) Train with self-collisions ON (filtering nested-link ghost pairs) - the reward wall prevents, the physics makes cheating impossible
self-collision-physics-plus-reward-wallNever train a contact-risk behavior with self-collisions disabled; enable them with an audited filter list for nested/overlapping pairs (zero contacts across a pose sweep), record the fps cost, and keep a calibrated distance penalty as the preventive layer on top.
Symptom
walk_v8 logged 107 frames of leg-on-leg contact while still earning 0.751 tracking score - because training-side self-collisions were OFF, leg clipping was literally imperceptible to the policy ("碰腿在训练里 根本感知不到").
Context
Enabling self-collisions naively is its own trap: an Isaac audit had shown PhysX auto-filters adjacent bodies (base-hip clean for free) but nested links generate ghost forces - calf and ankle_roll overlap 65 mm at the zero pose, producing 12x body-weight phantom forces. The v10 recipe: enable self-collisions, explicitly filter only the two nested pairs (l/r calf-ankle_roll), then run a zero-contact audit at three poses (nominal stand, walk crouch, swing-extreme) requiring contact count = 0, adding any residual pair to the filter and re-auditing; a 500-iter sanity run for NaN and an fps-cost record (measured -8.8%). Redundancy with the reward-side foot-distance wall was argued, not assumed: "N2 离得远(奖励侧预防),SC 碰了疼(物理侧兜底)" - the reward keeps distance at range, the physics makes contact hurt - so the v8-style "clip legs and still score" outcome becomes physically impossible.
Change
enabled_self_collisions=True + 2-pair filter + three-pose zero-contact audit (re-verified at 0.00 N after the later mass update) + fps budget recorded.
Outcome
Leg contact entered the training signal; the audit protocol caught the nested-pair ghost-force hazard before it corrupted training; combined with the calibrated distance wall, later versions held contact = 0 on hardware and in sim.
Mechanism
A hazard absent from the training physics cannot be learned about, no matter the reward; but collision meshes that interpenetrate at rest inject large fictitious forces if enabled blindly. Filtered enabling plus a pose-swept zero-contact audit gives true contact physics with no phantom energy - and layering prevention (reward) with consequence (physics) covers both learning and enforcement.
Applies when
- real robot self-contacts while training scored it healthy
- enabling self-collisions on a model with nested collision meshes
- deciding between reward-side and physics-side fixes for clipping
“PhysX 自动过滤相邻体(base↔hip_pitch 免费干净),幽灵力只在 calf↔ankle_roll(零位嵌套 65mm,12 倍体重)。… 与 N2 互补不冗余:N2 离得远(奖励侧预防),SC 碰了疼(物理侧兜底)—— v8 那种 107 帧互碰拿 0.751 跟踪分的事从此物理上不可能。”
train/WALK_V10_SPEC.md § 4. SC —— 训练侧自碰撞(范围已探明,比想象便宜) Four consecutive fixes were each continued from the previous fix's degraded state until the user stopped the ladder - "change parameters, don't stack errors" - rolled back to the last good checkpoint and audited the target geometry first
stop-stacking-roll-back-and-auditWhen successive rungs each start from the previous rung's output and the target symptom does not move, stop, roll back to the last good checkpoint and re-derive the next change from an audit; keep the measurements, discard the stacked remedies.
Symptom
After the real-robot splits, the V2.7 ladder tried to widen the stance: a term swap (A), a new stance knife (b), more iterations (甲), a doubled weight (乙). Stance barely moved while hip yaw ratcheted 46.7 -> 49.9 -> 52.5 deg toward its 60 deg limit.
Context
Each rung started from the previous rung's output. The user ruled on 2026-08-11 that things had gone wrong from V2.7-A: go back to v2_6 and rethink which parameters to change instead of stacking errors.
Change
乙 was killed at start and not counted; the product baseline rolled back to v2_6c model_29399; every measurement and law learned on the ladder was kept ("the data is real; what stacked was the treatment"). Before any new training, a zero-training kinematic audit of the stance targets was run.
Outcome
The audit found the stand_pose target itself rewarding the narrow stance (pose-target-geometric-audit) and showed geometrically why the yawed stance could not be widened with flat feet - so 乙 was proven unnecessary without running it. The next in-lineage attempts still failed, which is what established that the stance is set by the get-up path.
Mechanism
A rung continued from a degraded state inherits its compensations, so each new fix answers the previous fix's side effects; the yaw ratchet was the visible trace of that stacking.
Applies when
- three or more corrective rungs in a row without progress on the target metric
- a side-effect metric ratchets in one direction across rungs
- a new rung is being planned from the latest (not the best) checkpoint
“用户裁:"从 V2.7-A 开始就出问题了,应该回到 2.6 再思考如何改变参数而不是 错误叠加。"认账:A 的补丁 → b 的新刀 → 甲的加时 → 乙的加权,每级都从上级 的**退化状态**续(yaw 46.7→52.5° 的棘轮就是叠加痕迹)。 … 数据是真的,叠加的是处置。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 方法论裁定(2026-08-11,用户):V2.7 全阶梯叫停,回滚 v2_6 The advisor's "runtime five-stage state machine + per-stage reference poses + RL residual" appeared in none of the three papers it cited - reading the originals changed the plan and downgraded two widely repeated industry claims
advisor-paraphrase-vs-paperRead the primary source behind any piece of advice before adopting its architecture; record where the paraphrase and the original differ, downgrade claims the originals do not support to speculation, and adopt what the verified sources actually share.
Symptom
After the violent first real run, an advisor proposed re-architecting recovery as a runtime staged state machine with reference poses and an RL residual, citing HoST, HumanUP and StableMimic.
Context
The three papers were read in full on 2026-08-09 and tabulated (deployment form, what the "stages" really are, hard constraints, references). HoST: one end-to-end policy, height-gated rewards in training, action anchored as q + beta*a with a beta curriculum. HumanUP: two training stages with the same observation/action; the vendor's three-stage state machine is the baseline it beats (41.7% vs 78.3%); Stage II tracks an 8x slowed Stage I trajectory (4x too violent, 10x does not converge). StableMimic: a learned soft gate. On 08-10 more sources were checked the same way: the Agility page does not say Digit's self-righting was learned in simulation (only step recovery is stated as RL), so the claim was downgraded to speculation; HoST's support for the 12-DoF armless Mini Pi exists in its code repository, not in the paper text.
Change
Adopted only what the originals share: a hard action bound or anchored action space, strong smoothing including a second-difference term, a slowed hidden reference, heavy DR and real fallen states. The staged runtime state machine was not adopted.
Outcome
The next re-rooting candidates came straight from the verified material, and the beta-anchored action space (HoST, with the Mini Pi configuration as the nearest real-robot precedent) became V2, the lineage that later stood up on hardware.
Mechanism
A paraphrase compresses a paper into the advisor's own architecture; only the original shows what was actually deployed, what was a baseline, and which numbers came with which ablation.
Applies when
- an advisor, agent or summary proposes an architecture with citations
- an industry claim ("X learned it in sim") is about to justify a design
- several papers are cited for one combined recipe
“⇒ **顾问的核心形态"runtime 五阶段状态机 + 每阶段参考姿态 + RL residual"在三篇引文 里均不存在**,其中 HumanUP 还点名 state machine 是局限。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 三篇引文精读判决(2026-08-09 全文核对;顾问转述与原文有出入) Anchoring the action on the measured joint angle (target = q + beta*a), with a beta curriculum down to tau_limit/kp, bounded torque by construction, removed the re-falls and later stood the robot up on hardware
beta-anchored-action-targetFor large-motion skills on position-controlled actuators, bound the action relative to the measured joint angle with a per-joint authority of tau_limit/kp, curriculum the authority down from full range, keep the curriculum state out of the observation and pin acceptance at the deployed authority - and make the deployment code refuse to run the anchored contract without a measured q.
Symptom
The V0 full-range absolute action produced violent targets; the V1 command-anchored rate limit made standing oscillate. Both failure modes came from how the action becomes a target.
Context
V2.0 (user approved, from scratch): BetaAnchorJointPositionAction, target = q_measured + beta_j(m)*a, memoryless per step. beta_j(m) = floor + m*(beta0 - floor), beta0 = the contract half-range (m = 1 reproduces V0 authority), floor = min(tau_limit/kp, beta0): hip_pitch 1.309 -> 0.40, knee 1.047 -> 0.40, hip_yaw -> 0.917, the other joints unchanged - the tightening lands exactly on the joints the V0 torque account convicted. m drops 0.1 per step when a standing-share EMA exceeds 0.35. beta is NOT in the observation, so the 45-dim contract is untouched; acceptance is pinned at m = 0 because the Python curriculum state is not saved in the checkpoint. The deployment chain got a new profile (recovery_v2: action_anchor current_q, explicit per-joint beta written into the contract, independent of the gain profile), and policy_io raises if q is missing rather than silently falling back to the absolute contract; the old profile's check reproduced its pre-change deviation bit for bit.
Change
New action term and beta curriculum; later the RS06 floor was lowered 0.40 -> 0.30 -> 0.25 (kp*beta 7.5 N*m) and the stamped deployment profile was synced to 0.25.
Outcome
First acceptance at m = 0 (v2_0b): re-falls 0% in every category, the torque gate passed for the first time on the line (worst 69.9%), knee jitter 0.004; supine 98.8 / side 88.8% with prone and mid still failing (fixed by the conditional pull curriculum). MuJoCo showed demand at or under the limits (hip_pitch 11.7/12 against V0's 26.8). Lowering beta cut impact (hip_pitch demand 9.7 -> 8.5 N*m) but barely slowed the get-up - it had become coordination-limited. Enabling the policy moves the target only +/-beta around the current pose, so there is no homing fling; the 08-11 real get-up and the later v3_1p1c both run on this contract.
Mechanism
kp*beta caps the proportional torque in a single step with no build-up delay and no memory, giving both a hard impact bound and full balance bandwidth.
Applies when
- a skill needs full joint range but hardware torque limits are low
- absolute position targets cause impacts or saturation
- changing the action semantics of a contract that deployed policies share
“**动作项** `BetaAnchorJointPositionAction`:`target = q_实测 + β_j(m)·a`, 逐步无记忆 … **Play/验收钉 m=0(= floor = 部署档)**:python 课程状态不进 checkpoint, Play cfg 显式 `beta_m_start=0` … 判读:**结构赌注兑现** —— 站姿零再摔 + 力矩账首过(kp·β 封顶按构造)”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §33 V2.0 预注册(2026-08-10,用户点头开工):β 锚定动作空间,从零训 After seven patch-generations, freeze the lineage as a regression baseline, fix the structural debts, and retrain from zero
freeze-lineage-fix-structure-restartWhen successive rungs keep trading one symptom for another, ask whether the remaining problems are structural (contracts, latency, sampling, reward-table architecture); if so, freeze the lineage as regression baselines, pay the structural debts, and restart minimal - carrying forward laws and instruments, not weights and weights' patches.
Symptom
The v5-v11 walk lineage had accumulated interacting patches (reward terms, gates, clamps, per-joint scales) faster than it converged on the user's goal; v12's spec itself was superseded before training by an external review's verdict that the remaining problems were structural, not parametric.
Context
The 2026-08-05 status banner records the pivot: the walk profile was rolled back wholesale to v10b parameters, the v5-v11 lineage frozen "只作回归对照" (kept only as regression baselines), and four structural debts were named as prerequisites for a from-zero straight-walk baseline: the action-latency FIFO (fixed with its own test), the ONNX manifest contract, discrete command sampling, and a minimal reward table. The v12 spec - fully designed, partially implemented - was suspended: "本规格挂起,不再按此开训".
Change
Strategy switched from "one more patch generation" to freeze-fix-restart: lineage checkpoints retained as comparison anchors, infrastructure hardened first, then a clean retrain with a minimal reward table (this restart produced the s* generation that later became the real-robot SOTA line).
Outcome
A designed-and-ready training generation was deliberately not run - the review's structural findings outranked sunk design cost; the restart line inherited seven generations of laws (calibrations, gate batteries, falsified fixes) without inheriting their entangled reward table.
Mechanism
Patch lineages accumulate coupled terms whose interactions eventually cost more to reason about than a restart costs to train; the knowledge worth keeping is the laws and instruments (measured plant values, calibrated gates, falsified directions), not the entangled weights. A restart on hardened structure converts the lineage's lessons into a clean initial design instead of another delta.
Applies when
- repeated rungs shuffle symptoms without net progress
- an external review flags infrastructure/contract debts
- deciding between another patch generation and a clean retrain
“同日外部评审定调换路线:冻结 v5~v11 血统(只作回归对照),修结构性问题(latency FIFO 已修 tests/test_action_latency.py、ONNX manifest 契约、离散命令采样、最小奖励表)后从零训直行基线。本规格挂起,不再按此开训。”
train/WALK_V12_SPEC.md § ⚠️ 状态 (2026-08-05) Friction DR was demoted after a measurement (94% success at mu 0.4 with no friction randomization) and promoted again when the action contract changed and mu 0.4 fell to 76% - DR priorities belong to a plant and contract, not to a task
friction-priority-re-measured-after-plant-changeRe-measure transfer along the friction axis for every new action contract or plant, not once per task; a DR priority settled under one action parameterization does not carry to the next.
Symptom
Getting up is all scraping and pushing against the ground, and training pinned friction at 1.0, so friction looked like the first thing to randomize.
Context
The MuJoCo gate on R0.5 (5 categories x 10 seeds x 4 friction levels) measured 100/100/98/94% at mu 1.0/0.8/0.6/0.4: degradation showed first as time (prone 3.2 -> 5.3 s), not failure, so friction DR was demoted and the DR budget earmarked for mass/COM. After the switch to the beta-anchored action space, V2.2 read 90/94/90/76%: mu 0.4 was now the weak row.
Change
V2.3 (single variable): friction DR static (1.0, 1.0) -> (0.2, 2.0), dynamic (0.15, 1.6), the HiFAR range keeping the base dynamic/static ratio; restitution untouched. Continued from v2_2.
Outcome
Isaac nominal 99.8% (DR did not hurt the nominal plant); MuJoCo 98/98/96/92% - mu 0.4 76 -> 92%, mu 1.0 back to R3.1's 98% with bounded torque.
Mechanism
How much a policy leans on friction depends on how it moves; the spec records that the sensitivity rose after the action contract changed but does not establish why.
Applies when
- changing the action space, gains or authority of an existing skill
- deciding which DR axis to spend the next rung on
- an earlier sweep justified leaving an axis unrandomized
“**μ 砍到 0.4(训练值的 40%)仍有 94%**,退化先体现在**用时**(prone 3.2→5.3 s) 而不是成败。μ≥0.8 完全无损。→ **§17 曾把"摩擦随机化提到 R4 第一项"当作优先 事项,这条实测把它降级了**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §21 MuJoCo 复核门 ② 摩擦依赖