Training Coach
Doctrine
A report may cite any of these as doctrine-N.
doctrine-1Contract freeze and fingerprint disciplineThe policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.
Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).
Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.
doctrine-2Attribution by resolved training params - never eval-override knobsCapability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.
Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).
Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.
doctrine-3PASS gates become constraints; FAIL gates become objectivesOnce a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.
Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.
Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".
doctrine-4One variable per ladder rung - counted against what the checkpoint sawA rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.
Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.
Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.
doctrine-5Pre-register risks, readings, and stop criteria before the ladderBefore a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.
Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).
Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.
doctrine-6Plant parameters are measured, never inventedEvery plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.
Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).
Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.
doctrine-7Sim2sim gate before sim2real - under deployment conditionsEvery checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).
Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).
Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.
doctrine-8Observation honesty - the actor's inputs are a hardware contractThe actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.
Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).
Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.
doctrine-9Reward economics are audited in realized currencyReward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.
Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).
Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.
doctrine-10The zero-cost option must be the desired behaviorFor every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.
Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).
Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.
doctrine-11Measurement discipline: independent referees, signs, distributionsA disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.
Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).
Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.
doctrine-12The deployment pipeline is plantIrreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.
Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).
Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.
doctrine-13DR budget is finite; its distribution is the measured supportRobustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.
Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).
Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.
doctrine-14Gates measure what hardware feels: posture, margins, stripped assistsAcceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.
Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).
Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.
doctrine-15Fork and root selection: recoverability, maturity, frozen rewardsChoose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.
Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).
Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.
doctrine-16Curricula: verified engagement, lineage counters, disease-phase gatingAutomatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.
Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).
Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.
doctrine-17Probe before training: feasibility first, hypotheses in tablesAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.
Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).
Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).
doctrine-18External advice is recomputed locally; values transfer as ratiosEvery external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.
Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).
Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.
doctrine-19Hardware sessions are scripted experiments, not tuning sessionsReal-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.
Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).
Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.
doctrine-20Close questions in writing; restart when the debt is structuralAudited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.
Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).
Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.
doctrine-21Name the quantity in the space it lives inA goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.
Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).
Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.
doctrine-22Continuation needs a live gradient; a release is chosen by a scanContinue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.
Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).
Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.
Experience cards
168 cards matching “gate-penalties-to-the-disease-phase”.
Curriculum-gate a penalty to the phase where its disease occurs - early on it only taxes exploration
gate-penalties-to-the-disease-phaseFor penalties aimed at late-stage pathologies (freezing, saturation, degenerate attractors), ramp the weight in only after exploration noise has decayed; anchor the terminal weight to measured healthy-vs-sick raw values, and shift all related tripwires to after the ramp completes.
Symptom
The action_saturation penalty, applied from iteration 0 in v8a, taxed exploration itself: with init_noise_std 1.2 the sampled actions paid ~-2.45/step before any policy had formed - while the disease it targets (clamp freezing) is a LATE pathology (v9 froze at iteration ~2624).
Context
v10 re-introduced the same penalty behind a curriculum gate: weight 0 until iter 1000, ramping linearly to -1.0 by iter 2000 - present only when the disease can occur, absent while exploration noise dominates. The trust argument was evidence, not hope: in v8a the term, while active, had pulled joint_pos_ref from 0.041 up to 0.155 and climbing - proof it can extract a policy from the frozen pit. Weight magnitudes were anchored to measured raw values (healthy v5 0.310 / v6 0.106 vs frozen v7 1.145 / v9 1.22 per step: at -1.0 healthy pays 6-18% of tracking, frozen pays 60%+, standing ~0). v10c then isolated the gated term as THE anti-freeze mechanism by single variable, upgraded to untouchable status in v11: "S 的门控机制(v10c 单变量铁案:任何情况下 不许撤,只许调终值)" - and v11 dared to relax other penalties only because S stood guard.
Change
action_saturation gated 0 -> -1.0 over iters 1000-2000 (later terminal value tuned -1.0 -> -0.5 with the gate mechanism itself frozen); tripwires adjusted to respect the gate's timing (freeze check moved to iter 2500-3000 to give the ramped term its effect window).
Outcome
Freezing stopped recurring while early training kept full exploration; the mechanism graduated from experiment to invariant within two versions.
Mechanism
A penalty's incidence depends on who occupies its support: early in training that is exploration noise (whose suppression starves learning), late it is the converged pathology. Time-gating aligns the penalty's presence with its target's presence, buying the constraint without the exploration tax - and tripwire timing must then be computed from the gate schedule, not from ungated precedents.
Applies when
- a structural penalty punishes exploration in early training
- a late-onset pathology (freeze/saturation) needs a standing guard
- deciding when a curriculum ramp should engage
“v8a 实锤它的病根是"罚在采样动作上"——init_noise_std 1.2 的早期等于罚探索(~−2.45/步);而冻结是晚期病(v9 速率 2624 才死平)… 门控让它只在病发期在场。… v8a 里它在场时 joint_pos_ref 从 0.041 爬到 0.155 且仍在升——有从低谷爬出的实证。”
train/WALK_V10_SPEC.md § 2. S 保险 —— action_saturation 课程门控 A time gate that had cured one lineage's rushing made the from-scratch lineage trade away its stance width twice (0.364 -> 0.235 m, 0.355 -> 0.251 m) - its disease was absent there, so the fix was retired and the pre-gate checkpoint shipped
time-gate-vs-wide-stance-retire-the-fixCarry a fix into a new lineage only if its disease is present there; a mechanism that cured one lineage can be net negative in another, and when doubling a term's weight recovers almost nothing, treat the two objectives as structurally in conflict and remove the one whose purpose is gone.
Symptom
V3.1's phase 2 (the 3 s zero gate on standing income, continued from P1b) kept 100% success on every friction level and slowed the get-up, but the lateral stance drifted 0.364 -> 0.235 m and hip yaw crept to 57 deg against its 60 deg limit. With the width band's weight doubled (P2c, after P1c) it drifted again, 0.355 -> 0.251 m, below the pre-registered 0.30 m failure line.
Context
The zero gate had been introduced in V2.5/V2.5b to slow the old lineage's get-up. In V3.1 the rushing was already absent: P1c got up in 0.90-1.06 s with a worst torque ratio of 73.3%, better than the stamped v2_6c, because the full beta curriculum, second-difference smoothing and pull curriculum had cured the violence inside training.
Change
Recorded as a candidate law with two data points - the zero gate and a wide stance are mutually exclusive here - and the zero gate was removed from the V3.1 recipe. P1c (the pre-gate checkpoint) went through the full stamp-level acceptance instead.
Outcome
P1c passed everything: all six criteria, lateral stance 0.355 m, foot tilt P75 2.0 deg, mu {1.0, 0.8, 0.6, 0.4} x 10 seeds all 100%. recovery_v3_1p1c.onnx was stamped and pushed to the robot channel.
Mechanism
The zero gate moves the income toward "stay stable until the end", and under low-friction DR a wide stance has a slip tail, so survival outbids the width band; doubling the band's price bought back only 0.016 m - an auction that does not converge signals structural conflict, not an under-priced term.
Applies when
- porting reward mechanisms from an old lineage into a fresh recipe
- a width, margin or posture metric erodes during a late training phase
- a weight increase produces a negligible change in its target
“**定律候选(二实证):归零门 × 宽站互斥**。 … 加价翻倍只挽回 0.016,竞拍不收敛)。 … **归零门是 v2_5 血统的历史包袱,对 V3.1 配方是净负资产,P2 阶段除名 —— P1c 即终点形态**。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §49 终验(2026-08-14) Every gate on a penalty is an exit - to stop paying a stance tax the policy parked 2 deg outside a 30 deg uprightness gate (lunging), and, from scratch, just under a height gate (crouching); a positive band and an always-on guard fixed both
penalty-gate-is-an-escape-hatchNever gate a penalty on a state the policy can leave by getting worse; use always-on guards for what must never happen and positive, gated bands for what you want, and count every gate on a penalty as one more escape route to check in the logs.
Symptom
V2.9 added stance_width_task = relu(0.34 m - foot spacing) x standing gates, weight -10. Three checkpoints scored 0% on acceptance: standing height reached, feet on the ground, angular rate low, but the torso leaned 32.5/32.3/31.9 deg in a fore-aft lunge. The training dashboard read "tax paid off, base_height at full value".
Context
The penalty was gated by uprightness (tilt < 30 deg) and standing height; its tax had no time gate (500 steps x -1.65) while the standing income sat behind a 3 s zero gate (~325 steps), so leaning just past 30 deg lost a little gated income and saved the whole tax. base_height has no upright gate, so the lunge still collected it. The first metric also measured full horizontal spacing, so a staggered lunge counted as "wide".
Change
Two laws written down: a penalty may only carry gates the policy cannot escape by getting worse (make it an always-on guard) or it becomes a positive band ("not earned" is not "escaped"); and width is measured laterally in the base yaw frame. From scratch (V3.1 P1) with the lateral metric but the same gates, the policy parked just under the height gate instead (base_height 1.176/1.5, h ~ 0.30 m against a 0.3264 gate), the beta curriculum never advanced in 1,700 iterations, and the run was stopped early. P1b flipped the penalty into a positive band +2.0 x clamp(lateral / 0.34) x standing gates; P1c added yaw_guard = -5 x relu(|hip_yaw| - 30 deg), always on, no gate, no exemption.
Outcome
P1b: lateral stance 0.364 m, all four categories 100%, MuJoCo mu 1.0/0.4 both 100%. P1c: all six acceptance criteria passed for the first time on the line (hip-yaw saturation 1.2%), the guard's tax converging to -0.016 (almost never touched).
Mechanism
A gated tax that is not paid is saved, so the policy moves to the cheapest state just outside the gate; a gated income that is not earned is simply lost, so a positive band has no exit. HoST's style penalties are ungated or binary - the same law seen from the other side.
Applies when
- adding a penalty multiplied by an uprightness, height, phase or contact gate
- a policy settles just beyond a gate threshold
- training metrics look paid-up while acceptance collapses
“**罚项的门 = 策略的逃生门**。带直立门的负项可以靠"变得更差"(倾出门外) 全时免税;HoST 的 style 罚全部无门控/二值恰是同一律的反面实证。修律: 负项只许挂"变差逃不掉"的门(上限护栏 always-on),或改正向 band (收不到 ≠ 逃掉)。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §46 定案(血统内第四败 + 两条新律) The walk phase machine structurally cannot express flight - rebuild the representation for duty < 0.5, teach flight with a mask tax, never a cliff bounty
phase-machine-structural-limitsWhen a new gait changes the contact pattern's structure, audit whether the phase/mask representation can express it and rebuild the representation if not; teach the new contact pattern with graded mask-mismatch pressure and count it in acceptance with artifact-proof definitions (minimum segment length), never with cliff bounties.
Symptom
Running requires both feet airborne, but the walk-era phase machine switches legs by the sign of sin(phase) - with duty < 0.5 the two swing windows must OVERLAP during flight, which a sign-switching representation cannot express at all.
Context
The run phase machine was re-architected rather than patched: per-leg phases (left = phi, right = phi+0.5 mod 1) with leg_phase < duty defining stance, aligned to the walk sin convention at duty=0.5 so the machines agree where their domains overlap. Flight is taught by the SAME mechanism that once cured foot-dragging, direction reversed: in the two planned flight windows the contact mask is (0,0) and feet_contact_number_duty charges -0.3 per foot still on the ground - a mild ~0.16/step tax, deliberately NOT a cliff: "悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文". The reference shape (half-sine bump over swing progress) is zero at window boundaries by construction, eliminating the clearing-window that the C4 probe measured to cost 13-19% on non-sinusoidal references. The shape self-check ("零代价选项是什么") was run on three behaviors: standing pays ref everywhere (known cmd=0 stepping risk, booked), walking pays only the flight-window tax, proper running collects full marks.
Change
New rewards.py run section (leg_phase_duty / stance_mask_duty / ref_run / clearance_run / contact_duty) with the walk versions untouched byte-for-byte; flight acceptance metric defined with a segment-length floor (>=40 ms to count) so numeric contact flicker cannot fake flight.
Outcome
Flight became expressible and taught by a calibrated mild pressure; the walk lineage's phase code stayed frozen as its own contract.
Mechanism
A phase representation defines which contact patterns exist in the reward's vocabulary; duty cycling below 0.5 introduces states (double-flight) outside a half-period sign convention's language, so no weight tuning can teach them. And rare desirable events taught by cliff-shaped bounties invite hacks (jumping in place); the graded mask tax prices the planned pattern without creating a jackpot.
Applies when
- extending a walking stack to running/jumping (duty < 0.5)
- a desired contact pattern never appears despite reward increases
- defining flight/contact acceptance metrics
“duty<0.5 时摆动窗 (1−duty)T > T/2,两腿摆动窗在腾空段重叠 —— sin 符号切腿的机制结构上表达不了"双脚同时在空中"。… feet_contact_number_duty 对"还踩着地"持续 −0.3/脚 —— 与 walk 治拖地同一机制,方向相反。不腾空的税 ~0.16/步 … 梯度温和不构成悬崖(悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文)。”
train/RUN_V0_SPEC.md § 5. 相位机设计 Removing a foot-spacing wall passed every simulated gate and made the feet collide on the real robot - nothing priced stance width in the two-foot phase, the policy narrowed to the simulator's self-collision floor, and real calibration offsets closed the last millimetres; the wall came back with a gate
removed-wall-returns-on-hardwareWhen a constraint is removed, name what will govern that quantity instead and add a gate for it; never let a simulator's collision floor be the margin, and when a gate is exceeded by a hair, record the exact numbers and hand the release decision to a person instead of quietly passing it.
Symptom
On the first real-robot try of oneleg_v0 (2026-09-16) the two feet collided; the user also judged the folded foot not high enough.
Context
The V0 reward table had dropped the feet_lateral_distance wall because it seemed to conflict with the hip adduction single support needs. In the two-foot command bucket no remaining term governed stance width, so the policy drifted narrower until the simulator's self-collision stopped it; the sim acceptance had no foot-spacing gate, so 40/40 said nothing about it. On the robot, calibration offsets consumed the margin.
Change
V0.1: the wall restored (-10, minimum 0.16 m), re-checked against measured numbers (a swing-phase lateral spacing of ~148 mm costs 0.12 per step, acceptable); fold weight 0.8 -> 2.0; a ninth gate: minimum foot spacing >= 100 mm and zero leg-contact frames. The removal was kept on record.
Outcome
oneleg_v0_1 (V0r2 model_2200) passed 39/40 with the spacing gate 40/40. The single miss (a 15.4 deg tilt transient against a < 15 deg limit during a side switch, steady 6.9 deg, everything else green) was recorded with its numbers and released for the user to overrule.
Mechanism
An unpriced degree of freedom drifts to wherever the simulator stops it; if that stop is the simulator's own collision model, the policy's margin on hardware is whatever the calibration error leaves.
Applies when
- dropping a reward term that looked redundant or conflicting
- hardware shows a failure no simulated gate measures
- a release candidate misses one gate row by a small amount
“V0 撤墙被真机证伪(2026-09-16):双脚桶没有任何项管站宽,策略贴 sim 自碰撞底线收窄,真机标定偏差一吃**双脚相碰**。 … min ≥ 100 mm 且腿碰 0 帧(eval_straight 同判据)—— … V0 真机双脚相碰暴露 sim 门未看脚距的缺口 … L s2 标称 tilt 瞬态 15.4°(门限 <15, 超 0.4°, 稳态 6.9°, 该跑其余全绿)——换侧瞬态蹭线, 判定放行留档, 用户可否决。”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §4 feet_lateral_distance 行 / §6 验收门 ⑨ / §8 核查单 7 Fix a too-deep nominal pose before adding any penalties - the default stance defines the basin training starts in
nominal-posture-before-penaltiesBefore tuning penalties on a degenerate gait, audit the nominal pose and height targets against morphology and published ratios; if the default stance encodes the degenerate behavior, fix it first - and recompute dependent quantities (init height) by FK, not by hand.
Symptom
Policy lived in a crouched shuffle; nominal knee angle was 0.5 rad (28.6 deg) - deeper than published configs (Unitree G1 0.3 rad / 17.2 deg, Booster T1 0.4 rad) - so the policy's starting point and its action-space center both sat inside the crouch basin.
Context
Initially ranked "secondary" in the local diagnosis, this was promoted to co-first priority by the cross-check against published reward tables, which states that with nominal knee flexion above ~0.4 rad, fixing the posture must precede adding any penalties ("改这个之前别加任何 惩罚都是白费"). Companion base-height items: walk profile had weakened base_height_l2 to -5.0 (base class -10, field standard -10 to -20, "the second most common cause of death"), and the height target must be the STANDING height (0.384), not the crouch height.
Change
Nominal knee 0.5 -> 0.3 rad with init_base_height recomputed by MuJoCo FK (0.3739 -> 0.3802); base_height_l2 restored to -10 with standing height target; both bundled as first-priority alongside the clearance term.
Outcome
Part of the v5/v6 package that lifted swing height to 34 mm and tracking to 87%; the crouch basin stopped being the default answer.
Mechanism
The nominal pose is the fixed point every regularizer pulls toward and the point where action=0 lands; if that point is itself the degenerate posture, every penalty fights the geometry. Correcting the attractor is prior to shaping the gradient field around it.
Applies when
- policy converges to a crouched or collapsed posture
- nominal joint angles were chosen for stability rather than gait
- base-height reward targets or weights were locally weakened
“研究明确说"nominal 膝屈超过 ~0.4 rad 必须先改,改这个之前别加任何惩罚"。我们是 0.50,超标。… base_height_l2 在 walk profile 里被减到 −5.0(基类是 −10)。研究说这是"第二常见死因"且应 −10 ~ −20。改回 −10。目标高度用站立高 0.384 是对的(研究要求 target 必须是*站立*高度而非蹲姿)。”
train/WALK_DIAGNOSIS.md § 修正 ①(升级优先级) / 修正 ④ An edge-triggered landing penalty missed the tail and fired after the harm - penalize overspeed continuously inside the contact window
penalize-tail-before-touchdownPenalties aimed at impact/violation events must (a) price the excess over a threshold, not the mean, and (b) be active on the approach (state-gated window), not triggered by the event - check your control rate can even see the event you are penalizing.
Symptom
The v7 landing penalty (vz^2 on the contact-force rising edge, weight -10) did not bite: landing-velocity 95th percentile stayed at 2.61 m/s against a 0.3 target.
Context
Two structural faults were identified: (1) it penalized the MEAN over sparse events - many soft landings dilute the occasional violent slam, while the damage (GRF peaks, motor peak load) lives in the tail; (2) it fired AFTER touchdown - at 50 Hz evaluation the rising edge is aliased by physics decimation, so the read vz is often the already-decelerated post-impact value: underestimated, and with no shaping gradient before contact. Replacement: continuous penalty while the sole is inside a height gate (h < 0.03 m): relu(-vz - 0.30) - only the excess over an allowed approach speed is penalized (tail only), and gradient exists for several frames BEFORE touchdown. The sole-height computation again subtracts the 0.0585 m link offset ("WALK_DIAGNOSIS 坑#1, 别再踩"); the edge-triggered version was kept as a diagnostic only.
Change
feet_landing_vel reformulated: edge-event vz^2 -> in-window relu(-vz - v_ok) with v_ok 0.30 (conservative vs the sqrt(L)-scaled human value ~0.19, to be tightened after passing), h_gate 0.03, weight unchanged -10.
Outcome
The failure analysis of the first form was written before the second was trained; the v_ok escalation path (0.30 -> 0.45 if the robot becomes afraid to land) was pre-registered in the risk table.
Mechanism
Sparse-event mean penalties optimize the average case while the constraint is a quantile; and any penalty evaluated only at/after a discrete event gives the optimizer no gradient along the approach trajectory that determines the event. A state-gated continuous excess penalty fixes both: it prices only violations and shapes the approach.
Applies when
- impact/landing penalties fail to move tail percentiles
- a penalty is triggered by contact edges at a coarse control rate
- designing constraint-style penalties for rare violent events
“罚的是均值路径:上升沿是稀疏事件 … 大量软着陆稀释偶发猛砸;而伤害在尾部 … 罚在触地后:50 Hz 评一次,上升沿被物理 decimation 混叠,读到的 vz 常是撞完已减速的值——既低估,又没有触地前的塑形梯度。”
train/WALK_V8_SPEC.md § 2. 改动 B — 落地惩罚改罚尾部、罚在触地前 PPO's Gaussian noise cannot compose phase-locked oscillations - deliver them as feed-forward and let the policy learn the residual
feedforward-for-phase-locked-skillsIf a skill needs a temporally coherent (phase-locked) action component, do not expect step-wise exploration to find it: inject a verified feed-forward and train the policy as a residual stabilizer, keeping the feed-forward inside the deployment contract.
Symptom
Four different reward arrangements (no reference / wrong-sign reference / correct-sign reference / cage released) all failed to elicit sidewalk, while open-loop probes proved the behavior existed and was safe on the same platform with the same policy as base.
Context
Producing lateral velocity requires a phase-locked hip_roll oscillation synchronized to the gait clock. PPO's exploration is per-step, zero-mean, uncorrelated Gaussian noise - it can never compose a sustained phase-locked component, so the behavior is unreachable by exploration regardless of how it is rewarded. The fix changed the delivery channel: target = default + scale*action + lat_ff(cmd_vy, phi). The policy's action becomes a residual on top of the feed-forward, retaining full balance authority (it can even cancel the feed-forward); the feed-forward supplies exactly the component exploration cannot. This mirrors why the sagittal joint_pos_ref worked (it also delivered phase structure), just via a different channel.
Change
Contract-level change, done cleanly: new profile omni_ff (= omni + lat_ff_gain -0.5), existing omni profile bit-identical; feed-forward applied after the action delay stage; missing cmd/phase raises instead of silently dropping; deployment must use the same phi as build_obs (recomputing gives a one-tick phase misalignment).
Outcome
From C2-700, +100 iterations sufficed: product omni_c4_ff800 scored vy +120%/+125% (from +4%/-1%), 260/260 cells at 20/20 survival, zero old-skill regression, left/right gap 5 pp - the entire C4 saga resolved by changing the delivery mechanism, not the reward.
Mechanism
Exploration noise spans only the subspace its correlation structure can express; skills requiring coherent oscillation lie outside the span of i.i.d. per-step noise. Feed-forward moves the required structure into the action pipeline where it needs zero probability mass to appear, reducing the learning problem to stabilizing around a demonstrated behavior - which PPO does well.
Applies when
- a periodic/oscillatory skill trains flat under every reward variant
- open-loop injection of the behavior already works
- considering GRU/curriculum/exploration tricks for a rhythmic skill
“病因不在奖励,在探索形式:产生侧向速度需要相位锁定的 hip_roll 振荡,PPO 的逐步高斯噪声零均值无相关,合不出相位锁定分量。… target = default + scale·a + lat_ff(cmd_vy, φ)。策略动作因此是前馈之上的残差,保留全部平衡权限”
train/C_LADDER_RUN.md § 3j. C4-redo4:唯一变量 = 侧步参考改为前馈注入(契约级) Four in-lineage attempts to widen the standing stance failed - remove a tax, add a joint-space knife, change the target, add a task-space metric penalty - because the stance was the end state of the get-up path; trained from scratch with the right terms it grew right from day one
stance-decided-by-get-up-pathA posture a skill ends in is shaped by the path the policy takes to reach it; if several single-variable edits to the terminal-phase reward cannot move it, stop editing that phase and retrain with the terminal constraint present from the start.
Symptom
v2_6c stood with its feet 0.159 m apart (task-space) and its hips yawed 45-47 deg the same way, which split on the real robot. Standing-phase reward edits did not move it.
Context
V2.7-A removed the flat-feet tax on compensated stances (stance unchanged); V2.7b added a hip-roll lower-bound hinge (+5 deg in 3,000 iterations, yaw ratchet); V2.8 changed the stand_pose target to a wide flat stance (stance unchanged, yaw not unwound, feet nearly overlapping, mu 0.4 transfer 2%); V2.9 penalized lateral spacing in metres (the policy parked just outside the penalty's gate in a lunge, 0% success). The v2_6c get-up goes through a split and closes the feet together as it rises.
Change
In-lineage stance surgery was formally closed. V3.1 trained from scratch with task-space stance terms present from the first iteration (and, after P1, a positive width band instead of a penalty).
Outcome
V3.1 P1b: lateral stance 0.364 m, foot tilt 0.0 deg, all four categories 100%, MuJoCo mu 1.0 and 0.4 both 100% - with a symmetric toe-out the kinematic audit had not enumerated. P1c (with a yaw guard): 0.355 m, all six acceptance criteria passing, mu 1.0-0.4 all 100%; it became the product.
Mechanism
A converged policy does not rebuild the path that produced its terminal posture; a standing-phase gradient only finds the nearest hack around the posture the get-up delivers.
Applies when
- the final posture of a transition skill is wrong and resists terminal-phase shaping
- repeated continuation rungs produce hacks instead of the intended posture
- deciding between another in-lineage fix and a from-scratch retrain
“窄站距 + yaw 扭是 v2_6c 起身策略(劈叉起身 → 双脚并拢收势)的**结构性 终态**,不是站立段的孤立参数 —— 站立形态由起身路径决定,在血统内只动 站立段奖励改不动它。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 结果:V2.8 判 FAIL —— 血统内站姿手术第三次证伪 The 1.5 Hz step-frequency gate was retracted - the author had misread his own actuator data, and the limit fought pendulum dynamics
gate-threshold-retracted-frequencyEvery gate threshold must cite its measurement and survive a first-principles sanity check; when a gate keeps failing otherwise healthy behavior, re-derive the threshold from the raw data before enforcing it again - and retract wrong gates in writing.
Symptom
An acceptance criterion "step frequency <= 1.5 Hz" kept failing healthy policies (walk_v4 at 2.33 Hz), and an earlier attempt to force slower stepping (v3) had killed stepping altogether.
Context
The threshold had been derived from the author's own actuator frequency-response measurements - but re-reading the raw table showed the misread: amplitude ratio at 2.0 Hz is 0.88 (knee) / 0.83 (ankle), acceptable; the genuinely bad point was walk_v2's 3.8 Hz at 0.58. Mechanism check agreed: the leg as a compound pendulum (L ~ 0.30 m) has a natural frequency ~1.1 Hz, swing half-period 0.45 s - the observed 2.1-2.4 Hz sits near where the leg wants to swing, and forcing 1.5 Hz "是跟摆动动力学对着干" (fights the swing dynamics). The frequency definition itself was pinned by two independent methods (contact-event counting 2.39 Hz vs FFT 2.33 Hz, agreeing): reported numbers are cycle frequency = steps per leg per second.
Change
Gate retracted in writing: "步频 ≤1.5 Hz 应删除或放宽到 ≤2.5 Hz"; frequency definition standardized before entering any config.
Outcome
walk_v4/v6's 2.1-2.4 Hz reclassified from disease to normal; the v3 failure got its probable explanation (suppressing stepping to meet a wrong gate).
Mechanism
A gate is only as good as the measurement and the reading behind it; thresholds inherited from a misread plot become invisible design constraints that later training obeys at real cost. Cross-checking a threshold against first-principles dynamics (pendulum frequency) is a cheap way to catch such misreads.
Applies when
- an acceptance threshold repeatedly fails policies that look healthy
- thresholds were set from a single person's reading of raw data
- a forced compliance with a gate degrades the behavior it guards
“我当初依据自己测的执行器频响定的,但看错了区间。… 2.0~2.4 Hz 的幅值比 0.83~0.88 是可接受的;真正不行的是 walk_v2 的 3.8 Hz。… 腿按复摆算(L≈0.30 m)自然频率约 1.1 Hz … 把它压到 1.5 Hz 是跟摆动动力学对着干(walk_v3 把迈步压没了,可能正是这个原因)。”
train/WALK_DIAGNOSIS.md § ② 撤回"步频 ≤1.5 Hz"这条验收标准 —— 是我定错了 Ungated phase shaping made standing 42x more expensive than stepping - and the stepping was cooking the hip motors
moving-gate-42x-stand-taxGate every phase/clock-driven shaping term on the command that justifies motion, price the cmd=0 case explicitly during design - and when a reward flaw is sim-only, book it with trigger conditions instead of operating immediately on a working lineage.
Symptom
At cmd = 0 the sim policy never stood still - it stepped in place and crept 0.98 m per 20 s; on the real robot the same lineage's stepping made hip_roll motors run 20 degC hotter than every other joint (43-48 degC vs 25-28).
Context
The arithmetic closed it: the three phase-shaping terms (joint_pos_ref 1.6, feet_contact_number 1.2, feet_clearance_swing 1.6) are driven by the gait clock with NO command gating, so standing at cmd=0 forfeits 3.32/step of shaping while honest standing earns only 0.078 of tracking - stepping wins 42x. The heat chain: perpetual stepping = perpetual single support = one hip_roll stalled at ~4.3 N*m (25% of torque limit) carrying the torso's frontal-plane moment - two hip_rolls = 90% of whole-machine steady-state I2R; measured stand-vs-step comparison: total heat -71% when actually standing. The suspended test acquitted the actuator (0.21 rad sag -> 0.0008 rad in air) and mechanics vetoed the easy fix ("降 kp 救不了热" - equilibrium torque equals the external load regardless of kp). The fix (moving_gate: hard-gate the three shaping terms on |cmd| > eps) was designed - then DEFERRED by the user because the real robot at the time stood fine: "真机不表现该问题, 为真机不存在的病改奖励表不划算", with written trigger conditions (C-ladder stand row persistently failing, or real robot starting to step/drift at cmd=0) and the known hazard tag (this is exactly the reward-change class that triggers the B-arm signature). When the real robot later DID step and cook, the booked trigger fired and moving_gate moved from debt to to-do with its benefit re-priced: "停止空烧 hip_roll,稳态发热降 ~71%".
Change
moving_gate designed with the gate_by_cmd convention; deferral, triggers, and expected heat recovery all pre-registered instead of patching the reward for a then-sim-only symptom.
Outcome
The cmd=0 stepping went from mystery to closed arithmetic; the thermal measurement (hip_roll +20 degC) quantitatively confirmed the 90%-of-heat prediction; the reward change waited for real-world justification instead of spending a risky revision early.
Mechanism
Clock-driven shaping terms define a perpetual-motion bounty unless gated by command; the resulting idle gait is not a training bug but the table's optimum. Its cost surfaces on hardware as stall-torque heating set by statics (mass x lateral offset), which no gain change can remove - only removing the motion (gating) or widening the stance can.
Applies when
- the policy steps in place or creeps at zero command
- specific joints run hot in idle behaviors
- deciding when a known reward flaw justifies a risky mid-lineage fix
“塑形合计 −3.32/步 … 净: 站定亏 42 倍 … 两颗 hip_roll 4.29 / 4.09 N·m(各占限扭 25%),占全机稳态 I²R 的 90% … 吊挂实测 hip_roll 跟踪误差 0.21 rad → 0.0008 rad … 降 kp 救不了热 … 真机不表现该问题, 为真机不存在的病改奖励表不划算”
train/README.md § C1 FAIL 节 (cmd=0 的 sim/真机分歧记账) / C2 真机 A/B 三b 发热定性 Before training a one-leg stand the spec named the cheapest cheats - hopping on the support foot, a raised foot resting unloaded, a leg tripod - and gave each a countermeasure and a gate; one still appeared and was caught by exactly those gates
enumerate-cheapest-cheats-before-trainingBefore training, list the cheapest behaviours that would satisfy each reward term without doing the task, give each a countermeasure in the reward and a gate in acceptance, and prove the intended behaviour is reachable with a probe - then treat any gate the policy games as evidence about the reward, not the gate.
Symptom
The literature's single-leg benchmark reports eight state-of-the-art general policies holding a clean one-leg stand 0 times out of 90 - they survive by sneaking steps and hops - so the task's first adversary was the policy's own cheating.
Context
The spec's shape self-check ("what is the zero-cost option?") listed, for the one-foot bucket: the cheapest cheat, a foot resting on the ground without load, countered by a 5 N contact threshold plus positive swing income; the second cheapest, small hops on the support foot to reset balance, countered by a continuous support-air penalty plus a gate of zero support-foot flight segments. The probe that preceded training had already seen a third: early low-lift postures "survived" by pressing the swing foot at 78-95 N, a leg tripod, removed by folding the shank back. The two-foot bucket was checked too: its zero-cost behaviour is ordinary standing, with no odd base state.
Change
Countermeasures and gates written before training: swing-contact and support-air penalties, gate 2 (zero swing-foot contact frames above 5 N), gate 3 (zero support-foot flight segments).
Outcome
The first run still found the unloaded-foot cheat (a binary reward band gave it no gradient to lift) - and it was caught, by the contact gates and the cross-simulator comparison, not discovered on hardware. The retrained V0 passed all gates 40/40, including zero support-foot flight after the flight detector was corrected.
Mechanism
A policy optimizes the reward, not the intent; the cheapest behaviours that satisfy the reward are predictable from the reward's structure, and a gate written for each before training turns a silent cheat into a failed row.
Applies when
- designing rewards for balance, contact or "hold still" tasks
- benchmark policies are known to cheat the task
- writing acceptance gates for a new skill
“文献里 8 个 SOTA 通用策略在单脚站基准上 0/90 干净保持, 全靠偷步偷跳活命,这是本任务的第一反作弊对象 … 单脚桶下最便宜的作弊是"脚虚放地上不受力"——接触判定 >5 N 力阈(沿用),配 swing_height_band 正收入拉开。 … 第二便宜是"支撑脚小跳重置"——support_air_penalty 连续罚 + 验收门支撑脚腾空段=0 双保险。”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §0 目标口径 / §5 形状自检(零成本选项是什么) A stand gate judged by survival passes a robot that wanders a meter - judge posture instead
stand-gate-posture-not-survivalFor every gate, ask what behavior the metric is a proxy for and validate its ordering against real observations; replace metrics whose ordering disagrees with reality, and demote them explicitly rather than silently.
Symptom
Real robot "standing" drifted 0.5-1.4 m across the floor while the sim stand gate scored a clean 20/20 - because the gate's metric was episode survival, which wandering does not violate.
Context
C2's stand condition was originally written as "survival regression <=2/20". A real-robot counter-example on 2026-08-08 forced the re-judgment: wandering robots survive. Cross-checking candidate sim metrics against real-robot feel showed max-tilt median ordering agreed with hands-on ranking, while displacement ordering was actually OPPOSITE to real impressions - so displacement was demoted to a reference quantity, not a gate.
Change
Stand PASS criterion rewritten from survival to posture: "stand tilt max median <= root baseline +1.5 deg"; displacement kept only as reference. Applied to all subsequent rungs (C4 and redo levels inherit it).
Outcome
Later rungs gated stand on tilt (e.g. C4 product: 7.7 deg vs parent 7.1 deg, +0.6 deg PASS); the wandering failure mode became visible to the battery instead of hidden by survival.
Mechanism
A gate metric is a proxy for an intended behavior; survival is a proxy for "did not fall", not "stood still". Metric choice must be validated against ground truth (real-robot feel/measurement), and a proxy whose ordering disagrees with reality on real data is worse than no metric - it steers selection backwards.
Applies when
- writing PASS conditions for stand/idle/hold behaviors
- a gate passes policies that visibly misbehave on hardware
- choosing between candidate metrics for an acceptance battery
“stand 必须用位姿判,不能用存活判(2026-08-08 真机反证改判):站着乱走 0.5~1.4 m 时存活照样 20/20 —— C2 的 stand 条件原写「存活退化 ≤2/20」,选错了指标。改为: stand 倾角 max 中位 ≤ 根基线 +1.5°(倾角与真机手感排序一致;位移排序与真机相反,降为参考量)”
train/C_LADDER_RUN.md § 3b. PASS 条件 ⚠️ stand 必须用位姿判 FK-verify a borrowed reference's structure, then size its amplitude by the reference's job - it pins phase, the policy adds lift
reference-structure-fk-amplitude-divisionWhen borrowing a reference trajectory: verify its structural claim against your own kinematics (an invariant like flat-foot), assign it the phase-pinning job, and size amplitude low enough that the policy contributes the lift - moving toward a proven foreign value in halves, not jumps.
Symptom
walk_v4 had big knee swing (40-46 deg) but only 18-24 mm foot lift - amplitude without hip/knee/ankle phase coordination; later, walk_v5's real-robot swing ballooned to 73.6 deg (sim 55.7) with violent footfalls - amplitude over-driven by the reference.
Context
Structure first: Humanoid-Gym's 1:2:1 hip:knee:ankle reference was verified on the local model before adoption - the ratio exactly satisfies the locally derived flat-foot constraint hip - knee + ankle = 0, FK-tested at multiple amplitudes with sole pitch 0.00 deg throughout. Amplitude second, and here the first reasoning failed honestly: FK said shorter legs need LARGER reference scale (0.30 for 30 mm lift), and the FK was correct - but the premise was wrong ("FK 没错, 但前提错了"): it assumed foot lift must come from the reference. HighTorque Pi, same scale, uses 0.08 with a 0.02 m foot-height target - proof that lift is added by the policy ON TOP of the reference, whose actual job is pinning the phase relationship. Scale 0.30 made the reference the entire gait: over-constrained and over-driven. The correction went to 0.15, deliberately not Pi's 0.08: "一次只走一半, 留退路" (walk half the distance, keep a retreat).
Change
target_joint_pos_scale 0.30 -> 0.15 as one of v6-minimal's three changes, treating both the footfall force and the lateral kicking (yaw momentum scales with leg swing amplitude).
Outcome
v6 improved landing force 1.72x -> 1.55x, suspended tilt 45.9 -> 23.0 deg, turn-gain asymmetry 70% -> 19%; the later v6-halved-shaping experiment (35 mm -> 4 mm collapse) confirmed the reference still carries the gait's existence on this machine - the division of labor is real but machine-specific.
Mechanism
A joint-space reference plays two separable roles: encoding structure (phase relations that keep the foot flat) and injecting amplitude (energy). Structure transfers across robots and is checkable by FK against an invariant; amplitude is a negotiation with the policy, and over-assigning it to the reference removes the policy's freedom to modulate lift with state.
Applies when
- importing a reference gait / imitation target from another codebase
- reference amplitude reasoning based on leg length alone
- real swing amplitude far exceeds sim's under a strong reference
“FK 没错, 但前提错了。我默认抬脚必须由参考轨迹产生。HighTorque Pi 同尺度机器人 … 用 0.08, 而它 target_feet_height = 0.02 m —— 说明抬脚是策略在参考之上加出来的, 参考只负责钉住髋/膝/踝的相位配合。我们取 0.30 等于让参考本身就是整个步态, 过约束 + 过驱动”
train/WALK_V6_MINIMAL.md § ① target_joint_pos_scale 0.30 → 0.15 A 10 s acceptance episode left 6-7 s of standing to observe - a narrow stance held for that window and split on hardware; a gate cannot see instability slower than its own horizon
episode-length-bounds-what-a-gate-seesSize the standing phase of an acceptance episode, and its disturbances and floor friction, to what deployment will impose; a pass on a short static window certifies only that window.
Symptom
v2_6 passed every simulation gate (success 99.6%, re-falls 0-1%) and then, on the real robot, stood up and slid into the splits several times; prone starts stood and then fell backwards.
Context
Acceptance ran 10 s episodes; a ~1-2 s get-up left roughly 6-7 s of static standing on a nominal floor with no disturbance. The narrow stance's lateral margin and the straight-knee stance's lack of any flex buffer are both failure modes that need time, disturbance or lower friction to show.
Change
The gap was booked as a known blind spot of the gate ("long-duration standing stability") alongside the task-space stance criterion; later rungs added MuJoCo friction sweeps at mu 0.4 to every checkpoint scan.
Outcome
The spec through §50 records the blind spot but no longer standing window or disturbance row in the recovery acceptance itself.
Mechanism
An acceptance episode observes only the dynamics that unfold within its horizon under its conditions; slow drifts and disturbance-triggered failures are outside it by construction.
Applies when
- a policy passes sim gates and fails on hardware after a delay
- acceptance episodes are short relative to deployment use
- stability is judged without pushes or friction variation
“**sim 门为什么没逮住**:10 s episode 起身后只站 ~6-7 s,静态窗口内窄站距 撑得住;真机站立时长/扰动谱在门口径之外 —— 长时站立稳定性记为口径缺口。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §43 判读:sim 门为什么没逮住 A single-signal contact detector lied in both directions - foot height flagged 40% false flight on a walking gait, contact force alone flagged false flight during low-friction slip - so flight became force < 5 N AND sole height > 5 mm
contact-detector-single-signal-liesDefine contact and flight events from two independent signals (force and geometry) in conjunction, validate the detector on a behaviour known not to contain the event before using it as a gate, and match the trainer's threshold when comparing across simulators.
Symptom
The run line needed a flight-fraction gate. The first MuJoCo version, "sole higher than 2 mm", measured 40% flight on a walking policy that never flies. Five weeks later the one-leg gate, using contact force alone, reported support-foot "flight" segments at mu 0.4 for a policy that was not hopping.
Context
During a walking step the toe lifts or the heel strikes with the foot pitched, so the ankle-roll origin rises a few mm while part of the sole still touches - height alone calls that flight. Switching to contact force < 5 N (the same threshold Isaac's contact reward uses) zeroed the false flight on walking. In the one-leg re-test every force-only "flight" segment had a measured sole height of 0.0 mm: the normal force chattered while slip corrections played out on low friction.
Change
Run line: flight = contact force < 5 N ("lift-off must be judged by contact force"). One-leg line (2026-09-16): flight = force < 5 N AND sole height > 5 mm, recorded as the same measurement lesson in the opposite direction; the gate's behavioural meaning was unchanged.
Outcome
With force-based detection the walking policy read 0.0 flight and the run policy's zero flight was confirmed by two plants; with the conjunctive definition the one-leg support-foot gate stopped reporting false hops.
Mechanism
Each signal has its own failure: geometry moves without leaving the ground (foot pitch), and forces drop without leaving the ground (slip chatter); requiring both removes both families of false positives.
Applies when
- writing a flight, lift-off, hop or slip detector for a gate
- a gate reports an event the video does not show
- reusing a detector on a different gait or floor friction
“首版用**足底高度>2mm** 判离地, 在 omni_s1e 走路策略上测出 40% 假腾空 … 改用**接触力 <5N**(与 Isaac feet_contact_number 同源阈值)后 空检归零 (walk 策略 flight_frac 0.0)。**课文: 离地判定必须用接触力, 高度判 会把脚的俯仰当腾空**”
train/README.md § run R1 立项 (2026-08-09): Mac 侧新工具 + 一次空检抓获 A joint-velocity penalty meant to slow the get-up cut joint speed 16% and left the get-up time unchanged - the knob never moved the variable, so the idea it was meant to test stayed untested
dof-vel-penalty-is-not-a-pacing-knobBefore reading a result as a test of an idea, check that the knob actually moved the independent variable; velocity regularizers smooth a schedule they do not set, and a schedule driven by per-step task income moves only when that income's time structure does.
Symptom
The get-up took 0.6-0.9 s in Isaac with large torque demand; the user proposed getting up more slowly so less torque would be needed.
Context
The idea had support in the accounts: the acceptance bound is an upper bound of 5 s (5-8x margin), the quasi-static squat path peaks at 25% of the limits, and rolling over needs no momentum. R3.2 raised dof_vel from -1e-3 to -5e-3 as the single variable.
Change
dof_vel -1e-3 -> -5e-3 (child-run from R3.1).
Outcome
Get-up medians moved +0.02-0.04 s (noise); raw joint velocity -16%; torque demand median got worse (hip_pitch 46-48% -> 63-67%) as the new term competed with torque_headroom on the same joints; MuJoCo 98 -> 96%. Verdict FAIL on the knob, not on the idea, and the rung was not adopted. When pace was later attacked through the income's time structure (V2.5/V2.5b), the MuJoCo get-up moved into the 3.5-4.5 s design band.
Mechanism
The pace was set by base_height_progress paying for every step spent high (stand earlier, earn more); a velocity regularizer only smooths motion along the same schedule and does not change when the robot stands up.
Applies when
- trying to make a skill slower or gentler with smoothness penalties
- an experiment's primary metric did not move and a verdict is being written
- two penalties act on the same joints
“**关键判读:`dof_vel` 罚只把关节速度压了 16%,而起身用时一点没变。** 也就是说**这一级根本没有把"慢下来"这个自变量推动起来** —— 所以它**不构成对 用户假说的检验** … 起身节奏由 `base_height_progress` 的逐步计酬决定(早站起来就多 拿),速度正则只在同一条时间轨迹上把动作抹匀,不改变何时站起来。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §25 R3.2(dof_vel −1e-3→−5e-3,慢一点起身) Cross-simulator gate (Isaac Lab to MuJoCo) comes before any hardware attempt
sim2sim-gate-before-sim2realGate every policy through a second simulator with an independently built plant before hardware; treat sim2sim failure as a contract or overfitting bug, and sim2real failure after a sim2sim pass as a plant/actuator gap.
Symptom
A policy that only ever ran in its training simulator carries untested dependencies on that simulator's solver, contact model, and defaults; the first place those dependencies surface should not be the real robot.
Context
Standing order of operations for the whole Lucen program, recorded as the opening line of the experience log: train in Isaac Lab, gate in MuJoCo, only then go to hardware. The MuJoCo side is the same plant used for evaluation batteries, so a sim2sim pass also validates the exported policy + contract (obs ordering, scales, defaults) outside the training stack.
Change
Pipeline rule adopted - every checkpoint must pass the MuJoCo evaluation battery (sim2sim) before it is considered for real deployment (sim2real).
Outcome
Used generation after generation as the cheap filter; hardware sessions only ever received policies that had already survived a second simulator.
Mechanism
Two simulators disagree exactly where a policy is overfit to simulator-specific artifacts (contact softness, integrator, default parameters, obs conventions); a cross-sim transfer catches contract bugs and solver overfitting at zero hardware risk, so real-robot failures that remain are attributable to genuine plant/actuator gaps.
Applies when
- planning the path from training to first hardware trial
- exported policy behaves differently outside the training framework
- triaging whether a real-robot failure is contract vs plant
“先sim2sim - 从isaaclab 到mujoco / 再sim2real”
Experience.md § opening lines (1-2) The hip_roll (l+r) asymmetry scalar predicted real-robot lateral drift - promote validated sim scalars into the gate
hip-roll-sum-predicts-lateral-driftHunt for cheap sim scalars that predict real-robot behaviors, validate them on direction AND ordering across multiple policies, then promote them into the acceptance battery; treat later violations as debt to justify in writing, not noise to ignore.
Symptom
A persistent hip_roll left/right asymmetry row in the sim2sim symmetry table had been dismissed as "calibration or mechanical asymmetry" noise; meanwhile real deployments drifted sideways by policy-dependent amounts.
Context
Forward-kinematics analysis reframed the scalar: both hip_rolls move the feet in +y for positive angle, so a same-signed (l+r) sum IS a lateral translation mode - the scalar is a direct lateral-drift bias estimate. Checked against real deployments: s1e (l+r = -0.0178, smallest magnitude) was the steadiest with least drift; 700 (+0.0253) drifted mildly left; A800 (+0.0267) drifted clearly left with the largest tilt 12.9 deg. Direction correct 3/3, ordering correct 3/3 (the log's heading calls it "四枚四中", four-for-four).
Change
The scalar was promoted into the acceptance battery as a posture-class criterion alongside tilt-max median: "hip_roll 左右不对称 |l+r| 不得比父代大" - doubling as a heat proxy (error ~ torque ~ heating).
Outcome
Used at every later gate; when the C4 product exceeded it by +0.005 rad (~+0.3 deg vs parent), the criterion was not silently waived - it was booked as explicit debt with a mechanism argument (the increment is task-required, far smaller than the sidewalk amplitude +/-2.2 deg) plus a related account (stand saturation 32.4% -> 37.2%).
Mechanism
A policy's static joint-angle bias in a translation-producing mode integrates into real-world drift; sim can measure that bias precisely and cheaply. A sim scalar earns gate status exactly when its predictions are validated against hardware in both direction and ordering - and a validated gate may only be exceeded with a written mechanism-level justification, never silently.
Conflicts
The log's heading says "四枚四中" (4/4) but the evidence table lists three policies and the text says "方向 3/3、排序 3/3"; the fourth instance is not shown in this file.
Applies when
- a real robot drifts or leans in a policy-dependent way
- deciding which sim measurements deserve gate status
- a validated gate criterion is marginally exceeded by a new product
“s1e | −0.0178(绝对值最小)| 微右、最不飘 | 三者中最稳、飘最小 ✓ … A800 | +0.0267 | 左、最飘 | 明显左飘、倾角最大 12.9° ✓ 方向 3/3、排序 3/3。 → 正式纳入验收表(与「倾角 max 中位」并列为姿态类判据)。”
train/C_LADDER_RUN.md § 3e. 顺带:hip_roll 左右不对称 (l+r) 就是横移偏置 —— 四枚四中 Decompose the offending quantity by channel first - then penalize the failure event, not the joints
penalize-the-slip-not-the-jointBefore penalizing motion to fix a side effect, measure which channels actually carry the offending quantity; prefer penalties conditioned on the failure event that are exactly zero for healthy behavior - and do not medicate behaviors that measurement shows are not sick.
Symptom
Heading drift with support-foot yaw slip (v5: 212-284 deg accumulated over 15 s); the previous v6 draft had attacked it by penalizing lateral joints (a roll 4.0 / yaw 2.0 "home" group) - which collapsed training into the standing basin.
Context
Before choosing the penalty target, the yaw angular momentum was decomposed by joint group with MuJoCo subtree_angmom weighted by real walking joint velocities: pitch-class joints (hip_pitch + knee) carry 95.3%, hip_roll 3.3%, hip_yaw 1.4%. The failed "home" group had been taxing 2.7/step to manage a 4.7% channel. The replacement, feet_yaw_slip (-0.2, |support-foot yaw rate| while in contact), targets the failure event itself and - decisively - costs a non-slipping gait exactly zero, which "横向回家组做不到". The same rung's do-not-do table applied the complementary principle to foot spacing: measured 196-214 mm, stable, no crossing - "没病不吃药" (no disease, no medicine).
Change
Removed joint-usage penalties for the drift problem; added the event-conditional slip penalty (-0.2, realized tax 0.141/step = 12% of tracking) alongside the existing linear-slip term.
Outcome
Turn-gain left/right difference improved 70% -> 19% and heading 185 -> 60.3 deg by v6 without a standing-basin collapse; the 2.7/step lateral tax never returned.
Mechanism
Penalizing joints taxes every use of a channel including healthy use, and if the channel carries little of the offending quantity the tax buys nothing while pushing the optimum toward immobility. An event-conditional penalty (slip while in contact) prices only the failure, leaving the healthy gait's cost surface untouched - and the channel decomposition tells you in advance whether a joint-side fix can even work.
Applies when
- choosing a penalty target for drift/slip/impact problems
- a proposed penalty taxes joints or motions rather than failure events
- a previous joint-penalty attempt collapsed the gait
“pitch 类 (hip_pitch + knee) 占偏航角动量 95.3% … hip_yaw 1.4% … 压 hip_yaw 是管 1.4% 的通道收 2.7/步 的税 —— 上一轮正是这样把策略推进了站立盆地。滑移项不惩罚走路: 不打滑的步态代价为零, 这是横向"回家"组做不到的。”
train/WALK_V6_MINIMAL.md § ① / ② 新增 feet_yaw_slip The shipped checkpoint was chosen by scanning checkpoints on the full gate - neighbours 100 iterations apart failed 1 and 38 cells, late checkpoints degraded - never by taking the last one, and training stopped on signals, not on a schedule
checkpoint-choice-is-a-full-gate-scanChoose a release checkpoint by running the full acceptance battery over a band of checkpoints (including the transfer axis), stop training on measured signals rather than a fixed iteration count, and expect adjacent checkpoints to differ sharply.
Symptom
Gate results moved sharply and non-monotonically between checkpoints of the same run, and the last checkpoint was often not the best.
Context
One-leg V0r1: the 2,000 neighbourhood was best; from 2,500 on the nominal gates degraded (late overtraining); 2,000 itself had one real micro-hop (17.7 mm over 5 frames); 2,300 was all green and shipped. V0r2: failed cells per checkpoint 2,000:19, 2,100:38, 2,200:1, 2,300:3, 2,400:27, 2,500:12, 3,000:18 - 2,200 shipped. The recovery line learned the same from the other side: stopping v2_6 early at a scheduled point left a policy whose re-fall rate had spiked to 9-22% before consolidation healed it ("stop on signals, not on the schedule"), and a continuation's transfer decayed checkpoint by checkpoint while Isaac stayed perfect.
Change
The acceptance rule "scan checkpoints, do not look only at the last one" is written into the one-leg gates (called the S1 discipline); release candidates are chosen from the scan.
Outcome
Both one-leg releases were mid-run checkpoints (2,300 and 2,200) chosen by the full 40-cell battery.
Mechanism
PPO keeps changing the policy after the gates saturate; with no gradient toward the gate's conditions, later checkpoints wander, so gate quality is a noisy function of iteration.
Applies when
- picking which checkpoint of a run to export and stamp
- a run is stopped at a fixed iteration budget
- final-checkpoint results are worse than mid-run smoke tests
“Isaac 侧 S1 纪律: 验收扫 checkpoint,不是只看最后一个。 … 扫描判决: 2000 邻域最优——2500+ 标称面退化(⑤③② 散挂, 晚期过训), 2000 有一例真微跳(L s100 μ1.2, 17.7mm/5帧), 2300 全绿。”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §6 验收门 / §8 核查单 5 与 7 Run acceptance under measured contact parameters - honest condim/torsional-friction flipped a false PASS into a real-matching FAIL
eval-plant-honesty-contact-paramsPin the evaluation plant's contact model to measured values (contact dimension, torsional/rolling friction, mu) before trusting any gate that involves slip, impact, or drift - a gate can only fail a policy for physics its simulator contains.
Symptom
walk_v5 passed the old acceptance battery yet failed on the real robot (footfall force, drift, kicking) - the evaluation plant was flattering the policy.
Context
The battery was re-run under "honest contact parameters" - condim 4 (adding torsional contact) with measured torsional friction 0.035 - and v5 then FAILED exactly the rows corresponding to its real problems: heading 185 deg (limit 30), support-foot yaw slip 284 deg (limit 80), landing force 1.72x (limit 1.5x), suspended tilt 45.9 deg (limit 10). The slip physics depends on torsional friction, which the default contact model (condim 3) does not even simulate - a slip problem is invisible to an evaluator that cannot represent yaw friction at the foot. Term-sizing measurements for the new rewards were likewise taken under the same honest parameters (cmd 0.45, skipping the 5 s start transient).
Change
Acceptance harness pinned to condim 4 / torsion 0.035 (measured); verdicts issued under defaults declared non-citable for these rows.
Outcome
Sim acceptance verdicts began agreeing with hardware ("现在失败, 与真机一致"); the v6 fixes could be developed and validated against an evaluator that could actually see the disease.
Mechanism
An evaluator is a plant model too: contact dimensionality and friction values decide which failure modes exist in the simulation at all. Evaluating under default contact parameters tests the policy in a world where its real failure is physically impossible, producing structurally false PASSes.
Applies when
- sim acceptance passes policies that fail on hardware
- slip/drift/impact gates run under default simulator contact settings
- setting up a cross-simulator evaluation harness
“accept_v2.py 已加三条判据, walk_v5 在诚实的接触参数下(--condim 4 --torsion 0.035)现在失败, 与真机一致:直行 15s 航向累计 <30° | 185° ✗ … 落脚力峰值 <1.5× 体重 | 1.72× ✗”
train/WALK_V6_MINIMAL.md § 5. 验收 Calibrate a wall penalty by measuring healthy and sick policies - healthy pays ~0, the disease pays a wall
calibrate-threshold-between-healthy-and-sickCalibrate every threshold penalty by evaluating its exact formula on replays of at least one healthy and one sick policy: place the threshold between their distributions, size the weight so the sick policy pays a decisive fraction of tracking while the healthy one pays ~0, and pre-compute neighboring thresholds for cheap adjustment.
Symptom
Real walk_v7 occasionally clipped its own legs (stance narrowed to 133 mm mean vs nominal 214.5); a foot-distance penalty was needed, but an uncalibrated threshold/weight risked either doing nothing or becoming a reverse barrier.
Context
The term (relu(d_min - lateral foot distance), measured in the base yaw frame because world-frame y is meaningless after turning) was calibrated before training by replaying three known policies through the exact reward formula at the acceptance operating point: healthy v5 (183 mm) pays 0.6% of tracking - effectively free; narrowed v7 (133 mm) pays 22% - effective widening pressure; collapsed v8 (93 mm) pays 55% - a wall. d_min 0.16 was placed deliberately between healthy and sick, with alternative thresholds (0.14/0.18) pre-computed in the tool output for later adjustment. The shape self-check was named as a standing question: "先问'零代价的选项是什么'" - the zero-cost region must be exactly the desired behavior.
Change
feet_lateral_distance added at d_min 0.16 / weight -10, with telemetry expectation pre-registered (should decay toward 0 as stance learns >160 mm; if bow-legged over-widening >214 appears, only then discuss an upper bound).
Outcome
v9_probe on hardware: no leg contact ("没碰腿(N2 兑现)"), stance min 145/126 mm green; the term's zero-cost design left healthy gait untaxed.
Mechanism
A relu threshold penalty defines a free region and a priced region; its correctness is entirely in where the boundary sits relative to the healthy and pathological distributions. Replaying known-good and known-bad policies through the exact formula measures both distributions in the term's own currency, making the threshold and weight a placement decision instead of a guess.
Applies when
- adding any relu/threshold-style wall penalty
- a safety margin (foot distance, joint limit, clearance) needs enforcement without taxing normal behavior
- choosing between candidate thresholds for a new term
“形状自检 (v6 横向组/v8-B 的教训 —— 先问"零代价的选项是什么"): 标称站距付 0, 健康步态付 ~0, 收窄才付费 … v5(健康) 183 mm … 0.6% ≈ 免费 | v7(收窄) 133 mm … 22% —— 有效推宽 | v8(塌陷) 93 mm … 55% —— 墙 … d_min=0.16 恰在 v5(183)与 v7(133)之间”
train/WALK_V9_SPEC.md § 2. N2 —— 脚距惩罚(已完成权重预标定) Gate a new reward term by its command so all old modes score pointwise identical
gate-new-reward-terms-by-commandWhen a reward term must be added mid-lineage, gate it on the condition that defines the new task so every pre-existing situation scores exactly as before - and still watch for value-rescale pathologies inside the new mode.
Symptom
Adding a lateral tracking reward (track_lin_vel_y_exp) ungated would have paid 0-2.0 per step even in modes with cmd_vy = 0 (healthy gait sway of vy ~0.1 already earns 1.28), shifting the whole reward table by a large bias and rescaling the value function - no longer "just adding one mode".
Context
C4 was the C ladder's only true reward surgery. Single-variable discipline required that the change be invisible to every existing mode. The chosen construction: gate_by_cmd=True - the term pays only when |cmd_vy| > 0.02, so for all modes with cmd_vy == 0 the term is pointwise zero, i.e. the reward is pointwise identical to before the change. The same trick appeared earlier in C1: replacing the vy L2 tax with a command-error version that is "对 cmd_vy≡0 逐点同值" (pointwise equal when cmd_vy is 0), explicitly classified as not-a-reward-change.
Change
track_lin_vel_y_exp added with gate_by_cmd=True (weight +2.0, std 0.15); the residual acknowledged honestly - inside the side bucket the values DO change, so the rung still watched the known reward-reshuffle pathology signature (s1c B-arm: scatter -> half-recover -> collapse) as a stop criterion.
Outcome
Old modes provably unaffected (pointwise-equal argument); attribution for any change in old-skill metrics stayed clean through the C4 redo series.
Mechanism
PPO's critic normalizes to the reward scale it sees; an ungated additive term shifts returns in every state and re-scales advantages globally, entangling the new skill with all old ones. Command-gating confines the new term's support to the new mode's state distribution, making "pointwise identical elsewhere" a provable property rather than a hope.
Applies when
- adding a tracking/shaping term for a new command or skill to a lineage that must not regress
- reward change proposed while other skills are still being gated
- reviewing whether a config diff counts as a reward change
“只在 |cmd_vy| > 0.02 时付。不门控的话它对 cmd_vy≡0 的老模式也给 0~2.0 分(健康摇摆 vy≈0.1 → 1.28),等于给整张奖励表加一个大偏置、值函数尺度全变 … 门控后老模式逐点得 0 = 与加项前逐点同值,单变量纪律成立。… 但 side 桶内的值确实变了 —— 这仍是奖励表改版,开级盯 s1c B 臂签名”
train/C_LADDER_RUN.md § 3d. gate_by_cmd=True(重要) Prove a new penalty actually fires - two ways a clearance term silently did nothing
inert-reward-term-auditBefore training with a new reward term, log its realized per-step value under the current policy and confirm it is nonzero where intended - check coordinate zero-points against FK and check who occupies the term's gate; and never weaken the term that creates the states your new term needs.
Symptom
A newly designed swing-height clearance penalty could have trained as a no-op twice over, and the companion advice to lower feet_air_time actively backfired when tried.
Context
Instance 1 (zero-point offset): the proposed code used body_pos_w of the foot link, but that is the ankle_roll_link frame origin, which sits 0.0585 m above the ground even with the foot flat on it - so (0.03 - 0.0585) is always negative and the penalty is永远 0; the 0.0585 offset must be subtracted (verified identical in MuJoCo FK and Isaac). Instance 2 (gate occupancy): the clearance penalty fires only in swing phase; a dragging policy keeps both feet in contact, so the penalty is constantly 0 for exactly the policy it was meant to fix - and worse, any slight lift immediately incurs it, a reverse threshold. Lowering feet_air_time to 0.5 on that advice measurably collapsed air time to 0.0002 (below v3). Corrected understanding: "clearance 是把已有的摆动相抬高, 造出摆动相仍要靠 air_time" - air_time creates the swing phase, clearance raises it.
Change
Fixed the height zero-point; kept feet_air_time as the swing-phase creator with clearance layered on top; both errors documented as corrections to the team's own earlier advice.
Outcome
With both fixed, swing height rose from 22-23 mm (v2) to 29 mm (v5) to 34 mm (v6); the inert-term failure class entered the standing checklist.
Mechanism
A penalty's gradient exists only where its gate is occupied and its argument crosses its threshold; frame offsets shift the threshold out of reach, and phase gates can have zero occupancy under exactly the policy being treated. Terms interact as an ecology - one term must create the states in which another can act.
Applies when
- adding any gated or thresholded penalty (clearance, impact, slip)
- a new term produces no behavioral change at any weight
- body-frame positions are used in reward code
“body_pos_w 是 ankle_roll_link 坐标系原点,平放触地时仍高出地面 0.0585 m。… (0.03 − 0.0585) 恒为负 → 惩罚永远是 0 … clearance 惩罚只在摆动相生效,拖地时两脚始终触地 → 惩罚恒 0;而一旦轻微抬脚就立刻扣分,对正在拖地的策略是反向门槛。… 正确认识:clearance 是"把已有的摆动相抬高",造出摆动相仍要靠 air_time。”
train/WALK_DIAGNOSIS.md § walk_v4 独立验收 — 本文档给的两处代码/建议是错的 Single-impulse push recovery is a binary chaotic quantity - cross-machine floating-point divergence can flip the outcome
single-impulse-recovery-is-chaoticNever gate or compare single-event recovery outcomes across machines or domains: evaluate disturbances as survival distributions over phases and seeds, compare longitudinally on one machine, and treat any single-point cliff as unconfirmed until it survives the statistical protocol.
Symptom
Mac evaluation found a hard "0.8 N*s cliff" (0/3 survival) that the training machine flatly contradicted: the identical protocol (0.8 impulse at 8 s, cmd 0.2) survived 3/3 there, and a 0.6/0.8/2/4 cross sweep survived everything.
Context
The verdict became a named lesson ("跨机混沌课文"): whether one specific push at one specific phase is survived depends on a trajectory that diverges across machines from floating-point differences alone - "单次冲量恢复是二值混沌量, 跨机浮点发散可翻结局". The boundary was drawn precisely: the 20-seed statistical gates DO agree across machines (established precedent), but that agreement cannot be extrapolated to single-point recovery tests. Protocol amended: disturbance evaluation uses multiple push phases (8/10/12 s), >=10 seeds, and only same-machine longitudinal comparisons; the Mac-side recommendation built on the unreproducible cliff was not adopted, while its directionally-consistent small-impulse data was kept.
Change
Push evaluation redefined from single-event pass/fail to multi-phase multi-seed statistics, with cross-machine comparison banned for event-level results and allowed for distribution-level ones.
Outcome
A false hardware-relevant "cliff" was prevented from steering the ladder (the s2e push rung decisions were made on same-machine statistics); the chaos lesson was cited again when real push tests were restricted to qualitative cross-domain use.
Mechanism
Perturbation recovery near the viability boundary has sensitive dependence on initial conditions; different BLAS/GPU reduction orders yield different trajectories from identical configs, so a binary outcome at one phase is machine-specific noise. Averaging over phases and seeds restores a quantity whose expectation is machine-stable.
Applies when
- a push/disturbance result differs between machines or sim and real
- designing push-recovery acceptance tests
- a sharp pass/fail cliff appears in a chaotic-regime evaluation
“训练机上 Mac 原协议 (0.8 @8s cmd0.2) 3/3 全活 … 与 Mac 的 +0.8 0/3 直接矛盾。定性: 单次冲量恢复是二值混沌量, 跨机浮点发散可翻结局;统计门 (20-seed 八门) 跨机吻合的先例不能外推到单点恢复测试。协议改判: 抗推评测多相位 (push 时刻 8/10/12s) + ≥10 seed + 只做同机纵向比”
train/README.md § s2e 支线终章 (跨机混沌课文) The standing-pose reward had been pulling toward the narrow stance the whole line was fighting - a zero-training kinematic audit of the target vector found it, after first auditing the wrong nominal
pose-target-geometric-auditBefore training on a pose target, audit it with forward kinematics - is it geometrically consistent (feet flat, intended stance, intended width) and is it the frame you think it is (action nominal vs standing default)? A posture term's target may itself be the attractor you are fighting.
Symptom
Several rungs aimed at widening the stance failed; the stance stayed narrow as if something kept pulling it back.
Context
The audit (MuJoCo forward kinematics, no training): every 5 deg of hip roll widens the stance ~5.5 cm (0.271 m at 5 deg, 0.383 m at 15 deg); at 47 deg of hip yaw a wide stance cannot be flat-footed (residual foot tilt ~0.7 x hip roll), which explained the stalled rungs. The first report also said the stand_pose nominal (hip roll 25, knee 60) has a 63 deg residual foot tilt - but that was the action frame's nominal (the limit-midpoint squat), which stand_pose never used, despite a docstring warning not to mix them. stand_pose's real target was DEFAULT_JOINT_POS: the contract's all-zero pose, legs parallel, ~0.22 m apart.
Change
The disease statement was corrected in writing: the narrow stance was not an accidental by-product of proxy traps but the target stand_pose had been actively rewarding. A stored "narrow the stance" knife was marked toxic. V2.8 moved the target to a flat 15-deg stance (sigma 3 -> 1.5, flat_feet margin 5 -> 20 deg).
Outcome
V2.8 still failed in-lineage (stance unchanged, feet nearly overlapping, mu 0.4 transfer 2%) - see stance-decided-by-get-up-path - and the from-scratch V3.1 removed both roll joints from stand_pose and put width into a task-space term, which is what finally produced a 0.355 m flat stance. The 63 deg finding was kept as a warning: an action nominal used as a standing target would be a ready-made pit.
Mechanism
A posture term with a sharp kernel around the wrong target is an active attractor; every other term fighting it pays twice.
Applies when
- a posture keeps returning despite penalties against it
- a reward uses a default or nominal pose as its target
- the contract has more than one "nominal" (action frame vs standing pose)
“上文"stand_pose 的 nominal (hip25/knee60) 残倾 63°"**审计错了对象**:那是 **动作参考系 nominal**(限位中点蹲),stand_pose 从未指向它(函数 docstring 原文即警告"两者别混",还是混了 —— 记档)。 … **修正后的病根陈述:窄站距不是代理陷阱的意外副产物,而是 stand_pose 一直在主动奖励的目标本身**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 勘误(实现时抓到):审计混了两个 nominal —— 真病根比误诊的更直白 Every acceptance run records video of the very rollout that produced the numbers - the seated basin, edge-standing feet, tangled legs and the narrow stance were all seen on video before, or instead of, a metric catching them
video-as-acceptance-recordMake video a default output of every acceptance run, rendered from the same rollout the metrics come from (fixed views including the feet), and watch it - posture failures are visible before any gate row exists for them; never let video replace or override the numeric gate.
Symptom
Posture failures in the recovery line kept arriving as things the numbers had no row for: a kneeling W-sit, feet standing on their outer edges, crossed legs, a stance too narrow to hold on hardware.
Context
From 2026-08-09 (user decision) accept_recovery renders offscreen by default, following one env for the whole episode and archiving the clip. The MuJoCo gate's --video renders three views (side, front, feet) of the same rollout the metrics come from, with all plant modelling (delay, push); the older replay-based renderer produced an independent trajectory without delay and was not used for acceptance. Video never gates: if rendering breaks, --no-video keeps the numeric gate running.
Change
Video as a default acceptance artifact, named per policy, category and view, reviewed by the user.
Outcome
The R0.1 prone clip showed the same kneel-sit as R0; the R3.1 failure clip showed the crossed legs; the v2_5 feet view showed edge standing and led to V2.6; the v2_6 videos led the user to order a real-robot A/B between v2_5b and v2_6. Once, the recorder did not start (P1c final acceptance) and the visual material had to be produced separately.
Mechanism
Metrics exist only for failure modes someone anticipated; video shows the unanticipated ones, and rendering the metric rollout itself guarantees the picture and the numbers describe the same episode.
Applies when
- setting up an acceptance pipeline for posture-sensitive skills
- numbers pass but a human reviewer is uneasy
- sim videos are rendered by a separate replay tool
“**验收存视频(用户定 2026-08-09)**:`accept_recovery.py` 默认开 Isaac 离屏 渲染,跟拍一个 env 的整局并归档 … 跪坐这类盆地在数字表出现前肉眼先看见,视频是验收的 定性存档,数字门不受它影响”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §5 R0 验收门(预注册)验收存视频(用户定 2026-08-09) Size each joint's action authority to its measured working range - lock a channel to zero only when its job is provably elsewhere
per-joint-action-scale-lockdownSet per-joint action scales from measured target ranges and momentum decompositions: full authority for working channels, working-range authority for balance channels, zero for channels whose contribution is measured negligible - and prefer this structural quieting over perpetual reward penalties, keeping the contract dimensions intact.
Symptom
"Walks crooked" - roll and yaw channels wandered; hip_yaw peak-to-peak reached 22.3 deg in v11 while contributing essentially nothing to locomotion; a uniform action scale of 0.5 gave every joint the same authority regardless of its actual job.
Context
The v12 design replaced the scalar action scale with per-joint scales justified by measurements: pitch-class 0.5 (the gait's entire working channel - untouched); hip_yaw 0.0 - lossless because it carries only 1.4% of yaw momentum (v6 decomposition) and straight-line targets use merely +/-0.006-0.03 ("乱动纯属浪费"); roll 0.2 - NOT zero, because lateral balance and weight transfer are roll's unique job (locking it would degenerate into foot-edge rocking, "比现在更歪"), and 0.2 covers the measured working range +/-0.06-0.19 while "0.5 的另一半全是 '歪歪扭扭'的来源". Costs were accepted consciously: turning demoted to an observation row with a fallback (yaw 0 -> 0.1 in v13). Contract preserved: the 12-dim action interface unchanged, yaw values simply neutralized. The structural lockdown also RETIRED the reward-side hip_yaw_quiet penalty - "A 的 yaw=0 结构性取代,不再付奖励塑形成本".
Change
action_scale_joint pitch 0.5 / roll 0.2 / yaw 0.0 wired through robot.yaml -> policy_io (verified: all-ones action gives hip_yaw target exactly 0) -> Isaac action term, guarded by the contract checker ("它就是抓这种双侧不一致的").
Outcome
Designed and verified on the shared side before the lineage freeze; stands as the pattern for authority sizing: structure replaces reward shaping wherever a channel should simply not act.
Mechanism
Action scale is a per-channel authority budget; uniform budgets give noise channels the same voice as working channels, and reward-side quieting then pays a permanent shaping tax for what a zero scale provides for free. But zeroing is only lossless when decomposition proves the channel's contribution negligible AND no unique function (balance) lives there.
Conflicts
Wired and verified on the config/deploy side but never trained - the 2026-08-05 reset suspended v12 before the Isaac-side run.
Applies when
- some joints wander without contributing to the task
- a quieting penalty (deviation/L1) taxes every step forever
- deciding action-space authority for a new task or robot
“yaw=0 是无损的:实测它只贡献 1.4% 偏航动量、直行目标只 ±0.006~0.03,乱动纯属浪费。… roll 不能为 0:横向平衡/重心换脚是它的独有职责,锁死会退化成脚缘摇摆(比现在更歪)。0.2 的依据:各代实测 roll 目标只用 ±0.06~0.19,0.5 的另一半全是"歪歪扭扭"的来源。”
train/WALK_V12_SPEC.md § 3. A —— 逐关节动作幅度(用户"只动 pitch"的安全版) Audit rewards by realized contribution (weight x achieved value) - a weight of 2.0 was really paying 0.04
realized-contribution-auditEvaluate a reward table by each term's realized per-step contribution under the current policy, never by its weight column; if a term's realized value is ~0, escalating its weight is a no-op - change the term's structure instead.
Symptom
Foot dragging persisted through repeated weight escalation: feet_air_time had been raised 0.25 -> 1.0 -> 2.0 across versions with no behavioral change, and the training-side comment even recorded the fact ("几乎没有 单支撑相, 是拖着脚蹭") without the fix changing form.
Context
Computing realized per-term contributions in the trained state exposed the economy: feet_air_time contributed weight 2.0 x achieved 0.019 = 0.038 per step, against tracking's +1.20 - lifting the leg earned 3% of what tracking earned, so dragging was the rational optimum no matter the weight escalation. The same table acquitted the energy penalties (total -0.30 negative vs +1.74 positive) that a naive read of weights (-5.0 orientation!) would have blamed.
Change
Fix redirected from "raise the weight again" to "add a term whose realized contribution changes the optimum": a clearance penalty sized so its realized magnitude (~0.018/foot when dragging) is comparable to feet_air_time's, enough to flip the optimum without drowning tracking.
Outcome
With the term economy corrected (plus posture/range fixes), swing height reached 34 mm and tracking 87% by v6; weight escalation of the old term was abandoned.
Mechanism
A reward weight is only a multiplier on whatever the policy currently achieves on that term; when the achieved value is near zero (behavior absent), escalating the weight multiplies near-zero. Optimizer behavior is governed by realized per-step magnitudes, so audits must be conducted in that currency.
Applies when
- a behavior persists despite repeated weight increases
- auditing whether penalties are "too strong" or rewards "too weak"
- sizing a new reward term against existing ones
“把 feet_air_time 权重从 0.25 → 1.0 → 2.0 一路加,但没有加高度项。量级算下来:feet_air_time 权重 2.0 × 实得 0.019 = 0.038,而跟踪奖励是 1.2。抬腿的边际收益只有跟踪的 3%,拖地当然是最优解。… 正项 +1.74,负项 −0.30。能量惩罚不是瓶颈,抬腿没收益才是。”
train/WALK_DIAGNOSIS.md § 决定性证据(#1) / 各项奖励的实际量级 Never referee a suspect metric with another metric from the same code - they can share the disease
independent-referee-for-metric-disputesTo adjudicate a disputed measurement, compute the quantity by an independent method from raw state; never accept a sibling column from the same pipeline as the tiebreaker.
Symptom
A triple reversal on one question: sidewalk sign diagnosis (correct) was retracted using a second metric from the same script, then the retraction itself had to be retracted when that second metric turned out to be the buggy one - two opposite-direction errors on the same problem in one day, both written into the execution sheet.
Context
The probe's net-displacement metric suggested the sidewalk reference sign was inverted. Worried about yaw-drift pollution of net displacement, the author checked the same table's body-frame vy_mean column (~0.003 everywhere, 20-50x smaller) and retracted the sign diagnosis. But vy_mean came from mj_objectVelocity, which was silently reporting vertical velocity due to a frame bug - the "referee" was the diseased measurement. Re-measured with a truly independent computation (xmat.T @ qvel, world trajectory), the original diagnosis was confirmed: saw -0.5 gave vy +0.058/-0.130 (76%/106%), consistent with the net-displacement values all along (yaw pollution was real but only 10-21 deg, nowhere near reversal-sized).
Change
Lesson written twice, verbatim, as a hard rule: when questioning a measurement, the referee must be an independent algorithm (different code path, different physical derivation), e.g. rotate qvel by the body matrix directly, or inspect the raw world trajectory.
Outcome
With the independent referee in place the frame bug was confirmed, fixed, and the whole C4 line re-scored - revealing sidewalk had been working (see body-frame-velocity-api-audit).
Mechanism
Metrics sharing a code path (or an upstream API) share failure modes; agreement between them is evidence about the code, not the world. Only a measurement with an independent derivation can break the tie, because its errors are uncorrelated with the suspect's.
Applies when
- two metrics of the same quantity disagree
- about to retract a conclusion based on a second readout
- auditing evaluation code after a surprising result
“我用一个坏指标去质疑一个好指标,并把撤回写进了执行单。教训(写死):质疑一个测量时,不能用同一份代码里的另一个测量当裁判 —— 它们可能同源同病。裁判必须是独立算法(这次的裁判应该一开始就是 xmat.T · qvel[:3],或直接看世界轨迹)。”
train/C_LADDER_RUN.md § 3m. 二 我今天犯了两个方向相反的错 / 3n. 五 元教训 A training-side rate limit anchored on the last commanded target is an integrator inside the balance loop - two unrelated lineages converged to the same 34-43% re-fall rate, a soft penalty could not fix it, and the bandwidth arithmetic said safety and standing could not coexist
slew-anchor-is-an-integratorIf you constrain actions in training, anchor the constraint on the measured state, not on the previous command - a limiter with memory adds lag inside the balance loop; and when two different lineages converge to the same failure rate, treat the cause as structural and stop adding soft penalties.
Symptom
With the rate limit moved into training (V1), policies either could not stand up or stood up and kept falling again: the stand oscillated, fell and climbed back, 34-43% of the time.
Context
V1.0-B (user decision: explore new postures from scratch, hard constraint in training): target <- prev + clip(target - prev, +/-S*dt) at the TIGHT rates, anchored on the last issued target like the bridge. Four runs: v1_0 from scratch 0.8% - every category righted under the limit, then knelt (the limit also damped the exploration that had escaped the seated basin in V0); v1_0c continued from R3.1 99.8% get-up but 36-43% re-fall and knee jitter 0.604 rad/s; v1_0p with a pull-assist curriculum (56 N -> 0, fully withdrawn) 39.1% unassisted, re-fall 34-42%; v1_0w with a windup-gap penalty 25.4%, re-fall 18-43%. Removing the limiter from v1_0p gave 0.0%: the policy had co-adapted with it.
Change
The rate-limit route was declared dead after four runs and the line was re-rooted on a beta-anchored action space (V2, user approval required).
Outcome
V2.0's first acceptance at full authority already showed re-falls of 0-1% (V1: 34-43%), the structural bet paying off before any tuning.
Mechanism
Anchored on the previous command, a saturated policy becomes a rate controller - one more integrator in the loop - and active balance through that lag oscillates, while a kneeling sit needs no active control and is stable. The arithmetic: kp 30 needs 0.4 rad of error for 12 N*m; at 4 rad/s that takes 0.1 s, half the pendulum time constant sqrt(0.38/9.8) ~ 0.2 s; keeping standing bandwidth needs S of at least ~8 rad/s, within 20% of the 10 rad/s velocity limit - no bound at all. Anchoring on the measured angle (q + beta*a) makes the full kp*beta authority available in one step, with no memory, and caps the impact at the same time.
Applies when
- adding rate limits, slew limits or target filters to a policy's action path
- a policy stands but oscillates and re-falls after a constraint was added
- different lineages or curricula land on the same failure signature
“v1_0p(拉力课程):会站(prone 79.9%),再摔 34~42% —— **两条完全不同 血统、不同学习路径,收敛到同一失败率**。 … 动作饱和时它退化为 速率控制 = 环内多一个积分器;主动站姿平衡穿过该滞后必振荡(v1_0 的跪坐不需 主动控制,所以稳)。对照:**HoST 的 β 锚在当前实测 q,无记忆、无积分器**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §31 结构病定案:slew 的目标锚 = 控制环里的积分器 Add a termination that makes the degenerate strategy fatal - no height cut-off meant crouch-shuffling could live forever
termination-closes-degenerate-basinFor each known degenerate strategy, check whether the termination set makes it fatal; if the robot can live indefinitely inside the degenerate posture, add a termination just past the intended operating envelope rather than escalating penalties.
Symptom
Crouched foot-dragging survived indefinitely because the termination set contained only bad_orientation (40 deg) and base contact - there was no height termination at all, so a deep squat was a viable long-term strategy.
Context
The hypothesis audit found the missing termination (hypothesis 7); the cross-check against published configs found the field practice: Booster terminates at 0.45 m (38% of body height) and the research warning is that the termination height must not be so low that crouching survives it. The proposed value: 0.32 m, just below the walk crouch base height 0.3739 - a deep squat terminates immediately, "断掉蹲着蹭的活路" (cutting off the crouch-shuffle's livelihood).
Change
Add height termination at 0.32 m as a second-priority item of the walk fix package, alongside restoring base_height_l2 to -10.
Outcome
Entered the v5/v6 fix package under which the crouch-shuffle optimum disappeared (34 mm clearance, 87% tracking by v6).
Mechanism
Termination conditions define which strategies exist at all: a reward penalty prices a behavior, but a termination deletes its future returns entirely. Degenerate basins that are merely penalized can remain optimal under enough tracking pressure; a termination placed between the degenerate posture and the intended one makes the basin unreachable as a steady state.
Applies when
- a degenerate but stable behavior persists across reward tunings
- auditing termination conditions for a locomotion task
- a policy exploits the gap between penalized and terminated states
“加终止高度:研究第 6 条"终止高度不能低到让蹲着也能活"。我们完全没有高度终止。建议 0.32 m(略低于 walk 蹲姿基座高 0.3739,深蹲即终止)。… 加终止高度 0.32 m(深蹲即终止,断掉蹲着蹭的活路)”
train/WALK_DIAGNOSIS.md § 修正 ④ / 最终改动清单 第二优先 A constant-value plant rung passed every binary gate with record scores - and shipped 60% thinner posture margins that hardware exposed
constant-value-dr-overfits-marginRandomize deployment-critical axes over a narrow band spanning the measured real support - never a single value, never a fictitious tail - and report graded margin quantities (tilt margin) next to binary gates, because saturated gates rank thin-margin and thick-margin policies identically.
Symptom
s2_lag1 (trained at constant 1-frame latency) posted the strongest sim gate sheet in history (20/20 everywhere) yet was unstable on hardware, while s1e (trained across the full 0-3 frame band) was the every-run-stable SOTA at the same power.
Context
The sim autopsy (new --delay-jitter harness modeling the BusWorker's time-varying phase drift): 18 runs across constant and time-varying delays ALL survived - time variation alone does not kill - but the tilt-margin ordering reproduced hardware exactly: s1e 7.7-9.1 deg (thickest) < s2_lag1 10.7-15.0 < s1c 16.2-18.2. Attribution: constant-value training permits precise specialization to that one value; s1e's band diversity forced cross-value robustness - "恒定 1 帧训练 vs s1e 的 0~3 帧全带——分布多样性逼出跨值鲁棒,恒定值允许精确 特化" - so the constant-rung policy's margins were ~60% thinner, fine in sim's clean world, pushed over the line by real-world disturbances. Tool lesson booked: "存活门二值饱和后掩盖裕度差" - binary survival gates saturate and hide margin differences; graded margin columns (tilt-max) belong in the report. The synthesis with the opposite failure (wide tails cause drag-glide): the proposed resolution was a NARROW uniform band (0.02, 0.04) covering exactly the real 1-2 ticks - diversity inside the measured support, no tail, no single point. s1e's root selection later leaned on the same property: its full-band latency training "预装" the delay rungs and delivered "全工况稳定裕度" that survived power derating.
Change
DR-on-an-axis design refined to a three-way distinction: no wide fictitious tails (drag), no single constant values (thin margins), but a narrow band spanning the measured real support; acceptance reports gained graded margin columns alongside binary gates.
Outcome
The tilt-margin column entered the standard report; the s1e root (band-trained) carried the C ladder while the constant-value branch was archived with its three contributions credited.
Mechanism
Robustness margins are shaped by the diversity of the training distribution, not just its support: a point-mass distribution lets the optimizer trade margin for on-point performance, while a band forces solutions that keep margin across the band - and binary survival metrics cannot see the difference until the margin is spent on hardware.
Conflicts
The narrow-band (0.02,0.04) resolution was a pending recommendation ("裁决建议(待用户)") at the time of writing; the lineage instead moved root to s1e whose full-band training predated the staged ladder - the deterministic-staging card and this card record the two failure modes the final design must avoid simultaneously.
Applies when
- a rung trained at a fixed plant value aces sim but wobbles on hardware
- binary acceptance gates are all saturated across candidates
- choosing between constant, banded, and wide DR on one axis
“18 跑全活,时变性单独不足以击杀;但 tilt_max 裕度排序完整复现真机:s1e 7.7~9.1°(最厚)< s2_lag1 10.7~15.0 … 恒定 1 帧训练 vs s1e 的 0~3 帧全带——分布多样性逼出跨值鲁棒,恒定值允许精确特化 … 存活门二值饱和后掩盖裕度差(s2_lag1 sim 门 20/20 史上最强却真机不稳)”
train/README.md § s2_lag1 真机不稳 × s1e 稳的 sim 对拍(2026-08-07,时变延迟实验) The trainer read a stale USD after the URDF mass update - regenerate derived assets and gate on an automated equality instrument
derived-asset-staleness-checkFor every derived plant artifact (USD from URDF, generated value files), pair the generation step with an automated source-vs-derived equality instrument, prove the instrument can fail, gate training on its PASS, and re-run physical audits after every regeneration.
Symptom
Measured link masses had been committed to URDF/MJCF (total 9.58 -> 9.792 kg, weighed values), but Isaac reads the derived USD asset - which still carried the old masses: a silent 2.2% mass fork between the training plant and the evaluation plant.
Context
The v12 checklist made USD regeneration a hard precondition ("硬性 前置") and, crucially, backed it with an instrument: check_usd_mass.py compares USD vs URDF per-link mass AND inertia trace, validated by showing it FAILs the stale asset naming 9 offending links, then PASSes after re-conversion (13/13 links consistent, total 9.7920). The self-collision filter audit was re-run after the regeneration (three poses, 0.00 N) because a regenerated asset invalidates physical audits done on the old one. This milestone was also where the three plants first aligned: "三边 plant 首次对齐(armature+摩擦+ 实称质量)就在这一代".
Change
convert_urdf re-run on the training machine, regenerated asset committed, check_usd_mass.py PASS required as an acceptance gate for the generation; dependent audits repeated post-regeneration.
Outcome
The 2.2% plant fork was closed before it could distort a generation's acceptance numbers; the staleness class of bug now has a permanent detector instead of a memory.
Mechanism
Source-of-truth edits do not propagate to derived binary assets by themselves; any consumer reading the derivative silently trains or evaluates on the old plant. An automated equality check between source and derivative - proven able to fail - turns an invisible staleness into a red gate, and regeneration invalidates every audit performed on the old artifact.
Applies when
- editing masses/inertia/geometry in URDF or MJCF sources
- a trainer or evaluator consumes converted/derived assets
- plant numbers differ between simulators for no visible reason
“25ba997 把连杆质量更新为实称值(总重 9.58→9.792 kg,URDF/MJCF 已改),但 Isaac 读的是 train/assets/laika_v2.usd —— 仍是旧质量。… 否则 Isaac(9.58)与 MuJoCo(9.79)质量分叉 2.2%,v12 验收数字失真。验收门:python tools/check_usd_mass.py 必须 PASS … 对旧资产实测 FAIL/9 连杆点名,仪器已验证”
train/WALK_V12_SPEC.md § 7. 核查单 (⚠️ 先重转 USD) When the eval proxy has a known systematic bias, gate on within-proxy differences, not absolutes
relative-metrics-survive-proxy-biasWhere the evaluator is known-biased, design gates as within-evaluator contrasts (left vs right, A vs B, pre vs post) that cancel common-mode error; reserve absolute thresholds for quantities whose proxy calibration has been checked.
Symptom
MuJoCo systematically overestimated yaw turn gain (1.58/2.45 where Isaac measured 1.04/1.02), and friction alignment recovered only part of the gap (mu 1.0 -> 0.6 pulled it to 1.43/2.27) - so any absolute turn-gain acceptance threshold in MuJoCo would be judging the proxy's bias, not the policy.
Context
The acceptance criterion was rewritten to use only the left/right difference of the turn gain (standard: <20%): both directions pass through the same biased proxy, so the bias largely cancels in the difference while the policy's chirality - the thing being gated - survives. The absolute-gain row was dropped: "转向只判左右差,不判绝对值". A parallel task was still opened to align MuJoCo contact parameters to Isaac's (mu, restitution), since drift proved highly friction-sensitive (straight-line drift -65 deg at mu 1.0 vs +1 deg at mu 0.4) - bias reduction and bias-robust metrics proceeded together.
Change
Gate metric changed from absolute turn gain to left-right gain difference; proxy-alignment work scheduled separately rather than blocking acceptance.
Outcome
Turn acceptance became meaningful across proxy versions (v5 70% -> v6 19% difference measured the real improvement) while the absolute bias question was pursued without holding the ladder hostage.
Mechanism
A systematic multiplicative or additive proxy bias applies to both arms of a mirrored measurement; differencing (or ratioing) mirrored conditions cancels the common-mode bias to first order, leaving the asymmetry signal. Metrics built this way remain valid while the proxy is imperfect - which it always is somewhere.
Applies when
- a sim proxy disagrees with the trainer or hardware in absolute terms
- writing acceptance thresholds for direction-paired skills
- proxy calibration work would otherwise block a ladder
“转向只判左右差,不判绝对值 —— MuJoCo 的偏航增益系统性高估(实测 1.58/2.45 vs Isaac 1.04/1.02),摩擦只能解释一部分(μ 1.0→0.6 仅拉到 1.43/2.27)。差值是相对量。”
train/WALK_V5_SPEC.md § 6. 验收 The get-up kept getting faster because standing earlier paid more every step - lowering torque authority barely slowed it, and only zeroing the standing income for the first 3 s moved the pace into the design band
per-step-income-drives-speed-time-gateWhen a skill is too fast, find the term that pays for finishing early and gate that income by time; keep the "get into position" term ungated so the policy does not learn to wait, use a ramp instead of a cliff, and confirm with a paired same-level experiment that the drift is motivational before changing it.
Symptom
The user judged the get-up too fast (Isaac medians about 0.7-1.6 s) and suspected path dependence: the policy seemed to get faster the longer it trained.
Context
Lowering the beta authority 0.40 -> 0.30 cut impact but moved supine only 1.70 -> 2.00 s: coordination-limited, not torque-limited. A paired experiment inside one beta level (checkpoint 15,600 vs 18,499, +2,900 iterations, same ruler) measured the drift: get-up medians -7 to -10%. The spec concluded the motive, not the path, was the cause - per-step standing income pays for every early step, and any lineage (even one from scratch) races toward the fastest solution inside its constraints.
Change
V2.5: the standing income (base_height, stand_pose, still, feet_on_ground) multiplied by w(t) = clamp(t/3 s, 0, 1); upright deliberately NOT gated, so righting and sitting up early still pay and the policy is not taught to lie flat and wait; a ramp, not a step. V2.5b: zero before t0 = 3 s, then a 1 s ramp.
Outcome
V2.5: Isaac 100%, get-up +17-43% slower, MuJoCo 100/100/98/100% (the best cross-simulator reading yet), still short of the 3.5-4.5 s design band - a linear ramp only discounts early income. V2.5b: MuJoCo supine 2.04 -> 4.10 s and prone 3.18 -> 4.04 s, inside the band; the Isaac pace barely moved (a lineage habit on a gradient-free plateau). Later the zero gate proved harmful when trained from scratch (curriculum-history-is-part-of-the-product) and in the V3.1 lineage (time-gate-vs-wide-stance-retire-the-fix).
Mechanism
Constraints on authority or velocity change how the fastest solution looks; the time structure of the task income decides how fast the fastest solution is.
Applies when
- a policy is faster or more aggressive than wanted and constraints do not slow it
- progress-style rewards pay every step spent at the goal
- performance drifts faster with more training at fixed settings
“**V2.5 机制(唯一)**:站立收入(base_height/stand_pose/still/feet_on_ground) 乘时间斜坡 w(t)=clamp(t/T_gate,0,1),T_gate=3.0 s;**upright 刻意不门控** (翻正/坐直早期照常拿钱,防"躺平等门开" … **用户假设量化 证实:逐步计酬动机在 β 包络内持续压缩时间,约束挡不住动机 —— V2.5 动机层 修法为正解。**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §39 V2.5 预注册 / §39 补 配对实验 Two simulators disagreed 1.8x on one policy's torque demand - four explanations were eliminated with numbers, the surviving suspect was never tested, and neither reading was allowed to cancel the other
torque-disagreement-between-simulators-unresolvedWhen two simulators disagree on a safety-relevant quantity, eliminate the measurement explanations one at a time with numbers, name the survivor as a hypothesis, keep the pessimistic reading binding, and run the direct test (replay one action sequence open-loop through both plants) before the quantity is used to pass a gate for hardware.
Symptom
For R3.0, Isaac read hip_pitch torque demand at 75-80% of the limit (PASS); MuJoCo read the same policy at 148-152% (1.5x over the limit).
Context
Candidates were eliminated one by one: PD gains (identical, RS06 kp 30 / kd 1.5), the settle window (both skip the first 0.5 s), the statistic (worse-of-two-legs vs per-joint - explains ~15%), self-collision (the no-contact subset still read 152%), sampling (500 Hz peak vs 50 Hz sample - ~9%). About 1.8x remained. The prime suspect - Isaac's implicit actuator (PD solved inside the PhysX integration, chosen to mimic kHz firmware PD) against MuJoCo's 500 Hz explicit PD - was named but untested.
Change
The Isaac PASS was recorded as valid for the Isaac plant only; the prescribed decider was an open-loop replay of one action sequence through both plants, compared step by step, needing no training. It was upgraded to a precondition of R3.1.
Outcome
R3.1 ran in parallel with, not after, the replay, and the replay never appears as done. By §40 it was "the oldest open account" on the line, to be fed by real-robot logs - which the spec also never records. Meanwhile the 50 Hz sample under-read the peak by 34% as smoothing narrowed the spikes, so Isaac-side readings grew more optimistic exactly when the comparison mattered more.
Mechanism
Each measurement artifact explained a slice of the gap; what remained is a plant difference, and an untested plant hypothesis cannot license discarding the pessimistic simulator on a safety-relevant quantity.
Conflicts
The spec prescribes the open-loop replay (§23), upgrades it to an R3.1 precondition, then records that R3.1 ran in parallel without it (§24), and lists it as the oldest open account before hardware (§40); nothing through §50 (2026-08-14) records a result. Every later "torque gate PASS" on this line is an Isaac-plant reading plus a MuJoCo check, never a reconciled one.
Applies when
- Isaac and MuJoCo (or sim and hardware) report different torques, contacts or slips
- implicit vs explicit actuator models are in play
- a gate passes on one simulator only
“**同一个策略,一侧判 PASS 一侧判超限 1.5 倍。** … 口径项全部扣掉后仍剩 **~1.8×** 没有解释。 … 但**这条没有验证,不能拿它当结论去抵消 MuJoCo 的读数**。 … (做法:拿同一条 动作序列在两边开环回放,逐拍比 τ —— 不需要重训)”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §23 MuJoCo 复核门:成功率更稳,但 τ 与 Isaac 差 ~1.8× 没收敛 The real robot's right-leg kicking was over-trained-delay times loop gain - irreducible pipeline latency is plant, model it fully from day one
pipeline-latency-is-plant-not-drMeasure the end-to-end action pipeline delay and build it into the nominal plant and every acceptance gate from day one; treat power/scale deratings that "fix" oscillation as gain-reduction crutches flagging an unmodeled delay, and expect higher-feedback-gain policies to be MORE delay-fragile.
Symptom
On hardware, s1c/s1d at action scale 1.0 always kicked wildly with the right leg (s1c only ran as SOTA at power 0.8; s1d only at 0.7) - while sim showed nothing under default evaluation.
Context
Sim reproduced the incident item by item once the real pipeline delay was injected: s1d@1.0 with --delay 1 fell at 10.2 s, --delay 2 at 5.2 s; s1c@1.0 stressed (r_hip_roll saturation 5 -> 16%; "右脚" = the policy's chirality makes the right leg its high-gain leg); and the combos that worked on hardware (s1c@0.8+delay2, s1d@0.7+delay2) all survived in sim. Mechanism: the real pipeline is ~1-2 ticks (BusWorker next-cycle pickup + CAN round trip) but S1.1 trained only to 1 tick - "超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定" (delay beyond training x full loop gain = oscillation; the power derating had been buying stability by compressing loop gain). s1d was MORE fragile than s1c because its yaw 3-layer stack had learned higher feedback gain - higher gain, lower delay tolerance. Three changes: latency DR widened to cover reality; acceptance gates and smoke runs moved permanently to --delay 2 ("门必须在真机条件下预测 真机"); and the doctrine written twice-paid: "不可约的管线属性(延迟、 限速)不是'随机化选项',是 plant 本体,第一天就该全额建模" - S1's nominal-then-robust staging falsified by hardware for the second time. The later s1e hardware run at power 1.0 (no kicking, normal force) closed the loop: "0.8 = 旧代拐杖" - the derating had been a crutch for the under-modeled delay, not a real requirement.
Change
Latency modeled as plant from day one of any lineage (measured 1-2 ticks covered, bridge-layer rate limits likewise modeled by default); every gate and smoke evaluation issued under --delay 2.
Outcome
Kicking reproduced, explained, and eliminated in the s1e generation at full scale and full power; the deploy-side crutches (0.7/0.8) retired for the new lineage.
Mechanism
Feedback oscillation onset is a product of loop gain and phase lag; a policy trained below the real delay learns gains that sit past the real stability margin, and any output derating masks it by scaling gain down. Since pipeline delay is deterministic hardware property - not an uncertainty - it belongs in the nominal plant, and every evaluation must include it or the gate predicts a robot that does not exist.
Applies when
- hardware oscillation/kicking that sim only reproduces with added delay
- a policy only runs on hardware at reduced power/scale
- defining what belongs in the nominal plant vs the DR list
“真实链路延迟 ~1~2 拍 … S1.1 只训到 1 拍——超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定。s1d 比 s1c 更脆 = yaw 三层栈学出更高反馈增益,增益越高延迟容忍越低。… 教训入账:S1「先标称后鲁棒」第二次被真机证伪——不可约的管线属性(延迟、限速)不是"随机化选项",是 plant 本体,第一天就该全额建模。”
train/OMNI_V0_SPEC.md § 3. S1.4(真机右脚乱踢事故强制) A frame-history observation under zero DR memorizes the trainer's plant fingerprint - the estimator must see variation to learn estimation
history-obs-needs-plant-variationIf the observation carries history (stacked frames, RNN), keep at least minimal plant variation (gain/latency jitter) on from the first iteration - "nominal first, robust later" is structurally invalid for estimator-bearing contracts.
Symptom
omni_s1 (fresh 215-dim contract with a 5-frame history window, trained with DR fully off): training all green, yet the MuJoCo gate scored 0/20 on all eight doors - falls within 2 s, seven checkpoints, not one transferred.
Context
The history window exists precisely to let the actor implicitly estimate line velocity and actuator dynamics (the actor is denied base_lin_vel by observation honesty). Under a constant plant that implicit estimator has nothing to estimate - it learns the trainer's exact response fingerprint instead, and any other simulator's micro-differences are out-of-distribution: "5 帧窗按设计就是隐式估计器, plant 恒定时它学到 Isaac 精确响应的指纹". The planned "nominal-first-robust-later" staging was declared STRUCTURALLY incompatible with history observations: "估计器要见过变化才学估计, 否则学背诵" (an estimator must see variation to learn estimation, otherwise it learns recitation). Honest confound note kept: this is mixed with "zero DR does not transfer, period" - but both attributions prescribe the same fix, so no control was run.
Change
S1.1: minimum actuator jitter turned on from day one - kp/kd +/-10%, latency 0-1 frame (friction/COM/mass still nominal, no push - those stay for the S2 ladder).
Outcome
Transfer restored: survival 0/20 -> 20/20, speed 19/20, foot distance 20/20 (remaining failures moved to gait quality, a different disease); the staging doctrine was amended - history-carrying contracts never train under a frozen plant.
Mechanism
A recurrent/history channel fits whatever temporal structure minimizes loss; with a deterministic plant the cheapest structure is the plant's own impulse-response signature, yielding features that are simulator-specific rather than physics-general. Plant variation forces the channel to carry state-estimation features that transfer.
Conflicts
Attribution is explicitly confounded with the simpler "zero DR never transfers" reading ("与「零 DR 本身就不迁移」混杂 … 两种归因处方相同, 不做对照") - the source chose not to spend a control run separating them.
Applies when
- adding frame stacking or recurrence to an actor observation
- a nominal-plant policy fails a cross-simulator gate within seconds
- planning DR staging for a new contract
“frame_hist × 零 DR = plant 指纹过拟合——5 帧窗按设计就是隐式估计器, plant 恒定时它学到 Isaac 精确响应的指纹, MuJoCo 的微小差异即 OOD, 2 s 内摔, 七个 checkpoint 无一迁移。「先标称后鲁棒」的分段与历史观测结构性冲突:估计器要见过变化才学估计,否则学背诵。”
train/OMNI_V0_SPEC.md § 3. S1.1 修订记录 ①