Skip to content

Training Coach

The coach reads a training run and returns a diagnosis, in-whitelist proposals and an experiment plan — and every claim cites one of the cards below. This is that corpus: 22 methodology rules distilled from a real sim2real programme, and 171 episode cards behind them, each with the sentence in the war history it came from.

Doctrine

A report may cite any of these as doctrine-N.

  1. doctrine-1Contract freeze and fingerprint discipline

    The policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.

    Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).

    Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.

    contract-freeze-and-checkerlegacy-profile-pinningderived-asset-staleness-checkgain-profile-belongs-in-the-stamprecovery-two-policies-and-a-state-machinewalk-recovery-fsm-handoff

  2. doctrine-2Attribution by resolved training params - never eval-override knobs

    Capability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.

    Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).

    Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.

    kd-bandwidth-mu-law-attributiondeploy-knob-attribution-before-retrainingcycle-time-override-is-oodresolved-config-is-source-of-truth

  3. doctrine-3PASS gates become constraints; FAIL gates become objectives

    Once a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.

    Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.

    Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".

    fixed-acceptance-matrix-per-rungpreregistered-stop-criteria-per-rung

  4. doctrine-4One variable per ladder rung - counted against what the checkpoint saw

    A rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.

    Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.

    Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.

    resume-state-dr-auditorthogonal-batch-with-ablation-orderfill-the-missing-factorial-cell

  5. doctrine-5Pre-register risks, readings, and stop criteria before the ladder

    Before a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.

    Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).

    Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.

    preregister-risks-and-fork-readingspreregistered-stop-criteria-per-rungpreregistered-real-expectationsfeasibility-accounts-lock-design-point

  6. doctrine-6Plant parameters are measured, never invented

    Every plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.

    Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).

    Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.

    friction-measured-not-guessedarmature-n2-rotor-inertiatorque-limit-shape-by-measured-peakslatency-lerp-reverse-extrapolationfeasibility-accounts-lock-design-pointsim-api-friction-columnsget-up-feasibility-accounts-before-trainingsingle-support-gain-authority-probe

  7. doctrine-7Sim2sim gate before sim2real - under deployment conditions

    Every checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).

    Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).

    Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.

    sim2sim-gate-before-sim2realeval-plant-honesty-contact-paramspipeline-latency-is-plant-not-drbody-frame-velocity-api-audittorque-penalty-bought-by-leg-bracingtorque-disagreement-between-simulators-unresolved

  8. doctrine-8Observation honesty - the actor's inputs are a hardware contract

    The actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.

    Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).

    Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.

    observation-honesty-critic-onlyhistory-obs-needs-plant-variationreward-observability-limitdeploy-heading-loop-and-align-training

  9. doctrine-9Reward economics are audited in realized currency

    Reward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.

    Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).

    Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.

    realized-contribution-auditreward-cost-of-ignoring-auditgate-new-reward-terms-by-commandignore-floor-diagnosiscalibrate-threshold-between-healthy-and-sickcalibration-threshold-with-withdrawal-clauseinert-reward-term-auditseated-basin-dead-exp-kerneltail-torque-needs-hinge-on-computed-demand

  10. doctrine-10The zero-cost option must be the desired behavior

    For every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.

    Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).

    Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.

    penalize-the-slip-not-the-jointsaturation-cheating-zero-rate-costmoving-gate-42x-stand-taxtermination-closes-degenerate-basinpenalty-gate-is-an-escape-hatchsoft-limit-penalty-charges-nominal-poseunpriced-foot-attitude-is-a-free-variableenumerate-cheapest-cheats-before-trainingbinary-band-reward-fake-touchdown

  11. doctrine-11Measurement discipline: independent referees, signs, distributions

    A disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.

    Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).

    Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.

    independent-referee-for-metric-disputesbody-frame-velocity-api-auditsame-sign-response-is-yaw-biasmedian-hides-bimodal-distributionratio-metrics-need-absolute-checkheading-integral-not-body-ratesame-distribution-reward-comparisonmultiseed-sign-test-for-driftsingle-impulse-recovery-is-chaotic

  12. doctrine-12The deployment pipeline is plant

    Irreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.

    Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).

    Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.

    pipeline-latency-is-plant-not-drpower-scale-hurts-nonforward-axesdeploy-scaling-not-training-equivalentteleop-command-band-per-axislatency-dr-covers-measured-pipelinedeploy-rate-limiter-windupslew-anchor-is-an-integratorbeta-anchored-action-targetpower-derating-cuts-full-range-contract

  13. doctrine-13DR budget is finite; its distribution is the measured support

    Robustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.

    Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).

    Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.

    push-dr-conditional-budget-conservationdr-tail-plant-continuationconstant-value-dr-overfits-margintask-shaping-before-plant-hardeningcom-randomization-forces-leg-spreadcom-dr-rollback-on-symptomthin-dr-judged-by-channel-coveragefriction-priority-re-measured-after-plant-change

  14. doctrine-14Gates measure what hardware feels: posture, margins, stripped assists

    Acceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.

    Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).

    Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.

    task-metrics-vs-posture-metricsstand-gate-posture-not-survivalconstant-value-dr-overfits-marginsuspension-probe-removes-free-stabilizerhip-roll-sum-predicts-lateral-driftchirality-scored-separatelylow-speed-commands-reward-draggingend-state-confusion-matrixfrozen-acceptance-distribution-and-pinned-seedvideo-as-acceptance-recordepisode-length-bounds-what-a-gate-seesremoved-wall-returns-on-hardware

  15. doctrine-15Fork and root selection: recoverability, maturity, frozen rewards

    Choose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.

    Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).

    Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.

    fork-root-recoverable-shortfallroot-maturity-vs-product-qualityfine-tune-reward-change-falsified

  16. doctrine-16Curricula: verified engagement, lineage counters, disease-phase gating

    Automatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.

    Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).

    Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.

    auto-curriculum-engagement-checkcurriculum-counter-lineage-stepsgate-penalties-to-the-disease-phaseaggregate-metrics-mask-subgroup-failurebucket-share-is-not-a-gradient-levercurriculum-criterion-conditioned-on-lagging-categoryper-step-income-drives-speed-time-gatetime-gate-vs-wide-stance-retire-the-fix

  17. doctrine-17Probe before training: feasibility first, hypotheses in tables

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.

    Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).

    Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).

    open-loop-probe-before-reward-tuninghypothesis-table-code-auditperiod-doubling-evidence-racesuspended-test-isolates-actuator-authorityconfiguration-probe-wall-not-slopeprone-dead-end-is-foot-placementdof-vel-penalty-is-not-a-pacing-knobamplitude-cut-falsified-yaw-fix

  18. doctrine-18External advice is recomputed locally; values transfer as ratios

    Every external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.

    Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).

    Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.

    external-advice-audit-against-own-arithmetictransfer-ratios-not-absoluteslatency-dr-covers-measured-pipelinereference-structure-fk-amplitude-divisioncycle-average-tracking-for-gait-quantitiesadvisor-paraphrase-vs-paper

  19. doctrine-19Hardware sessions are scripted experiments, not tuning sessions

    Real-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.

    Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).

    Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.

    risk-ordered-real-deploymentbattery-bracketed-real-abknow-zero-command-behaviorpush-test-chirality-protocolno-field-tuning-protocolreversible-single-variable-field-experimentssim-veto-needs-real-confirmationfirst-real-get-up-violent-stage-one-policystaged-hang-mat-floor-for-get-uppower-cycle-preflighttwo-machine-config-disciplinefall-guard-becomes-a-statehardware-log-is-the-attribution-input

  20. doctrine-20Close questions in writing; restart when the debt is structural

    Audited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.

    Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).

    Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.

    frozen-verdicts-semantic-boundariesstale-verdicts-under-old-stackzero-offset-calibration-shifts-envelopeplant-swap-invariants-vs-shiftsfreeze-lineage-fix-structure-restartminimal-reward-table-with-provenancewrite-hardware-verdicts-back

  21. doctrine-21Name the quantity in the space it lives in

    A goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.

    Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).

    Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.

    joint-space-proxy-for-task-space-quantitycontact-detector-single-signal-lieszero-partial-credit-is-not-an-iteration-problemheading-integral-not-body-rate

  22. doctrine-22Continuation needs a live gradient; a release is chosen by a scan

    Continue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.

    Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).

    Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.

    converged-continuation-is-poisoncheckpoint-choice-is-a-full-gate-scanstance-decided-by-get-up-pathstop-stacking-roll-back-and-auditcurriculum-history-is-part-of-the-productcontinuation-budget-not-from-zeroroot-maturity-vs-product-quality

Experience cards

73 cards matching “curriculum-counter-lineage-steps”.

  • A curriculum ramp keyed to the process step counter re-fires on every resume - and shipped policies never saw the penaltycurriculum-counter-lineage-steps
    Mechanism understoodomnicurriculumcurriculumattribution

    Key every curriculum/ramp schedule to lineage-cumulative progress, not per-process counters; and audit where your shipped checkpoints sit relative to every ramp - a penalty that no product ever experienced is not part of your training.

    Symptom

    vx+0.30 died at a fixed relative time in every resumed run: resume at 500 -> slide at 1100, zero at 1300; resume at 700 -> slide at 1300, zero at 1500 - absolute depths offset by exactly the resume offset, relative timetable identical.

    Context

    ramp_reward_weight (the saturation penalty ramp) read env.common_step_counter, which restarts at 0 for every run including --resume. So start_step=600 meant "600 iters after THIS resume", not "lineage iteration 600". The A/B arm pair was the clean proof: their env.yaml differed only in log_dir, only resume point distinguished them, and the omni CurriculumManager had exactly one active term - nothing else could produce that timetable. Second consequence: every shipped checkpoint (s1e-500 at +500, C2-700 at +200, A800 at +100) was selected before its run's +600, so the saturation penalty weight was 0.000 for every product ever shipped - explaining saturation 33% and raw |action| 1.9 against clip 1.0 (hip_roll in bang-bang), i.e. half the heat budget.

    Change

    Two independent recommendations recorded: (1) make the ramp count lineage steps (add the checkpoint's iteration offset on resume) or pin terminal weights in downstream rungs instead of ramping; (2) give the saturation penalty its own rung - never mixed into a skill-learning rung (that would be two variables again).

    Outcome

    Explained the recurring +600 death of the highest-amplitude command and the persistent actuator saturation of all shipped products with one root cause; honest caveat booked (at +800 the weight is only -0.086, small, but vx+0.30 is the command demanding the largest action amplitude, so it is squeezed first).

    Mechanism

    Resumable training splits "the lineage" from "the process"; any schedule keyed to process-local counters silently re-applies its transient to every descendant run, and any product-selection habit that picks checkpoints early systematically samples the pre-ramp regime - the curriculum exists in the config but never in any shipped policy.

    Applies when

    • resumed/forked training with any scheduled reward or DR ramp
    • a metric dies at a fixed offset after each resume
    • shipped policies show behavior a late-schedule penalty should prevent
    “ramp_reward_weight 读的是 env.common_step_counter,它每个 run 从 0 开始,--resume 也不例外。… 原始 run / 臂B | 500 | 1100 = +600 | 1300 = +800;臂A | 700 | 1300 = +600 | 1500 = +800 … 所有出品其实从没见过饱和罚。… 这解释了 sat_max_pct 33%、raw |a| 最大 1.9(clip 是 1.0)—— hip_roll 一直在 bang-bang,而罚它的那一项权重恒 0。热账的一半在这里。”
    train/C_LADDER_RUN.md § 3g. 系统性问题:saturation_ramp 每次 resume 归零
  • Training the final recipe from scratch in one run - every mechanism the lineage had accumulated - produced 0% and a seated robot; the order in which the lineage acquired those mechanisms was part of why it workedcurriculum-history-is-part-of-the-product
    Observed oncerecoverytraining-runcurriculumprocess

    A recipe that ends a lineage is not a recipe for a from-scratch run: consolidate it as ordered curriculum phases matching how the lineage acquired its mechanisms, check that every curriculum criterion is reachable from the starting policy, and read the run's raw term values, not the total reward, before calling it green.

    Symptom

    V3.0 trained the lineage's whole final recipe from scratch in one 9,000 iteration run - full beta curriculum, prone-conditioned pull assist, friction DR, the 3 s zero gate on standing income, flat_feet - testing the proposition "the product is defined by its configuration, not by its training history". Training looked all green (reward 26.33, episode length 500, 100% time-outs).

    Context

    Read in raw units against v2_6c at the same weights, the green was a seated equilibrium: base_height 0.377 vs 0.640, stand_pose 0.205 vs 0.516, flat_feet 0.0000 (zero because it sits outside its height gate, not because the feet were flat). Acceptance: 0.0% in Isaac at the deployed authority, 0.0% on every MuJoCo friction level, and still 0.0% at the training-time authority (100% seated at 0.222 m, upright and still).

    Change

    The full stdout (316k lines) was read: the beta curriculum's criterion (standing share over 0.35) was met zero times, so beta never left the wide setting and the policy had no experience at the deployed authority; the pull curriculum was stuck on the same criterion. The lineage had escaped the seated basin with immediate income (V2.0-V2.2) and only then added the zero gate to cure rushing (V2.5); from scratch, the zero gate removed the early "stand fast, earn more" gradient needed to escape. Verdict "the curriculum history is part of the product", limited to n = 1. v2_6c stayed the product.

    Outcome

    V3.1 kept the order as explicit phases: P1 from scratch with immediate income (zero gate off) until the curricula advance, P2 adding the zero gate. P1b/P1c escaped the seated basin and passed; P2 was later judged net negative and dropped (time-gate-vs-wide-stance-retire-the-fix).

    Mechanism

    Mechanisms that refine a competent policy (time gates, tight authority) can delete the gradient a naive policy needs, and a curriculum whose advancement criterion the naive policy never meets freezes at its first level.

    Conflicts

    The spec limits the falsification to "this recipe + this curriculum criterion" (n = 1, no seed sweep, no criterion tuning). V3.1's phased run succeeding is consistent with the ordering reading but changed other terms too.

    Applies when

    • consolidating a long lineage of continuation fixes into one clean recipe
    • a from-scratch run with all mechanisms enabled plateaus early
    • curriculum state is not logged or never advances
    “命题:产物由配置定义,而非训练史定义。 … 训练 log(完整 stdout 316k 行)`[beta_anchor]` 仅初始 1 行,**达标 0 次** … V2.0~V2.2 靠**即时计酬**爬出坐姿盆地(§33),站立巩固后 V2.5 才装归零门 治"过快"(§39)。**课程史是产品的一部分。**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §42 V3.0 判决(2026-08-11):从零单 run 全机制 FAIL 于坐姿盆地
  • An auto-curriculum ratchet capped out at iter 248 and never engaged - stage difficulty manually or verify engagementauto-curriculum-engagement-check
    Observed oncewalkcurriculumcurriculumdomain-randomizationprocess

    Prefer manually staged difficulty with gated transitions; if you use an automatic curriculum, instrument its internal state and alarm when it stops engaging - a saturated curriculum is constant DR wearing a curriculum's name.

    Symptom

    A curriculum mechanism intended to grow difficulty adaptively (s1f's ratchet) hit its cap at iteration 248 and never bit again - for 96% of the run its effect was equivalent to constant DR, i.e. the curriculum existed in name only.

    Context

    When external advice suggested graded wz bands (start ±0.15, then ±0.30), the team agreed with grading but explicitly rejected automatic curriculum, citing the s1f episode. The same logic had already been paid for with push grading: ±0.6 failed twice, ±0.3 was feasible - grading matters, but the grade transitions were made by hand at verified checkpoints.

    Change

    Ladder policy: difficulty staged manually, one band per rung, each transition gated by the acceptance battery; automatic ratchets not used unless their engagement is monitored and demonstrated.

    Outcome

    Every C-ladder band change (wz ±0.12-0.25 first, wider later) was an explicit, attributable rung; no silent constant-DR-in-disguise runs recurred.

    Mechanism

    Adaptive curricula couple their own state machine to noisy training metrics; a ratchet that saturates early stops adapting but keeps its name, so the operator believes difficulty is progressing when it is frozen. Manual staging costs more decisions but each decision is observable and reversible.

    Applies when

    • choosing between auto-curriculum and staged bands for a new skill
    • a curriculum's difficulty parameter plateaus early in training
    • post-hoc attribution of what difficulty a lineage actually saw
    “C2 的 wz 分级(±0.15 → ±0.30,不要一上来 ±0.6)—— 与我们付过学费的 push 分级同形(±0.6 两轮 FAIL,±0.3 才可行)。但必须手动分级,不做自动课程 —— s1f 的课程化棘轮 iter 248 封顶未咬合,96% 时长等价常量 DR”
    train/C_LADDER_RUN.md § 1. 采纳 3 条 (C2 的 wz 分级)
  • Push DR helped one lineage and hurt another at the same dose - robustness budget is conserved and gets borrowed, not createdpush-dr-conditional-budget-conservation
    Replicatedomnidr-tuningdomain-randomizationcurriculumattribution

    Before opening a disturbance-DR rung, measure whether the untrained policy already meets the spec; if training it anyway, expect the benefit to be conditional on the lineage's existing DR load, grade the intensity, and audit retained margins - budget spent elsewhere will be borrowed back.

    Symptom

    The push rung's outcome flipped with the lineage: direct +/-0.6 m/s push failed outright on first attempt (base walking collapsed - kd1.2 scan 0/3 from iter 3300, sim2sim self-falls with pushes OFF - no PASS point existed); staged +/-0.3 then gave the narrow-kd single-working-point lineage real gains (push survival 1/5 -> 4/5) while the SAME dose made the dual-working-point balanced-band lineage WORSE (20-seed survival 18 -> 12/20 plus across-the-board push regression).

    Context

    The four-ladder verdict ("四梯定案", s2e/s2f at both intensities) named the pattern: "push DR 收益条件性" - the benefit is conditional on how much robustness budget the lineage has already spent. The law candidate: "DR 总预算守恒, 平衡带鲁棒性从抗扰余量借" - total DR budget is conserved; a lineage already covering a wide plant band pays for push tolerance out of its disturbance margin. Both S2 ladders therefore closed at the friction rung, with the decisive numerator: untrained push tolerance already met the 4-6 N*s requirement, so the rung was not needed at all ("⑥ push 不训(收益条件性,免训 ±0.6 已达 标)"). The same accounting later justified the C-before-S2 ordering ("push/μ 两轮已实证 DR 预算有限且会被重分配") and trimmed the second S2 pass to three rungs.

    Change

    Push removed from the standing ladder; graded intensity retained as the method IF a lineage ever needs push training; "does the untrained policy already meet the disturbance spec" instituted as the first check before opening any disturbance rung.

    Outcome

    Two rungs (push, ground mu) deleted from the second S2 pass on measured grounds; the ladder's real yield was re-stated honestly as precision, not robustness (speed gate 0 -> 20/20, zero-command drift 0.98 -> 0.06 m, but push 159 -> 125/160).

    Mechanism

    A fixed-capacity policy allocates representation and margin across the training distribution; adding a disturbance axis to a lineage that already spans a wide plant family forces reallocation - the new tolerance is bought with existing margins. Lineages with narrow plant coverage have free budget, so the identical DR dose lands as gain. Benefit is a property of (dose x lineage state), never of the dose alone.

    Applies when

    • proposing push/perturbation training on a hardened lineage
    • the same DR rung helped one lineage and hurt another
    • accounting where a ladder's robustness gains actually came from
    “push DR 收益条件性 —— s2e⑥a (单工作点血统 kd 窄带) ±0.3 得抗推 1/5→4/5; s2f⑥ (双工作点平衡带血统) 同档反而 20-seed 存活 18→12/20 且抗推全面倒退。规律候选: DR 总预算守恒, 平衡带鲁棒性从抗扰余量借。两阶梯均以 ⑤ 摩擦级收官 … 抗推 4~6 N·s 免训已达标。”
    train/OMNI_V0_SPEC.md § 4. ⑥ push 四梯定案 (2026-08-07)
  • A time gate that had cured one lineage's rushing made the from-scratch lineage trade away its stance width twice (0.364 -> 0.235 m, 0.355 -> 0.251 m) - its disease was absent there, so the fix was retired and the pre-gate checkpoint shippedtime-gate-vs-wide-stance-retire-the-fix
    Replicatedrecoveryreward-shapingreward-shapingcurriculumfork-selection

    Carry a fix into a new lineage only if its disease is present there; a mechanism that cured one lineage can be net negative in another, and when doubling a term's weight recovers almost nothing, treat the two objectives as structurally in conflict and remove the one whose purpose is gone.

    Symptom

    V3.1's phase 2 (the 3 s zero gate on standing income, continued from P1b) kept 100% success on every friction level and slowed the get-up, but the lateral stance drifted 0.364 -> 0.235 m and hip yaw crept to 57 deg against its 60 deg limit. With the width band's weight doubled (P2c, after P1c) it drifted again, 0.355 -> 0.251 m, below the pre-registered 0.30 m failure line.

    Context

    The zero gate had been introduced in V2.5/V2.5b to slow the old lineage's get-up. In V3.1 the rushing was already absent: P1c got up in 0.90-1.06 s with a worst torque ratio of 73.3%, better than the stamped v2_6c, because the full beta curriculum, second-difference smoothing and pull curriculum had cured the violence inside training.

    Change

    Recorded as a candidate law with two data points - the zero gate and a wide stance are mutually exclusive here - and the zero gate was removed from the V3.1 recipe. P1c (the pre-gate checkpoint) went through the full stamp-level acceptance instead.

    Outcome

    P1c passed everything: all six criteria, lateral stance 0.355 m, foot tilt P75 2.0 deg, mu {1.0, 0.8, 0.6, 0.4} x 10 seeds all 100%. recovery_v3_1p1c.onnx was stamped and pushed to the robot channel.

    Mechanism

    The zero gate moves the income toward "stay stable until the end", and under low-friction DR a wide stance has a slip tail, so survival outbids the width band; doubling the band's price bought back only 0.016 m - an auction that does not converge signals structural conflict, not an under-priced term.

    Applies when

    • porting reward mechanisms from an old lineage into a fresh recipe
    • a width, margin or posture metric erodes during a late training phase
    • a weight increase produces a negligible change in its target
    “**定律候选(二实证):归零门 × 宽站互斥**。 … 加价翻倍只挽回 0.016,竞拍不收敛)。 … **归零门是 v2_5 血统的历史包袱,对 V3.1 配方是净负资产,P2 阶段除名 —— P1c 即终点形态**。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §49 终验(2026-08-14)
  • The latency DR range must cover the measured deployment pipeline - 0-20 ms could not even reach the real 1-2 control stepslatency-dr-covers-measured-pipeline
    Mechanism understoodwalkactuator-modelingactuator-modelingdomain-randomizationhardware

    Measure end-to-end action latency in control steps on your own stack (including cross-process queue boundaries), set the DR range to cover it with margin, and never import a delay count without its control frequency.

    Symptom

    Action latency was randomized over 0-20 ms (0-1 control step at 50 Hz), but the measured deployment path is 1-2 steps: the deploy process writes the target, an independently running BusWorker picks it up on its NEXT cycle, plus CAN round-trip - the training range could not cover the robot's actual latency at all.

    Context

    Fix: widen action_latency_s to 0-0.06 (0-3 steps). The external reference's "uniform 6 steps" was explicitly NOT copied - that number depends on his unknown control frequency; locally, a sweep at 0/1/2/3 steps showed walk_v5 survives all with insensitive metrics, so 6 steps "在我们这里没有依据" (has no local basis). The range was set from the measured pipeline with margin, not from a foreign constant.

    Change

    action_latency_s (0, 0.02) -> (0, 0.06), justified by pipeline analysis (writer/worker cycle boundary + bus time) and bounded by the local latency sweep.

    Outcome

    The DR band now brackets the true deployment latency; the policy trains against the delay it will actually face instead of a fictional sub-step world.

    Mechanism

    Latency DR only immunizes against delays inside its support; a range below the physical pipeline guarantees an untrained distribution shift at deployment. The correct range comes from tracing the pipeline's worst case (queueing boundaries + transport), and foreign step-counts are meaningless without the control rate they were measured at.

    Applies when

    • setting or auditing action-delay randomization
    • deployment uses a separate bus/worker process from the policy loop
    • importing delay-modeling numbers from other projects
    “现行 0~20 ms = 0~1 个 50Hz 控制步, 而实测部署链路是 1~2 步(deploy 写 STATE.target 后, 独立跑的 BusWorker 下一轮才取走下发, 再加 CAN 往返)——现在的区间覆盖不到真机的实际延迟。… 不照抄参考来源的"统一 6 步": 那取决于他的控制频率(未知), 而我们扫过 0/1/2/3 步 … 6 步在我们这里没有依据。”
    train/WALK_V7_SPEC.md § ⑤ action_latency_s 0~0.02 → 0~0.06
  • A pull-assist curriculum keyed to a global success share was satisfied by the categories that already worked and withdrew before prone learned anything - conditioning the criterion on prone took it from 2.5% to 98.7%curriculum-criterion-conditioned-on-lagging-category
    Mechanism understoodrecoverytraining-runcurriculumgate-battery

    Measure a curriculum's advancement criterion on the population the scaffold is meant to help; a global success share is met by whatever already works, and the help is withdrawn before the lagging case learns.

    Symptom

    Under the beta-anchored action, supine and side stood reliably while prone still sat (1.3%). A pull-assist curriculum added to help it was withdrawn completely within ~790 iterations and prone moved only to 2.5% (noise).

    Context

    The pull assist follows HoST: an upward force on the base, active only when the torso is within 30 deg of vertical, scaled by body weight (HoST's 200 N on G1 = 0.583 BW -> 56 N here, steps of 5.6 N, ten levels to zero); the product must pass with no assist. It had already taught sit-to-stand in V1. In V2.1 its advancement criterion was the standing-time share over all envs (threshold raised to 0.55 because the share was already ~0.53).

    Change

    V2.2: PullAssistForce with gate_category = "prone" - only envs whose first step classifies them as prone count toward the criterion - and the threshold back at 0.35. A feasibility signal was pre-registered: if prone's share stayed near zero under the full 56 N, return to the roll-over path instead of adding force.

    Outcome

    The prone-conditioned curriculum withdrew level by level only as prone itself passed: prone 98.7%, and the four-category gate passed for the first time on the line (98.6% overall, re-falls 0%, torque gate PASS).

    Mechanism

    A pooled success share is filled by the categories that already succeed (supine/side ~53%), so the scaffold is removed on their account before the lagging category has used it.

    Applies when

    • an assist, guide force or easier setting is withdrawn by a success threshold
    • one task category lags while the pooled metric looks healthy
    • a curriculum ran to completion without changing the lagging category
    “**教训:全局站立占比阈会被存量类别(supine/side ~53%)凑够,拉力在 prone 学会前就撤光了 —— metric 设计失误,不是拉力机制失效**(它在 V1 教会过 坐→站)。 … **V2.2(已启动)**:`PullAssistForce` 加 `gate_category="prone"` —— 达标判据 只统计 prone 类 env”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §34 V2.1 判决(2026-08-10)
  • After seven patch-generations, freeze the lineage as a regression baseline, fix the structural debts, and retrain from zerofreeze-lineage-fix-structure-restart
    Observed oncewalkprocessprocesscontract-freezecurriculum

    When successive rungs keep trading one symptom for another, ask whether the remaining problems are structural (contracts, latency, sampling, reward-table architecture); if so, freeze the lineage as regression baselines, pay the structural debts, and restart minimal - carrying forward laws and instruments, not weights and weights' patches.

    Symptom

    The v5-v11 walk lineage had accumulated interacting patches (reward terms, gates, clamps, per-joint scales) faster than it converged on the user's goal; v12's spec itself was superseded before training by an external review's verdict that the remaining problems were structural, not parametric.

    Context

    The 2026-08-05 status banner records the pivot: the walk profile was rolled back wholesale to v10b parameters, the v5-v11 lineage frozen "只作回归对照" (kept only as regression baselines), and four structural debts were named as prerequisites for a from-zero straight-walk baseline: the action-latency FIFO (fixed with its own test), the ONNX manifest contract, discrete command sampling, and a minimal reward table. The v12 spec - fully designed, partially implemented - was suspended: "本规格挂起,不再按此开训".

    Change

    Strategy switched from "one more patch generation" to freeze-fix-restart: lineage checkpoints retained as comparison anchors, infrastructure hardened first, then a clean retrain with a minimal reward table (this restart produced the s* generation that later became the real-robot SOTA line).

    Outcome

    A designed-and-ready training generation was deliberately not run - the review's structural findings outranked sunk design cost; the restart line inherited seven generations of laws (calibrations, gate batteries, falsified fixes) without inheriting their entangled reward table.

    Mechanism

    Patch lineages accumulate coupled terms whose interactions eventually cost more to reason about than a restart costs to train; the knowledge worth keeping is the laws and instruments (measured plant values, calibrated gates, falsified directions), not the entangled weights. A restart on hardened structure converts the lineage's lessons into a clean initial design instead of another delta.

    Applies when

    • repeated rungs shuffle symptoms without net progress
    • an external review flags infrastructure/contract debts
    • deciding between another patch generation and a clean retrain
    “同日外部评审定调换路线:冻结 v5~v11 血统(只作回归对照),修结构性问题(latency FIFO 已修 tests/test_action_latency.py、ONNX manifest 契约、离散命令采样、最小奖励表)后从零训直行基线。本规格挂起,不再按此开训。”
    train/WALK_V12_SPEC.md § ⚠️ 状态 (2026-08-05)
  • Anchoring the action on the measured joint angle (target = q + beta*a), with a beta curriculum down to tau_limit/kp, bounded torque by construction, removed the re-falls and later stood the robot up on hardwarebeta-anchored-action-target
    Mechanism understoodrecoverytraining-runactuator-modelingcontract-freezecurriculum

    For large-motion skills on position-controlled actuators, bound the action relative to the measured joint angle with a per-joint authority of tau_limit/kp, curriculum the authority down from full range, keep the curriculum state out of the observation and pin acceptance at the deployed authority - and make the deployment code refuse to run the anchored contract without a measured q.

    Symptom

    The V0 full-range absolute action produced violent targets; the V1 command-anchored rate limit made standing oscillate. Both failure modes came from how the action becomes a target.

    Context

    V2.0 (user approved, from scratch): BetaAnchorJointPositionAction, target = q_measured + beta_j(m)*a, memoryless per step. beta_j(m) = floor + m*(beta0 - floor), beta0 = the contract half-range (m = 1 reproduces V0 authority), floor = min(tau_limit/kp, beta0): hip_pitch 1.309 -> 0.40, knee 1.047 -> 0.40, hip_yaw -> 0.917, the other joints unchanged - the tightening lands exactly on the joints the V0 torque account convicted. m drops 0.1 per step when a standing-share EMA exceeds 0.35. beta is NOT in the observation, so the 45-dim contract is untouched; acceptance is pinned at m = 0 because the Python curriculum state is not saved in the checkpoint. The deployment chain got a new profile (recovery_v2: action_anchor current_q, explicit per-joint beta written into the contract, independent of the gain profile), and policy_io raises if q is missing rather than silently falling back to the absolute contract; the old profile's check reproduced its pre-change deviation bit for bit.

    Change

    New action term and beta curriculum; later the RS06 floor was lowered 0.40 -> 0.30 -> 0.25 (kp*beta 7.5 N*m) and the stamped deployment profile was synced to 0.25.

    Outcome

    First acceptance at m = 0 (v2_0b): re-falls 0% in every category, the torque gate passed for the first time on the line (worst 69.9%), knee jitter 0.004; supine 98.8 / side 88.8% with prone and mid still failing (fixed by the conditional pull curriculum). MuJoCo showed demand at or under the limits (hip_pitch 11.7/12 against V0's 26.8). Lowering beta cut impact (hip_pitch demand 9.7 -> 8.5 N*m) but barely slowed the get-up - it had become coordination-limited. Enabling the policy moves the target only +/-beta around the current pose, so there is no homing fling; the 08-11 real get-up and the later v3_1p1c both run on this contract.

    Mechanism

    kp*beta caps the proportional torque in a single step with no build-up delay and no memory, giving both a hard impact bound and full balance bandwidth.

    Applies when

    • a skill needs full joint range but hardware torque limits are low
    • absolute position targets cause impacts or saturation
    • changing the action semantics of a contract that deployed policies share
    “**动作项** `BetaAnchorJointPositionAction`:`target = q_实测 + β_j(m)·a`, 逐步无记忆 … **Play/验收钉 m=0(= floor = 部署档)**:python 课程状态不进 checkpoint, Play cfg 显式 `beta_m_start=0` … 判读:**结构赌注兑现** —— 站姿零再摔 + 力矩账首过(kp·β 封顶按构造)”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §33 V2.0 预注册(2026-08-10,用户点头开工):β 锚定动作空间,从零训
  • Ungated phase shaping made standing 42x more expensive than stepping - and the stepping was cooking the hip motorsmoving-gate-42x-stand-tax
    Mechanism understoodomnireward-shapingreward-shapinghardwarereal-acceptanceprocess

    Gate every phase/clock-driven shaping term on the command that justifies motion, price the cmd=0 case explicitly during design - and when a reward flaw is sim-only, book it with trigger conditions instead of operating immediately on a working lineage.

    Symptom

    At cmd = 0 the sim policy never stood still - it stepped in place and crept 0.98 m per 20 s; on the real robot the same lineage's stepping made hip_roll motors run 20 degC hotter than every other joint (43-48 degC vs 25-28).

    Context

    The arithmetic closed it: the three phase-shaping terms (joint_pos_ref 1.6, feet_contact_number 1.2, feet_clearance_swing 1.6) are driven by the gait clock with NO command gating, so standing at cmd=0 forfeits 3.32/step of shaping while honest standing earns only 0.078 of tracking - stepping wins 42x. The heat chain: perpetual stepping = perpetual single support = one hip_roll stalled at ~4.3 N*m (25% of torque limit) carrying the torso's frontal-plane moment - two hip_rolls = 90% of whole-machine steady-state I2R; measured stand-vs-step comparison: total heat -71% when actually standing. The suspended test acquitted the actuator (0.21 rad sag -> 0.0008 rad in air) and mechanics vetoed the easy fix ("降 kp 救不了热" - equilibrium torque equals the external load regardless of kp). The fix (moving_gate: hard-gate the three shaping terms on |cmd| > eps) was designed - then DEFERRED by the user because the real robot at the time stood fine: "真机不表现该问题, 为真机不存在的病改奖励表不划算", with written trigger conditions (C-ladder stand row persistently failing, or real robot starting to step/drift at cmd=0) and the known hazard tag (this is exactly the reward-change class that triggers the B-arm signature). When the real robot later DID step and cook, the booked trigger fired and moving_gate moved from debt to to-do with its benefit re-priced: "停止空烧 hip_roll,稳态发热降 ~71%".

    Change

    moving_gate designed with the gate_by_cmd convention; deferral, triggers, and expected heat recovery all pre-registered instead of patching the reward for a then-sim-only symptom.

    Outcome

    The cmd=0 stepping went from mystery to closed arithmetic; the thermal measurement (hip_roll +20 degC) quantitatively confirmed the 90%-of-heat prediction; the reward change waited for real-world justification instead of spending a risky revision early.

    Mechanism

    Clock-driven shaping terms define a perpetual-motion bounty unless gated by command; the resulting idle gait is not a training bug but the table's optimum. Its cost surfaces on hardware as stall-torque heating set by statics (mass x lateral offset), which no gain change can remove - only removing the motion (gating) or widening the stance can.

    Applies when

    • the policy steps in place or creeps at zero command
    • specific joints run hot in idle behaviors
    • deciding when a known reward flaw justifies a risky mid-lineage fix
    “塑形合计 −3.32/步 … 净: 站定亏 42 倍 … 两颗 hip_roll 4.29 / 4.09 N·m(各占限扭 25%),占全机稳态 I²R 的 90% … 吊挂实测 hip_roll 跟踪误差 0.21 rad → 0.0008 rad … 降 kp 救不了热 … 真机不表现该问题, 为真机不存在的病改奖励表不划算”
    train/README.md § C1 FAIL 节 (cmd=0 的 sim/真机分歧记账) / C2 真机 A/B 三b 发热定性
  • Curriculum-gate a penalty to the phase where its disease occurs - early on it only taxes explorationgate-penalties-to-the-disease-phase
    Replicatedwalkcurriculumcurriculumreward-shapingaction-rate

    For penalties aimed at late-stage pathologies (freezing, saturation, degenerate attractors), ramp the weight in only after exploration noise has decayed; anchor the terminal weight to measured healthy-vs-sick raw values, and shift all related tripwires to after the ramp completes.

    Symptom

    The action_saturation penalty, applied from iteration 0 in v8a, taxed exploration itself: with init_noise_std 1.2 the sampled actions paid ~-2.45/step before any policy had formed - while the disease it targets (clamp freezing) is a LATE pathology (v9 froze at iteration ~2624).

    Context

    v10 re-introduced the same penalty behind a curriculum gate: weight 0 until iter 1000, ramping linearly to -1.0 by iter 2000 - present only when the disease can occur, absent while exploration noise dominates. The trust argument was evidence, not hope: in v8a the term, while active, had pulled joint_pos_ref from 0.041 up to 0.155 and climbing - proof it can extract a policy from the frozen pit. Weight magnitudes were anchored to measured raw values (healthy v5 0.310 / v6 0.106 vs frozen v7 1.145 / v9 1.22 per step: at -1.0 healthy pays 6-18% of tracking, frozen pays 60%+, standing ~0). v10c then isolated the gated term as THE anti-freeze mechanism by single variable, upgraded to untouchable status in v11: "S 的门控机制(v10c 单变量铁案:任何情况下 不许撤,只许调终值)" - and v11 dared to relax other penalties only because S stood guard.

    Change

    action_saturation gated 0 -> -1.0 over iters 1000-2000 (later terminal value tuned -1.0 -> -0.5 with the gate mechanism itself frozen); tripwires adjusted to respect the gate's timing (freeze check moved to iter 2500-3000 to give the ramped term its effect window).

    Outcome

    Freezing stopped recurring while early training kept full exploration; the mechanism graduated from experiment to invariant within two versions.

    Mechanism

    A penalty's incidence depends on who occupies its support: early in training that is exploration noise (whose suppression starves learning), late it is the converged pathology. Time-gating aligns the penalty's presence with its target's presence, buying the constraint without the exploration tax - and tripwire timing must then be computed from the gate schedule, not from ungated precedents.

    Applies when

    • a structural penalty punishes exploration in early training
    • a late-onset pathology (freeze/saturation) needs a standing guard
    • deciding when a curriculum ramp should engage
    “v8a 实锤它的病根是"罚在采样动作上"——init_noise_std 1.2 的早期等于罚探索(~−2.45/步);而冻结是晚期病(v9 速率 2624 才死平)… 门控让它只在病发期在场。… v8a 里它在场时 joint_pos_ref 从 0.041 爬到 0.155 且仍在升——有从低谷爬出的实证。”
    train/WALK_V10_SPEC.md § 2. S 保险 —— action_saturation 课程门控
  • Low-friction robustness traced to kd DR bandwidth, not friction training - by digging resolved params across 8 lineages, 3840 cellskd-bandwidth-mu-law-attribution
    Mechanism understoodomniattributionattributiondomain-randomizationprocess

    Attribute capability differences by tabulating every lineage's resolved training params and eliminating zero-variance and non-aligned columns first; never let an eval-side override knob serve as the explanation axis, and never write a mechanism into a law before it survives a targeted test.

    Symptom

    Lineages differed wildly in low-ground-friction survival, and the intuitive explanation - "some trained ground friction, some didn't" - was about to steer the ladder toward a ground-mu training rung.

    Context

    The attribution ran as a full parameter-vs-result cross: 8 lineages x 4 eval kd levels x 6 mu levels x 20 seeds = 3840 cells, with each lineage's RESOLVED training params dug out and compared item by item. First kill: all 8 lineages had ground mu pinned at (1.0,1.0) - zero variance - so low-mu differences cannot come from friction training at all. The only training parameter aligned with the mu score was kd DR bandwidth: narrow (<=0.24) lineages scored 19.9/19.5/19.5, wide (>=0.40) scored 17.1/15.2/14.6/14.2/12.8 - the two groups completely non-overlapping. Every rival was excluded item by item (kd center no; kp band no; COM small-beneficial non-driving; friction rung a clean double null 19.5->19.5 and 15.2->14.6; iteration count non-monotonic), and the one clean single-variable causal link confirmed it: the s2e-3 kd surgery (0.7,1.3)->(1.08,1.32) moved the score 17.1->19.5. Counter-proof against "each best at its own operating point": the narrow-band lineage evaluated OUT of band (18.2) still beat the wide-band lineage at its own band center (9.2). Two axes were ordered never to be conflated (the first attribution's own error): training kd bandwidth is a parameter axis / lineage property; the eval-side --kd-scale knob is a plant axis (more damping physically helps on slippery floors for ALL policies) - "plant 轴只能当部署缓解,不能当 归因". A tempting mechanism story ("drag vs step attractor") was tested and falsified, and explicitly kept OUT of the law: "机制未定, 不入定律".

    Change

    The planned ground-mu training rung was recommended closed ("建议 不开") in favor of a kd band-narrowing rung (0.8,1.2)->(0.9,1.1) centered on the deployed value - with a pre-registered risk that the law demands "bandwidth = measured dispersion" and the real robot's kd dispersion was not yet measured; if it exceeds +/-10%, narrowing sacrifices real coverage and the rung must yield.

    Outcome

    A whole training rung was deleted from the ladder by attribution alone (the second S2 pass dropped mu and push, 5 rungs -> 3); floor material became a deployment-selection input (mu <~0.6 -> deploy the kd1.2 gain profile) rather than a training target.

    Mechanism

    Cross-lineage performance differences must be attributed over the actual training-parameter table, not over eval knobs or plausible stories: eval knobs act on the plant for every policy (a physical effect), while lineage properties come only from training-time parameters. Zero-variance columns are free eliminations, and one clean single-variable rung is worth more than any correlation.

    Applies when

    • explaining why lineages differ on a robustness axis
    • an eval-side knob (gain scale, power) changes results and invites misattribution
    • deciding whether to open a DR rung for an axis never actually varied in training
    “8 血统地面 μ 训练带全部钉 (1.0,1.0) 零方差,低 μ 差异与「训没训地面摩擦」无关,是 kd DR 带宽的副产物 … 宽 ≤0.24 → 19.9/19.5/19.5;宽 ≥0.40 → 17.1/15.2/14.6/14.2/12.8, 两组完全不重叠。… 训练 kd 带宽 = 参数轴/血统属性;评测部署 --kd-scale = plant 轴 … plant 轴只能当部署缓解, 不能当归因。… 机制未定, 不入定律。”
    train/OMNI_V0_SPEC.md § 4. 地面 μ 鲁棒性 = kd DR 带宽的副产物 (2026-08-08)
  • Four in-lineage attempts to widen the standing stance failed - remove a tax, add a joint-space knife, change the target, add a task-space metric penalty - because the stance was the end state of the get-up path; trained from scratch with the right terms it grew right from day onestance-decided-by-get-up-path
    Replicatedrecoveryattributioncurriculumfork-selectionreward-shaping

    A posture a skill ends in is shaped by the path the policy takes to reach it; if several single-variable edits to the terminal-phase reward cannot move it, stop editing that phase and retrain with the terminal constraint present from the start.

    Symptom

    v2_6c stood with its feet 0.159 m apart (task-space) and its hips yawed 45-47 deg the same way, which split on the real robot. Standing-phase reward edits did not move it.

    Context

    V2.7-A removed the flat-feet tax on compensated stances (stance unchanged); V2.7b added a hip-roll lower-bound hinge (+5 deg in 3,000 iterations, yaw ratchet); V2.8 changed the stand_pose target to a wide flat stance (stance unchanged, yaw not unwound, feet nearly overlapping, mu 0.4 transfer 2%); V2.9 penalized lateral spacing in metres (the policy parked just outside the penalty's gate in a lunge, 0% success). The v2_6c get-up goes through a split and closes the feet together as it rises.

    Change

    In-lineage stance surgery was formally closed. V3.1 trained from scratch with task-space stance terms present from the first iteration (and, after P1, a positive width band instead of a penalty).

    Outcome

    V3.1 P1b: lateral stance 0.364 m, foot tilt 0.0 deg, all four categories 100%, MuJoCo mu 1.0 and 0.4 both 100% - with a symmetric toe-out the kinematic audit had not enumerated. P1c (with a yaw guard): 0.355 m, all six acceptance criteria passing, mu 1.0-0.4 all 100%; it became the product.

    Mechanism

    A converged policy does not rebuild the path that produced its terminal posture; a standing-phase gradient only finds the nearest hack around the posture the get-up delivers.

    Applies when

    • the final posture of a transition skill is wrong and resists terminal-phase shaping
    • repeated continuation rungs produce hacks instead of the intended posture
    • deciding between another in-lineage fix and a from-scratch retrain
    “窄站距 + yaw 扭是 v2_6c 起身策略(劈叉起身 → 双脚并拢收势)的**结构性 终态**,不是站立段的孤立参数 —— 站立形态由起身路径决定,在血统内只动 站立段奖励改不动它。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 结果:V2.8 判 FAIL —— 血统内站姿手术第三次证伪
  • Brief the operator on the lineage's measured zero-command and untrained-axis behavior before handing over the joystickknow-zero-command-behavior
    Replicatedomnireal-deployreal-acceptanceprocess

    Before any teleop/demo, measure and write down the policy's zero-command behavior and per-axis competence, label untrained axes explicitly as not-bugs, and set the floor/procedure to accommodate the known drift.

    Symptom

    A teleop session was about to start on a policy that does not stand still at zero command and has never been trained on lateral commands - behaviors an unbriefed operator would report as bugs or emergencies.

    Context

    Three measured facts were written into the teleop instructions ("都有 实测依据, 不是猜"): (1) A/D (lateral) keys will get essentially no response - probe-measured sidewalk tracking ~3%, an untrained axis: "这正是 C4 要解决的事, 不是 bug"; (2) no keypress = cmd 0, and this lineage does not stand still at zero command - a three-generation lineage property: paces in place, drifts right ~5 cm/s, net rotation -30 deg/20 s; sim survival is 20/20 (it will not fall) but it walks away slowly, so leave floor margin especially on the right; (3) S (backward) WILL respond - probe-measured 20/20 survival, 67% tracking untrained, which is also why this root was chosen for the C ladder. Plus a keybinding dry-run while suspended before touching down.

    Change

    Operator briefing became part of the deployment artifact: expected response per key, expected idle behavior with magnitudes and directions, and the distinction between untrained (expected, not a bug) and abnormal.

    Outcome

    The session proceeded with correct interpretations available in advance; the known zero-command wander was handled by floor margin and start-with-command procedure rather than misdiagnosed on the spot.

    Mechanism

    A learned policy's off-nominal behaviors (idle drift, untrained axes) are lineage properties, stable and measurable in sim beforehand; operator surprise converts known properties into false incident reports and unsafe reactions. A briefing transfers the measured behavior model to the person holding the controller.

    Applies when

    • handing a learned policy to an operator or demo audience
    • the policy idles in a non-stationary way at zero command
    • some command axes are untrained in the current lineage
    “A/D 基本不会有反应 —— s1e 从未训过非零 vy, 选根探针实测侧走跟踪率 ~3% … 这正是 C4 要解决的事, 不是 bug。… 不按键 = cmd 0, 而 s1e 在零指令下不站定 —— 血统属性, 三代实录: 原地踏步 + 右漂 ~5 cm/s + 净旋 −30°/20s。”
    train/REAL_RUN_S2.md § 附: WSAD 遥控 上机前必须知道的三条
  • Fix the task first, harden the plant second - DR budget spent on a dying task is wastedtask-shaping-before-plant-hardening
    Mechanism understoodomniprocessdomain-randomizationprocesscurriculum

    Freeze the task/command distribution before spending DR budget on plant robustness; if the task will still change, schedule plant hardening as a final pass and book the interim robustness gap explicitly.

    Symptom

    Tempting default ordering was to keep the plant-hardened (S2) lineage and teach it new commands; but the S2 plant adaptation had been earned on the straight-walk task, and the new omni tasks (sidewalk, in-place turn) use completely different contact patterns.

    Context

    The team had direct evidence that DR robustness is a budget that gets reallocated when the data distribution changes ("push/μ 两轮已实证 DR 预算有限且会被重分配") - robustness trained under one task/command distribution does not persist when training continues under another.

    Change

    Ladder order set to: first C (task shaping - add command modes until the task family is final), then a second S2 pass (plant hardening) on the C product. The plant-robustness gap this creates mid-ladder is accepted and booked explicitly ("此处不欠账" - the debt is assigned to the second S2 pass, not denied).

    Outcome

    The first S2 pass was not wasted: its laws (kd bandwidth <-> low mu, push need not be trained, ground mu need not be trained, bistability) let the second pass drop from five rungs to three. The C ladder itself ran on the softer plant band without incident.

    Mechanism

    DR robustness is carried by the policy's visited-state distribution; changing the task changes that distribution, so robustness bought under the old task partially dissolves. Hardening before the task is final means paying for robustness on states that will no longer be visited - "给一个即将不存在的任务花预算" (spending budget on a soon-to-not-exist task).

    Applies when

    • deciding ordering between skill/command expansion and DR hardening
    • a hardened lineage is proposed as the root for a task change
    • robustness regressions appear after adding new command modes
    “S2 的 plant 适应是为直行步态调的,C4 侧走/C3 原地转是完全不同的接触模式,先硬化再改任务 = 给一个即将不存在的任务花预算(push/μ 两轮已实证 DR 预算有限且会被重分配)。故顺序改为 先 C(任务定型)→ 再 S2(plant 硬化)。”
    train/C_LADDER_RUN.md § 0. 决策逻辑 = 短板可不可恢复 (末段)
  • PPO's Gaussian noise cannot compose phase-locked oscillations - deliver them as feed-forward and let the policy learn the residualfeedforward-for-phase-locked-skills
    Mechanism understoodomnicurriculumcurriculumreward-shapingcontract-freeze

    If a skill needs a temporally coherent (phase-locked) action component, do not expect step-wise exploration to find it: inject a verified feed-forward and train the policy as a residual stabilizer, keeping the feed-forward inside the deployment contract.

    Symptom

    Four different reward arrangements (no reference / wrong-sign reference / correct-sign reference / cage released) all failed to elicit sidewalk, while open-loop probes proved the behavior existed and was safe on the same platform with the same policy as base.

    Context

    Producing lateral velocity requires a phase-locked hip_roll oscillation synchronized to the gait clock. PPO's exploration is per-step, zero-mean, uncorrelated Gaussian noise - it can never compose a sustained phase-locked component, so the behavior is unreachable by exploration regardless of how it is rewarded. The fix changed the delivery channel: target = default + scale*action + lat_ff(cmd_vy, phi). The policy's action becomes a residual on top of the feed-forward, retaining full balance authority (it can even cancel the feed-forward); the feed-forward supplies exactly the component exploration cannot. This mirrors why the sagittal joint_pos_ref worked (it also delivered phase structure), just via a different channel.

    Change

    Contract-level change, done cleanly: new profile omni_ff (= omni + lat_ff_gain -0.5), existing omni profile bit-identical; feed-forward applied after the action delay stage; missing cmd/phase raises instead of silently dropping; deployment must use the same phi as build_obs (recomputing gives a one-tick phase misalignment).

    Outcome

    From C2-700, +100 iterations sufficed: product omni_c4_ff800 scored vy +120%/+125% (from +4%/-1%), 260/260 cells at 20/20 survival, zero old-skill regression, left/right gap 5 pp - the entire C4 saga resolved by changing the delivery mechanism, not the reward.

    Mechanism

    Exploration noise spans only the subspace its correlation structure can express; skills requiring coherent oscillation lie outside the span of i.i.d. per-step noise. Feed-forward moves the required structure into the action pipeline where it needs zero probability mass to appear, reducing the learning problem to stabilizing around a demonstrated behavior - which PPO does well.

    Applies when

    • a periodic/oscillatory skill trains flat under every reward variant
    • open-loop injection of the behavior already works
    • considering GRU/curriculum/exploration tricks for a rhythmic skill
    “病因不在奖励,在探索形式:产生侧向速度需要相位锁定的 hip_roll 振荡,PPO 的逐步高斯噪声零均值无相关,合不出相位锁定分量。… target = default + scale·a + lat_ff(cmd_vy, φ)。策略动作因此是前馈之上的残差,保留全部平衡权限”
    train/C_LADDER_RUN.md § 3j. C4-redo4:唯一变量 = 侧步参考改为前馈注入(契约级)
  • The action-delay was implemented as lerp - beyond one step it extrapolated BACKWARD, so a whole lineage trained on a fictitious actuatorlatency-lerp-reverse-extrapolation
    Mechanism understoodinfraactuator-modelingactuator-modelingplant-calibrationsim2sim

    Unit-test plant-model code (delays, filters, randomizers) against hand-computed truth across its FULL configured range, not just the nominal case - a delay must be a queue, and any interpolation used outside [0,1] is a silent plant corruption that training will faithfully absorb.

    Symptom

    Every walk/stand model up to v11 had been trained on a silently wrong plant: the action-latency implementation lerp(cur, prev, lag) is only an interpolation for lag <= 1 - at lag 3 it computes 3*prev - 2*cur, a REVERSE extrapolation. With the configured (0, 0.06) s at 50 Hz (lag in [0,3]), about 2/3 of environments were adapting to actuator dynamics that do not exist.

    Context

    Listed as evidence item #1 for the full restart: "全部旧模型训在错误 plant 上" and the head suspect for the real robot's wild kicking. The fix replaced it with a true FIFO delay line (commit 7f13793) plus its own regression test (tests/test_action_latency.py) - but every exported ONNX predated the fix, which is part of why the lineage was frozen rather than patched.

    Change

    Delay implementation rewritten as an honest FIFO with unit tests; the restart baseline trained on the corrected plant from day one.

    Outcome

    A generation-scale training investment was revealed to have a corrupt plant underneath; the class of bug (plausible-looking math that silently changes meaning outside its valid range) got a permanent test.

    Mechanism

    lerp(a, b, w) leaves the segment for w > 1; used as a delay it fabricates high-gain inverted dynamics precisely in the largest-delay draws, so the policy learns compensation for an actuator that cannot exist - and DR then trains robustness to the artifact rather than to reality. No training metric can catch this: the sim is self-consistent, just wrong.

    Applies when

    • implementing or auditing action delay / filtering in a trainer
    • a lineage behaves as if compensating dynamics nobody modeled
    • deciding whether old checkpoints are salvageable after a plant bug
    “动作延迟旧实现 lerp(cur, prev, lag) 在 lag>1 时是反向外插(w=3 → 3·prev−2·cur),配置 [0,0.06]s@50Hz 即 lag∈[0,3],约 2/3 env 在适应不存在的执行器动态。7f13793 已换真 FIFO,但所有 ONNX 均训于修复之前 —— 真机"乱踢"的头号嫌疑。”
    train/OMNI_V0_SPEC.md § 0. 为什么从零 (1)
  • Fine-tuning through a reward-table change was falsified (scatter, half-recover, collapse) - continuation training is legal only with the reward frozenfine-tune-reward-change-falsified
    Replicatedomnicurriculumcurriculumfork-selectionreward-shapingprocess

    Never fine-tune through a reward-table change - retrain from zero; reserve checkpoint continuation for frozen-reward plant/DR widening, reset noise_std when branching, and watch for the scatter/half-recover/collapse signature as the abort trigger.

    Symptom

    The s1c A/B experiment: arm B fine-tuned from an existing checkpoint under the revised reward (same contract, same network, changed reward + small DR) and failed with a characteristic signature - scatter, half-recover, fall back ("打散→半恢复→摔回"); arm A trained from zero under the same config won decisively (full shaping lifted swing to 21.6 mm within 500 iters; shipped at 5500).

    Context

    Verdict recorded: "从零 + 强塑形是本机唯一验证过的发育路径" (from-zero plus strong shaping is this machine's only validated development path). The signature became a standing stop criterion in every later rung that touched a reward ("s1c B 臂签名,出现即停"). Crucially the boundary of the law was drawn explicitly when S2 continuation training was proposed: "当年证伪的是「奖励表中途改版的 fine-tune」… S2 续训奖励表全程冻结、只逐级加宽 plant/DR,属另一类" - continuing a checkpoint with the reward FROZEN while widening plant/DR one rung at a time is a different class and was allowed (and then worked, powering the whole S2/C lineage) - with the honest fallback that if frozen-reward continuation ever collapses, that rung retrains from zero and the doctrine gets re-examined with data. Fine-tune arms also need mechanical care: reset the checkpoint's collapsed noise_std (terminal 0.033 "会杀死探索") and account for iteration counters re-zeroing (curriculum gates fire immediately).

    Change

    Reward changes and lineage continuation permanently separated: reward revisions -> from-zero retrain; plant/DR widening -> frozen-reward continuation with per-rung gates; the B-arm signature promoted to a universal tripwire.

    Outcome

    No later reward revision was attempted by fine-tune; frozen-reward continuation carried S2 (PD/COM/friction rungs) and the C command ladder successfully from the s1e root.

    Mechanism

    A trained policy sits in an optimum of its reward's geometry; changing the reward moves the optimum but leaves the policy's exploration noise near-zero and its value function calibrated to the old returns - it disassembles the old solution faster than it can assemble the new one. Widening DR under a frozen reward instead keeps the optimum's identity and asks only for local robustification.

    Applies when

    • proposing to fine-tune an existing policy under a revised reward
    • planning a robustification ladder from a validated checkpoint
    • a continued run scatters then partially recovers then collapses
    “B 臂 fine-tune 证伪(打散→半恢复→摔回——从零 + 强塑形是本机唯一验证过的发育路径)。… 当年证伪的是「奖励表中途改版的 fine-tune」(B 臂,塑形突变致终盘摔回);S2 续训奖励表全程冻结、只逐级加宽 plant/DR,属另一类;若 s2_lag1 续训本身塌方,回退方案 = 该级从零重训,续训教义再议(拿数据说话)。”
    train/OMNI_V0_SPEC.md § 3. S1.3 / 4. 与 s1c fine-tune 证伪的关系
  • Gate a new reward term by its command so all old modes score pointwise identicalgate-new-reward-terms-by-command
    Mechanism understoodomnireward-shapingreward-shapingprocess

    When a reward term must be added mid-lineage, gate it on the condition that defines the new task so every pre-existing situation scores exactly as before - and still watch for value-rescale pathologies inside the new mode.

    Symptom

    Adding a lateral tracking reward (track_lin_vel_y_exp) ungated would have paid 0-2.0 per step even in modes with cmd_vy = 0 (healthy gait sway of vy ~0.1 already earns 1.28), shifting the whole reward table by a large bias and rescaling the value function - no longer "just adding one mode".

    Context

    C4 was the C ladder's only true reward surgery. Single-variable discipline required that the change be invisible to every existing mode. The chosen construction: gate_by_cmd=True - the term pays only when |cmd_vy| > 0.02, so for all modes with cmd_vy == 0 the term is pointwise zero, i.e. the reward is pointwise identical to before the change. The same trick appeared earlier in C1: replacing the vy L2 tax with a command-error version that is "对 cmd_vy≡0 逐点同值" (pointwise equal when cmd_vy is 0), explicitly classified as not-a-reward-change.

    Change

    track_lin_vel_y_exp added with gate_by_cmd=True (weight +2.0, std 0.15); the residual acknowledged honestly - inside the side bucket the values DO change, so the rung still watched the known reward-reshuffle pathology signature (s1c B-arm: scatter -> half-recover -> collapse) as a stop criterion.

    Outcome

    Old modes provably unaffected (pointwise-equal argument); attribution for any change in old-skill metrics stayed clean through the C4 redo series.

    Mechanism

    PPO's critic normalizes to the reward scale it sees; an ungated additive term shifts returns in every state and re-scales advantages globally, entangling the new skill with all old ones. Command-gating confines the new term's support to the new mode's state distribution, making "pointwise identical elsewhere" a provable property rather than a hope.

    Applies when

    • adding a tracking/shaping term for a new command or skill to a lineage that must not regress
    • reward change proposed while other skills are still being gated
    • reviewing whether a config diff counts as a reward change
    “只在 |cmd_vy| > 0.02 时付。不门控的话它对 cmd_vy≡0 的老模式也给 0~2.0 分(健康摇摆 vy≈0.1 → 1.28),等于给整张奖励表加一个大偏置、值函数尺度全变 … 门控后老模式逐点得 0 = 与加项前逐点同值,单变量纪律成立。… 但 side 桶内的值确实变了 —— 这仍是奖励表改版,开级盯 s1c B 臂签名”
    train/C_LADDER_RUN.md § 3d. gate_by_cmd=True(重要)
  • Record-high training reward hid a fully-failing DR subgroup - aggregate metrics average over draws, gates must test per conditionaggregate-metrics-mask-subgroup-failure
    Mechanism understoodomnitraining-rundomain-randomizationgate-batterycurriculum

    Never gate on metrics aggregated across DR draws: evaluate at fixed representative conditions (especially the deployment-critical stratum), and if a difficulty axis matters, ramp it on measured per-stratum success rather than sampling the full range from iteration zero.

    Symptom

    omni_s1e trained under constant-wide latency DR (0, 0.06 s) posted the lineage's highest-ever Isaac reward (129) - while the --delay 2 smoke evaluation showed 3/3 falls from iter 1500 onward, persisting to early stop; the usable checkpoint window shrank to iters 500-1000.

    Context

    Diagnosis written plainly: "聚合奖励掩盖重延迟尾部子群体失败" - the aggregated reward averages over latency draws, so the majority of light-delay environments can mask the total failure of the heavy-delay tail. The remedy for the training side was a survival-gated ratchet curriculum (survival_gated_latency): the sampling cap starts at 0.02 s and rises +0.01 only when a 4096-reset window's survival (time_out share) reaches >=90%, capped at 0.06, ratchet up-only - "增益与延迟耐受一起长,不升到策略撑不住的 地方" (gain and delay tolerance grow together; never raise past what the policy can hold). The detection side was already in place from the noise-crutch episode: the per-condition smoke curve, not the training reward, is the health readout.

    Change

    Latency exposure made curriculum-gated on measured subgroup survival instead of uniform-from-zero; per-condition (--delay 2) smoke evaluation kept as the authoritative curve; watcher scoring adjusted (survival weighted 3x) so recovery during hard phases is not early-stopped away.

    Outcome

    The failure mode was caught by the smoke curve within one generation; the follow-up redesign (deterministic staged latency) superseded the ratchet, but the aggregate-masking lesson held through both.

    Mechanism

    Expected-return training weights each DR draw by probability, so a subgroup can contribute bounded loss while being catastrophically failed; any scalar averaged over the randomization cannot distinguish "uniformly decent" from "great on easy draws, dead on hard ones". Only conditioning the evaluation on the stratum reveals the split, and curricula should raise difficulty on measured stratum success, not on schedule.

    Applies when

    • training reward hits records while a fixed-condition eval degrades
    • wide DR on an axis where deployment sits at one known value
    • designing curricula for difficulty axes (delay, push, terrain)
    “常量 latency DR (0,0.06) 从零训被证伪——Isaac reward 129 历代最高,但 --delay 2 冒烟 iter1500 起 3/3 全摔持续到早停(聚合奖励掩盖重延迟尾部子群体失败,可用窗口只剩 500/1000)。… 采样上限 0.02 起步 … ≥90% 才 +0.01s,0.06 封顶,棘轮只升不降。”
    train/OMNI_V0_SPEC.md § 3. S1.5(s1e 训练塌方复盘)
  • Ideal PD is not enough - add a delay buffer and fit armature/friction/delay per jointactuator-delay-buffer-fitting
    Observed oncewalkactuator-modelingactuator-modelingplant-calibration

    Never ship ideal PD to hardware: add a measured delay (in control steps) and per-joint armature/friction fitted from step and sine responses, and treat remaining actuator mismatch as your standing largest sim2real residual.

    Symptom

    Standard ideal PD actuator model transfers poorly; sim assumes targets take effect instantly and joints reach arbitrary acceleration.

    Context

    A developer with a successful on-hardware Isaac Lab biped modified the actuator model in two ways and calibrated it against the real robot: step-response plus sine-sweep tests (positive step, negative step, sine tracking), overlaying sim curves on measured curves and hand-tuning.

    Change

    (1) Delay buffer: action targets take effect after a uniform 6 time-step delay on all joints; (2) acceleration limiting so the actuator cannot reach arbitrary acceleration; (3) per-joint fit of armature / friction / delay - different joints genuinely needed different values.

    Outcome

    Hip joints fit worst, knee best; the developer rated the result "not perfect, the best I could do" and still listed actuator-model improvement as next work - i.e. even the fitted model remained the dominant residual.

    Mechanism

    Real actuation is a lagged, bandwidth-limited system; a delay buffer and acceleration cap are the two cheapest structures that reproduce its phase and magnitude response. Per-joint differences come from differing load, wiring, and friction states, so a single global constant underfits.

    Applies when

    • actuator model in sim is ideal PD with no delay
    • step-response of real joint visibly lags or overshoots the sim's
    • budgeting which sim2real gap to attack first
    “标准 ideal PD actuator 不够用,他改了两处:延迟缓冲:目标不是立即生效,全部关节统一 6 个 time step 延迟 / 加速度曲线:执行器不能瞬间达到任意加速度 … 用 armature / friction / delay 三个参数逐关节拟合,标定方法是阶跃响应 + 正弦扫描 … 髋部关节偏差最大,膝关节最好。”
    Experience.md § 执行器建模 —— 最值得抄的一条 (lines 50-59)
  • The deploy-side walk/recovery switch - into recovery at tilt > 65 deg held 0.3 s, back at tilt < 15 deg with angular rate < 1 rad/s and straight knees held 1 s, a 15 s timeout, last action cleared both ways - and the handoff steps the design required are only partly implementedwalk-recovery-fsm-handoff
    Observed oncerecoveryreal-deployreal-acceptancecontract-freezeprocess

    Specify a deploy-time controller switch as hysteretic, time-filtered predicates the robot can measure (proxy what it cannot, e.g. straight knees for height), a timeout that ends in a safe stop, and a complete handoff (history, clock, last action, command ramp) - then test that the code performs every handoff step, because the design document is not the implementation.

    Symptom

    With a recovery policy and a locomotion policy as separate networks, the robot needs a switch: when is it "fallen", when is it "up", and what state must be reset so the next policy does not act on the previous one's history.

    Context

    The 08-09 design: enter recovery when fallen (tilt > 55 deg or height < 0.60 x 0.384 m) for 150 ms, leave for a stand-hold when upright (tilt < 12 deg, height > 0.85 x 0.384 m, feet steady, |omega| < 0.8) for 400 ms - wide entry, strict exit, hysteresis - then a mandatory handoff trio before walking resumes (reset the walking policy's observation history, restart its phase clock at 0, clear its previous action and latency buffer) and a command ramp instead of a jump. The 08-14 implementation in deploy_policy (--recovery-policy): each policy under its own manifest contract (walk: nominal + scale; recovery: beta-anchored), the same rl_default gains, the tilt cutoff disabled; RECOVERY at tilt > 65 deg for 0.3 s; LOCO again at tilt < 15 deg and |omega| < 1 and knees straight (< 0.35 rad) for 1 s - the robot's computer has no height estimate, and straight knees stand in for height so a V3.0-style upright kneel cannot pass as standing; RECOVERY longer than 15 s ends in a safe stop; last_action cleared on both switches; command forced to 0 during RECOVERY; power scaling applies to LOCO only.

    Change

    An open account was written down with it: the recovery end state (0.355 m stance, hip yaw -/+27 deg) is outside the walking policy's training start distribution, so the first test must use stand / zero command as LOCO, and the long-term fix is to widen the walking policy's initial states rather than bend recovery's stance to suit walking.

    Outcome

    The spec records only a successful compile (py_compile); the hang test was left to be done on site. The operator runbook carries the three-step procedure (hang with stand as LOCO, mat and push, then omni walk as LOCO) but no outcome.

    Mechanism

    Two policies trained separately each assume their own history, clock and last action; a switch that carries any of them across feeds the next policy a state it never saw - the design called this "the walking policy seeing a ghost history".

    Conflicts

    The 08-09 design requires resetting history, clock and previous action plus a command ramp; the 08-14 implementation records last_action clearing and a zero command during recovery; the one-leg spec of 2026-09-14 lists the handoff hygiene as specified but not implemented - reset_history() is called by nothing (a 215-dim policy would carry four frames of pre-fall history), the phase clock is not zeroed, and there is no command ramp back to LOCO. No hardware run of the FSM is recorded in either source.

    Applies when

    • switching between separately trained policies on hardware
    • a policy with history or phase observations is re-enabled mid-run
    • the robot lacks a sensor the switching criterion was designed around
    “**判据**:进 RECOVERY = 倾角 >65°(`--fall-tilt-deg`)持续 0.3 s;回 LOCO = §5 真机可测子集:倾角 <15° ∧ |ω|<1 ∧ **膝直 <0.35 rad(NX 无高度观测, 高度门用膝直代理 —— 防 V3.0 型"跪坐但直立"误判)** 持续 1 s (`--recover-hold`);RECOVERY 单次 >15 s(`--recovery-max-time`)安全停。 … **切换卫生**:两向切换 last_action 清零;RECOVERY 态 cmd 强制 0”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §50 FSM 双策略调度(2026-08-14,用户令):deploy_policy --recovery-policy
  • Training at scale 0.4 is NOT the twin of deploying 0.5 at power 0.8 - the algebra matches, the learned policy does notdeploy-scaling-not-training-equivalent
    Mechanism understoodomniattributionattributioncontract-freezesim2simcurriculum

    Never assume deploy-side scalings can be folded into training-time constants ("burning the crutch into training"): the learned optimum depends on the training-time authority, so treat such conversions as full experiments with pre-registered expectations and a sim2sim gate before any hardware.

    Symptom

    s1g (S1.6) trained from zero at action_scale 0.4 - meant as the "training twin" of the hardware-proven s1c-at-power-0.8 (0.8 x 0.5 = 0.4) - was all green in Isaac (zero falls, reward 117) yet scored 0/3 across all eight checkpoints and 0/20 at 20 seeds in the MuJoCo gate, falling forward at median 1.57 s with a 2.9x speed overshoot.

    Context

    The pre-registered expectation (survival gate should pass, since the conviction matrix showed s1c@0.8+delay2 all-survive) was cleanly falsified, and the harness was acquitted by controls: --delay 0 fell identically (not a delay fragility), check_contract all green, and s1c through the same harness survived 2/3. The verdict: "「s1c@0.8 = 0.4 训练孪生」的代数等价不成立" - a policy deployed with a derated output still LIVES in the 0.5 internal model it trained under (its value function, its expectations of its own authority), while a policy that starts training with reduced authority learns a different, clip-hugging gait with zero margin for plant differences ("部署端打折的策略活在 0.5 的内模里,训练起点收权限学出的是贴 clip 的 另一套步态,对 plant 差异零余量"). Result: the policy was withdrawn before hardware ("撤回——不上真机"), the lineage root moved back to the 0.5-contract s1c-5500, and this became the C ladder's cited fact-check ("s1g 是 0/20 证伪出局的那一代").

    Change

    The amplitude-surgery route abandoned; contract kept at scale 0.5; the deploy-side 0.8 crutch later retired on its own merits when the delay-complete s1e generation ran at full power.

    Outcome

    One training run bought a clean falsification of a plausible algebraic identity; no hardware time was spent on it because the sim2sim gate caught it.

    Mechanism

    Output scaling commutes with the network arithmetic but not with learning: the training-time scale shapes which gait solutions are reachable and how much clip headroom the optimum keeps. A derated mature policy retains the wide-authority solution executed softly; a from-zero narrow-authority policy finds a different optimum that saturates its smaller envelope - the two are not the same controller in different units.

    Applies when

    • proposing to move a deployment derating into a training constant
    • a scaled-down contract policy hugs the action clip
    • Isaac-green / cross-sim-zero results on a re-scaled lineage
    “预注册 a) 证伪——Isaac 全绿(零摔/reward 117)但 MuJoCo --delay 2 八档 checkpoint 扫描全数 0/3、iter6500 20-seed 0/20 … 「s1c@0.8 = 0.4 训练孪生」的代数等价不成立: 部署端打折的策略活在 0.5 的内模里,训练起点收权限学出的是贴 clip 的另一套步态,对 plant 差异零余量。”
    train/OMNI_V0_SPEC.md § 3. S1.6 判决(2026-08-07 验收)
  • Adapting a lineage to one plant increment needs hundreds of iterations, not thousands - long runs only buy specializationcontinuation-budget-not-from-zero
    Mechanism understoodomnitraining-runcurriculumprocess

    Budget continuation rungs by increment class (hundreds of iterations for plant pins and smooth shifts, ~1000-1500 only for behavior-demanding changes like push), enforce a hard cap with frequent evaluation, and treat remaining budget as a reason to stop, not to continue.

    Symptom

    The default "6000 iterations per rung" (a from-zero-scale budget) was about to be applied to continuation rungs whose only change is one plant/DR increment - overspending compute and, worse, giving each rung thousands of iterations to specialize away retained skills.

    Context

    The 2026-08-07 budget table replaced the default with "最低适应窗口 + 每 100 iter 验收 + hard cap" scaled to the increment's difficulty: fixed-latency levels 300-500 (cap 500-800; the base has already seen in-band values, this only pins the plant); PD full-band 700 (cap 1000; kp+/-20%/kd+/-30% clearly widens the actuator family); COM +/-20 mm 500 (cap 800; a smooth dynamics shift); friction DR 700 (cap 1000; contact and actuator friction change the gait/contact solution together); push 1000 (cap 1500; a non-static plant change requiring recovery behavior - hardest). Rationale: "续训适应一个 plant 增量不需要从零量级的预算,跑长了只是给特化时间". The C ladder reused the scheme (per-rung caps 500-2000 by increment type), and the deep-training hazard got its own name when long runs sold quality ("深适应卖质量" - deep adaptation sells quality).

    Change

    Per-rung iteration budgets set by increment class with hard caps and 100-iter watch loops; checkpoint selection inside the window by the smoke curve, never "run to cap because budget remains".

    Outcome

    S2/C rungs completed in 300-1500 iterations each; the recurring late-run degradations (collapse valleys at 1500+, vx+0.30 decay) fell outside most rungs' caps instead of inside their runs.

    Mechanism

    A continuation rung's learning problem is local robustification around an existing optimum - low sample complexity; iterations past adaptation are spent sharpening onto the current distribution, which is exactly how retained skills and margins erode. Budgets sized to the increment bound both compute and the specialization damage window.

    Applies when

    • planning iteration budgets for a robustification or command ladder
    • a continuation run keeps improving its training metric late
    • retained skills decay in the back half of long continuation runs
    “「最低适应窗口 + 每 100 iter 验收(watch_ckpt --every 100)+ hard cap」—— 续训适应一个 plant 增量不需要从零量级的预算,跑长了只是给特化时间 … ⑥ push | 1000 | 1500 | 非静态 plant 变化,要学 recovery 行为,最难”
    train/OMNI_V0_SPEC.md § 4. 每级 iter 预算(2026-08-07 用户定)
  • The smoke watcher panicked at the wrong operating point and its score goes blind when gates saturate - treat it as a survival sentinel, not a judgesmoke-watcher-operating-point
    Replicatedomnisim-evalmeasurementprocessgate-battery

    Configure every automated evaluator at the lineage's declared deployment operating point, give it graded metrics that cannot saturate, and until then scope its authority to catastrophe-detection - never let a mis-configured watcher stop or rank a rung on its own.

    Symptom

    Two watcher misfires in one ladder: (1) during the friction rung the watcher (evaluating at kd 1.0) reported panic-level 1/3 survival from iter 3300 - falsified by the official kd 1.2 scan, because the lineage's design operating point was kd 1.2 and the watcher lacked the --kd-scale passthrough; (2) during the PD rung the watcher's early-stop score froze at iter 1050 despite ongoing drift improvements, because with all eight gates passing (constant 0/3 failures) the score has no gradient left - "八门全过恒 0/3 时 score 对漂移改善盲, s1f 课文三现".

    Context

    Both are the same category: the in-training smoke loop is an instrument with its own configuration (operating point, score design), and its verdicts are only as aligned as that configuration. The booked doctrine: "冒烟只当存活哨兵" - until the watcher evaluates at the deployment operating point with graded metrics, its role is detecting catastrophes, not ranking checkpoints; ranking belongs to the full battery at the design operating point (and the watcher's scoring was separately patched to weight survival 3x so recoveries during hard phases are not early-stopped away).

    Change

    Watcher debt booked (--kd-scale passthrough); score saturation acknowledged with graded columns planned; selection authority kept with 20-seed batteries at the declared operating point.

    Outcome

    A false panic did not abort a rung that was actually passing at its design point; a frozen score did not hide real drift gains; the instrument's authority was scoped to what its configuration can actually see.

    Mechanism

    An evaluator is itself configured (gain profile, delay, metrics); evaluating a policy away from its design operating point measures a counterfactual robot, and bounded scores saturate once binary gates pass, losing all sensitivity. Instruments need the same operating-point discipline as deployments and graded outputs to retain gradient.

    Applies when

    • an automated smoke loop contradicts the official battery
    • early-stop scores freeze while graded metrics still improve
    • lineages with non-default deployment gain/delay profiles
    “watcher (kd1.0 口径) 3300 起 1/3 恐慌被 kd1.2 正式扫描证伪为考纲外假象 —— 工作点评测口径教训: watch_ckpt 缺 --kd-scale 透传 (待补), 冒烟只当存活哨兵。… watcher score 饱和误停 @1050(八门全过恒 0/3 时 score 对漂移改善盲, s1f 课文三现)”
    train/README.md § omni_s2e_fric (watcher 恐慌被证伪) / omni_s2e_pd (500 臂)
  • Raising a command bucket's share does not strengthen its per-state gradient - it only starves the other modesbucket-share-is-not-a-gradient-lever
    Mechanism understoodomnicurriculumcurriculumreward-shapingdomain-randomization

    When a skill is not learning, first prove its per-state signal is nonzero (ignore-floor and probe checks); only rebalance sampling shares to fix genuine sample starvation, and account the regression risk to the diluted modes before doing it.

    Symptom

    Sidewalk was not learning, and the reflex proposal was to give the side bucket a larger share of sampled commands.

    Context

    The C4-redo3 rung explicitly kept the 20/40/20/20 bucket (stand/forward/turn/side) with the reasoning written out: PPO computes advantages per state, so bucket proportion does not change the per-state gradient of side states; at 4096 envs x 20% x 24 steps the rollout already contained ~19.7k sidewalk states - sample count was not the bottleneck. And the cost side was already measured: cutting forward from 60% to 40% had made vx+0.30 die at +400 in an earlier run - more cuts would only collapse it sooner.

    Change

    Bucket proportions held constant across the entire C4 redo series; the actual bottlenecks (metric frame bug, reward variance penalty, exploration form) were pursued instead.

    Outcome

    Sidewalk was eventually fixed with zero bucket changes (feed-forward delivery, +100 iters); forward/turn skills never suffered starvation-induced regressions during the redo series.

    Mechanism

    Policy-gradient credit is assigned per visited state; oversampling a mode multiplies its states in the batch but not the informativeness of each, so if the per-state gradient is ~0 (behavior unreachable or reward indifferent), N times zero is still zero - while the displaced modes genuinely lose data and regress.

    Applies when

    • proposing to oversample a failing task/command mode
    • a majority mode regresses after share rebalancing
    • budgeting env count vs mode share for a multi-skill policy
    “比例不动:PPO 逐状态算优势,桶占比不改变单状态梯度;4096 env × 20% × 24 = 每 rollout 已有 1.97 万个侧走状态,样本数不是瓶颈;而 forward 60%→40% 已实测让 f30 在 +400 处死掉,再加码只会更早塌。”
    train/C_LADDER_RUN.md § 3i. 桶 20/40/20/20 不动(比例不动)
  • Drop the frozen policy into chosen configurations - a squat 2.7 cm lower than the stuck pose stood 52% of the time, the stuck W-sit 0%, and the interpolation between them showed a wall, not a slopeconfiguration-probe-wall-not-slope
    Mechanism understoodrecoveryattributionattributionmeasurementcurriculum

    When a policy is stuck, probe the frozen policy from a grid of hand-placed start configurations, including interpolations between the stuck state and a nearby state it escapes from; one read-only experiment separates height, torque, sampling and configuration and tells you whether to prevent entry or train the exit.

    Symptom

    After R0.3 the policy stood from 62% of starts and never from the W-sit it fell into; height, torque, missing samples and reward were all plausible suspects.

    Context

    A read-only probe placed the R0.3 policy directly into specified configurations. Squats (hip, knee, ankle) = (-.65,-1.3,-.65) stood 100%, (-1.0,-2.0,-1.0) 89.8%, (-1.2,-2.4,-1.2) at 0.176 m 52.3%; the measured W-sit at 0.203 m 0.0%; the account-(3) hand-over state (146 deg tilt) 34.4%; linear interpolations from the W-sit toward the squat at 25/50/75% stood 0.0/0.0/3.1%. The squat family's quasi-static torque is 16% of the limits, and the W-sit was visited ~9 s per episode in training. FK showed the squat family (-a,-2a,-a) keeps the torso vertical, the feet flat and the COM over the feet all the way from 0.146 m to 0.384 m.

    Change

    Height, torque and sampling were eliminated in one experiment; the next rungs targeted entering the W-sit (foot placement) instead of escaping it, and seeding the dead point itself was ruled out because it was already visited every episode.

    Outcome

    Pure configuration: the W-sit (hips externally rotated +/-47 deg, knees folded 110 deg, shins flat, feet beside the body) is a different place from the sagittal squat (feet flat under the COM). The policy's standing skill was bound to a narrow sagittal family, and the wall was confirmed by the interpolation. The foot-placement rungs that followed took prone from 0/159 to 158/159.

    Mechanism

    A learned skill covers the neighbourhood of the states it succeeded from; a start state outside that neighbourhood fails regardless of height or torque, and an interpolation that stays at zero until close to a working state shows the boundary is sharp.

    Applies when

    • a policy stalls in a specific posture and several causes are plausible
    • deciding between reverse-curriculum seeding and entry-prevention shaping
    • a feasibility account says a path exists but the policy does not take it
    “**决定性对比:比死点矮 2.7 cm 的蹲姿站立 52.3%,死点 0.0%。** 所以不是高度、 不是力矩(蹲姿族准静态力矩膝 1.96/12、踝 1.24/17,只占 16%)、也不是训练采样 (死点每局被访问 ~9 s)。**是纯位形问题** … 插值实验进一步显示这**不是坡是墙** —— 走到 75% 仍只有 3.1%”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §15 死点位形实验(只读探针,同一个 R0.3 策略放进指定位形)
  • Audit which joints your imitation term constrains - a task that needs deviation is fighting the referenceimitation-term-scope-audit
    Mechanism understoodomnireward-shapingreward-shapingcurriculum

    List which joints your imitation/deviation terms actually constrain and check the new skill's required motion against that list; for balance-coupled joints deliver references as feed-forward residuals, not absolute-position targets - and never assume "reference = 0" is neutral.

    Symptom

    Sidewalk would not learn despite a dedicated tracking reward; meanwhile the gait-shaping imitation term (joint_pos_ref) computed its error norm over ALL 12 joints while its reference covered only the 6 sagittal joints - roll/yaw reference was constantly 0.

    Context

    Two prior generations had shown the forward gait itself was taught by joint_pos_ref, not discovered by PPO (v6 halved the shaping and swing height collapsed 35 mm -> 4 mm). So the reference is load-bearing - but sidewalk requires hip_roll to deviate from nominal, and the all-joints norm punished exactly that deviation: "一边悬赏一边罚过程" (posting a bounty while punishing the process). A follow-up experiment (C4-redo3, free_roll=True releasing the 4 roll joints from the norm) raised the regularization headroom 6x -> 27x yet sidewalk stayed flat and released hip_roll wandered, killing other skills - net negative, withdrawn. A --roll-absolute probe showed the converse failure: pinning roll to a clock-driven absolute trajectory drove tilt 6.9 -> 13.7 deg. Conclusion recorded: absolute-position imitation cannot teach actions that must be superimposed on state feedback.

    Change

    The audit reframed the problem: neither punishing roll deviation nor freeing roll nor absolute roll tracking works; the reference for a balance-coupled joint must be delivered as feed-forward under the policy's residual control (see feedforward-for-phase-locked-skills).

    Outcome

    free_roll rung: joint_pos_ref term rose 0.887 -> 1.104 (release confirmed effective) but vy stayed flat; regularization hypothesis eliminated by experiment.

    Mechanism

    An imitation error norm defines a cage: joints inside it are pulled to the reference in absolute position, so any skill requiring systematic deviation is taxed per step; but joints carrying active balance cannot follow absolute references either, since their correct position depends on state. The scope and the delivery mechanism of the reference are therefore design decisions per joint, not defaults.

    Applies when

    • adding a skill that moves joints your reference sets to zero/nominal
    • an imitation or deviation penalty coexists with a new tracking reward
    • considering releasing joints from a shaping term mid-lineage
    “前进步态也不是 PPO 自己发现的,是 joint_pos_ref 教出来的(v6 砍半塑形 → 抬脚 35 mm 塌到 4 mm…)。而 ref_joint_offset 原本只写 6 个矢状面关节,roll/yaw 参考恒 0 —— 侧走既没被教,roll 一偏离 nominal 反被 joint_pos_ref 扣分。一边悬赏一边罚过程。”
    train/C_LADDER_RUN.md § 3e. 为什么首战 FAIL / 3i. 解锁笼子
  • The get-up kept getting faster because standing earlier paid more every step - lowering torque authority barely slowed it, and only zeroing the standing income for the first 3 s moved the pace into the design bandper-step-income-drives-speed-time-gate
    Mechanism understoodrecoveryreward-shapingreward-shapingcurriculum

    When a skill is too fast, find the term that pays for finishing early and gate that income by time; keep the "get into position" term ungated so the policy does not learn to wait, use a ramp instead of a cliff, and confirm with a paired same-level experiment that the drift is motivational before changing it.

    Symptom

    The user judged the get-up too fast (Isaac medians about 0.7-1.6 s) and suspected path dependence: the policy seemed to get faster the longer it trained.

    Context

    Lowering the beta authority 0.40 -> 0.30 cut impact but moved supine only 1.70 -> 2.00 s: coordination-limited, not torque-limited. A paired experiment inside one beta level (checkpoint 15,600 vs 18,499, +2,900 iterations, same ruler) measured the drift: get-up medians -7 to -10%. The spec concluded the motive, not the path, was the cause - per-step standing income pays for every early step, and any lineage (even one from scratch) races toward the fastest solution inside its constraints.

    Change

    V2.5: the standing income (base_height, stand_pose, still, feet_on_ground) multiplied by w(t) = clamp(t/3 s, 0, 1); upright deliberately NOT gated, so righting and sitting up early still pay and the policy is not taught to lie flat and wait; a ramp, not a step. V2.5b: zero before t0 = 3 s, then a 1 s ramp.

    Outcome

    V2.5: Isaac 100%, get-up +17-43% slower, MuJoCo 100/100/98/100% (the best cross-simulator reading yet), still short of the 3.5-4.5 s design band - a linear ramp only discounts early income. V2.5b: MuJoCo supine 2.04 -> 4.10 s and prone 3.18 -> 4.04 s, inside the band; the Isaac pace barely moved (a lineage habit on a gradient-free plateau). Later the zero gate proved harmful when trained from scratch (curriculum-history-is-part-of-the-product) and in the V3.1 lineage (time-gate-vs-wide-stance-retire-the-fix).

    Mechanism

    Constraints on authority or velocity change how the fastest solution looks; the time structure of the task income decides how fast the fastest solution is.

    Applies when

    • a policy is faster or more aggressive than wanted and constraints do not slow it
    • progress-style rewards pay every step spent at the goal
    • performance drifts faster with more training at fixed settings
    “**V2.5 机制(唯一)**:站立收入(base_height/stand_pose/still/feet_on_ground) 乘时间斜坡 w(t)=clamp(t/T_gate,0,1),T_gate=3.0 s;**upright 刻意不门控** (翻正/坐直早期照常拿钱,防"躺平等门开" … **用户假设量化 证实:逐步计酬动机在 β 包络内持续压缩时间,约束挡不住动机 —— V2.5 动机层 修法为正解。**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §39 V2.5 预注册 / §39 补 配对实验
  • Every rung gets written stop criteria - hit any one, stop; tightened when priors say results should come fastpreregistered-stop-criteria-per-rung
    Replicatedomnitraining-runprocessgate-battery

    Freeze per-rung stop criteria (old-skill floors, new-skill deadline, oscillation signature, known pathology signatures) before training, stop on first hit - and shorten the deadline in proportion to how fast your mechanism says results should appear.

    Symptom

    Continued training past the point of degradation had previously destroyed capabilities (C1 trained away the root's backward skill); without hard stop rules, sunk-cost reasoning keeps runs alive too long.

    Context

    Standard rung protocol - eval_c_matrix every 100 iters at 5 seeds with a fixed criteria list, frozen before training ("开训前写死,事后不许改"): (1) stand or vx+0.15 at <=3/5 for two consecutive points; (2) vx+0.15 tracking <70% for two consecutive points (vs recorded root baseline); (3) oscillation signature - survival repeatedly crossing zero between adjacent checkpoints, stop on first occurrence, do not wait for confirmation; (4) new-skill-no-progress deadline (e.g. wz tracking <30% or still same-signed after iter 400); (5) a known lineage pathology signature (s1c arm-B: scatter -> half-recover -> collapse).

    Change

    Stop budgets scale with prior knowledge: when C4-redo4's feed-forward was already proven open-loop, the no-progress deadline was tightened from +500 to +100 iters ("前馈是开环就给 71~96% 的 … 若 +100 还没有 … 不值得再烧").

    Outcome

    Multiple rungs were stopped exactly on criterion (C4-redo criterion 4 at iter 1200; redo2/redo3 at absolute 900), converting each into a clean hypothesis test instead of a drifting run; no rung burned its full cap on a dead hypothesis.

    Mechanism

    Degradation of retained skills is a lineage-level injury that later training often cannot undo, so detection latency is capital loss; and a stop rule written before the run cannot be bent by hope. Tightening deadlines when priors predict immediate results converts "no result yet" into evidence against the mechanism.

    Applies when

    • starting any resumed/curriculum training rung
    • deciding whether to keep training a run that shows early regression
    • a mechanism-backed change should produce results immediately
    “每 100 iter eval_c_matrix.py --seeds 5,命中任一即停 … 存活在相邻 checkpoint 间反复跨越 0(出现即停)… 收紧版:绝对 iter 900(= +200)时 vy±0.10 跟踪仍 <30% 即停”
    train/C_LADDER_RUN.md § 3h. 预注册停梯判据(开训前写死,事后不许改)
  • Score left and right separately - averages hide chirality breaking that mirror augmentation does not preventchirality-scored-separately
    Replicatedomnigate-batterygate-batteryreal-acceptance

    Report every mirrored skill as two numbers with an explicit gap budget; never accept an average, and never assume augmentation guarantees symmetry - measure it per lineage and treat breakage as hard to reverse.

    Symptom

    Policies developed quantified left/right asymmetry (e.g. C2-700 turned right at 82% but left at 67% - a 15 pp gap; push tolerance 40/40 symmetric on the root vs 17/40 on a deep-trained descendant), and averaged metrics would have reported healthy midpoints.

    Context

    The repo had policy-level symmetry-breaking evidence strong enough to make separate scoring a battery rule: "左右必须分开打分 … 平均 vy 跟踪会把它掩盖". Notably, chirality broke and never recovered even though mirror augmentation (command-level mirror_prob 0.5) was on the whole time - augmentation reduced but did not prevent asymmetry, and once broken it stayed broken through subsequent rungs. PASS conditions therefore carried explicit symmetry budgets (left/right tracking gap <=10 pp), and sim's predicted asymmetry (700: right faster than left) was flagged for direct real-robot timing confirmation.

    Change

    Battery rule: every directional skill reports left and right (CW/CCW) as separate rows with a max-gap budget; mirror augmentation treated as mitigation, not proof of symmetry.

    Outcome

    The 700-vs-A800 asymmetry gap (15 pp vs 7 pp) became a first-class selection criterion; C4 product shipped with a measured 5 pp gap.

    Mechanism

    Averaging over mirrored conditions cancels antisymmetric error exactly where it matters; and symmetry lost during training is a lineage injury (like plasticity loss) that later rungs do not spontaneously heal, so it must be gated, not assumed.

    Applies when

    • evaluating turn/sidewalk/push-recovery or any mirrored skill
    • relying on mirror/symmetry augmentation
    • selecting between checkpoints with similar average scores
    “左右必须分开打分(left/right lateral、CW/CCW turn 各自一行)—— 本仓已有 policy-level symmetry breaking 的量化证据,平均 vy 跟踪会把它掩盖。”
    train/C_LADDER_RUN.md § 5. 固定验收矩阵 (左右分开打分)
  • Narrowing the speed range to stop high-speed falls entrenched crouch-shuffling - judge gait quality at the speed that demands a gaitlow-speed-commands-reward-dragging
    Mechanism understoodwalkcurriculumcurriculumreward-shapinggate-battery

    Set command ranges to include speeds that physically demand the target behavior, and evaluate behavior-quality gates at those speeds; when a restriction is added to suppress a failure, check what new optimum it creates at the remaining commands.

    Symptom

    After the command range was narrowed to (0.15, 0.35) m/s (to treat walk_v1's 134% overspeed and 8.3 s fall at 0.5), the policy settled into crouched foot-dragging; tracking rose monotonically with speed (63% at cmd 0.2, 76% at 0.3, 87% at 0.45), showing low speeds were where the degenerate gait was optimal.

    Context

    The narrowing advice was the author's own and is retracted in the file: it treated the symptom (falls at speed) while reinforcing the root cause (at 0.15-0.35 m/s, shuffling in a crouch is globally optimal - the Froude number is so low that even humans would not lift their feet). A zero-cost experiment confirmed the flip side: at cmd 0.5 the same policy met BOTH tracking (81%) and clearance (23.0/23.2 mm) standards.

    Change

    Speed range widened back toward (0.15, 0.5) - upper bound deliberately slightly above the mechanically feasible ~0.44 m/s so the policy finds the boundary itself; acceptance re-pointed: gait-quality criteria (tracking, clearance) judged at 0.45-0.5 m/s, low speed kept only as a survival check.

    Outcome

    v4 -> v6 progression under the widened range delivered 87% tracking with 34 mm clearance; the "low command = drag" account was confirmed by the monotone tracking-vs-speed curve.

    Mechanism

    Command distribution is part of the reward: physics prices gaits per speed, and at very low speed the energetic optimum is no swing phase at all. Restricting training to that regime makes the degenerate gait the correct answer to the posed problem - and grading a gait at a speed that does not require stepping measures nothing.

    Applies when

    • a gait degenerates after a command-range restriction
    • quality metrics improve monotonically toward the range boundary
    • writing acceptance criteria for gait quality vs survival
    “现在看那个建议可能起了反作用:0.15~0.35 m/s 下蹲着蹭就是全局最优,抬腿反而亏。收窄治的是"高速摔倒"的症状,却强化了拖地的病根。… 验收标准里的 cmd 0.2 本身就是拖地速度(Froude 数极低,人在那个速度下也不抬脚)。accept_v2 应把速度跟踪与 clearance 的判定点改到 0.45~0.5 m/s”
    train/WALK_DIAGNOSIS.md § ② 放宽速度区间 / ① 零成本实验
  • Under continuous 3-axis uniform sampling, pure straight-line walking is a zero-measure event the policy never trainedzero-measure-commands-need-mode-sampling
    Mechanism understoodomnicurriculumcurriculumobservation-honestygate-battery

    Enumerate the exact command points users will actually issue (straight, stop, in-place turn) and give each explicit probability mass via mode sampling with off-axes pinned to zero - never assume a continuous sampler covers its measure-zero subsets.

    Symptom

    "The robot drifts even in sim when told to walk straight" persisted across reward tunings - because with commands drawn as vx in [0.15,0.5] x vy ~ U(+/-0.2) x wz ~ U(+/-0.6), the event vy=0 AND wz=0 has probability zero: pure straight-line walking was never sampled even once.

    Context

    Restart evidence item #3: "纯直行是零测度点 … 'sim 里直行就漂'是分布的 必然,不是 reward 没调好" - the drift metric was legitimately drowned by commanded turning (v11's own comment self-documented this). The structural fix is discrete mode sampling: a custom ModeVelocityCommand that first draws a mode by share (stand/forward/back/turn/side/mixed), then draws values only on that mode's axes with all others pinned to exact zero - which is also what preserves single-variable discipline in the C ladder (native 3-axis uniform "采不出'离散模式桶' … 把 C1~C4 的单变量纪律直接毁掉"). The mixed mode later got an ellipsoid constraint rather than a cube for the same reason in reverse - corner combinations of a cube are unrepresentative extremes.

    Change

    Command generation moved from independent per-axis uniforms to mode-bucket sampling with pinned-zero off-axes (plus 20% rel_standing); acceptance likewise evaluates per mode.

    Outcome

    Straight-line behavior became a trained, testable mode instead of a measure-zero hope; the C ladder could add one mode per rung with provable isolation.

    Mechanism

    A policy optimizes expected reward under the command distribution; events of probability zero contribute nothing to the objective, so exact-zero-command behaviors (straight walk, stand, in-place turn) are only learned if the sampler gives them mass. Product-of-uniforms distributions concentrate mass on mixtures and give none to the pure behaviors users actually command.

    Applies when

    • a "simple" command (straight, stop) underperforms mixtures in sim
    • designing command distributions for velocity-tracking tasks
    • a ladder needs per-mode isolation for attribution
    “纯直行是零测度点:最终 command 为 vx∈[0.15,0.5] × vy∈U(±0.2) × wz∈U(±0.6) 连续均匀,vy=0∧wz=0 从未被专门采样 —— "sim 里直行就漂"是分布的必然,不是 reward 没调好 … Isaac 原生 UniformVelocityCommand 是三轴各自 uniform,采不出"离散模式桶"”
    train/OMNI_V0_SPEC.md § 0. 为什么从零 (3) / 三件前置 (1)
  • Before training a one-leg stand the spec named the cheapest cheats - hopping on the support foot, a raised foot resting unloaded, a leg tripod - and gave each a countermeasure and a gate; one still appeared and was caught by exactly those gatesenumerate-cheapest-cheats-before-training
    Observed onceonelegreward-shapingreward-shapinggate-battery

    Before training, list the cheapest behaviours that would satisfy each reward term without doing the task, give each a countermeasure in the reward and a gate in acceptance, and prove the intended behaviour is reachable with a probe - then treat any gate the policy games as evidence about the reward, not the gate.

    Symptom

    The literature's single-leg benchmark reports eight state-of-the-art general policies holding a clean one-leg stand 0 times out of 90 - they survive by sneaking steps and hops - so the task's first adversary was the policy's own cheating.

    Context

    The spec's shape self-check ("what is the zero-cost option?") listed, for the one-foot bucket: the cheapest cheat, a foot resting on the ground without load, countered by a 5 N contact threshold plus positive swing income; the second cheapest, small hops on the support foot to reset balance, countered by a continuous support-air penalty plus a gate of zero support-foot flight segments. The probe that preceded training had already seen a third: early low-lift postures "survived" by pressing the swing foot at 78-95 N, a leg tripod, removed by folding the shank back. The two-foot bucket was checked too: its zero-cost behaviour is ordinary standing, with no odd base state.

    Change

    Countermeasures and gates written before training: swing-contact and support-air penalties, gate 2 (zero swing-foot contact frames above 5 N), gate 3 (zero support-foot flight segments).

    Outcome

    The first run still found the unloaded-foot cheat (a binary reward band gave it no gradient to lift) - and it was caught, by the contact gates and the cross-simulator comparison, not discovered on hardware. The retrained V0 passed all gates 40/40, including zero support-foot flight after the flight detector was corrected.

    Mechanism

    A policy optimizes the reward, not the intent; the cheapest behaviours that satisfy the reward are predictable from the reward's structure, and a gate written for each before training turns a silent cheat into a failed row.

    Applies when

    • designing rewards for balance, contact or "hold still" tasks
    • benchmark policies are known to cheat the task
    • writing acceptance gates for a new skill
    “文献里 8 个 SOTA 通用策略在单脚站基准上 0/90 干净保持, 全靠偷步偷跳活命,这是本任务的第一反作弊对象 … 单脚桶下最便宜的作弊是"脚虚放地上不受力"——接触判定 >5 N 力阈(沿用),配 swing_height_band 正收入拉开。 … 第二便宜是"支撑脚小跳重置"——support_air_penalty 连续罚 + 验收门支撑脚腾空段=0 双保险。”
    git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §0 目标口径 / §5 形状自检(零成本选项是什么)
  • The first real-robot get-up was "very violent, kicking on the floor, dangerous" - a sim-perfect policy with no reason to be slow, unbounded absolute targets, no domain randomization and a rate limiter that filtered nothing; the task was restated as "safe, slow, transferable"first-real-get-up-violent-stage-one-policy
    Observed oncerecoveryreal-deployreal-acceptanceactuator-modelingattribution

    Do not put a get-up policy on hardware until its action is bounded (hard bound or state-anchored targets), smoothed, randomized and tested at the real pipeline's latency, and say explicitly that the task is "safe, slow and transferable" - a simulation-perfect policy optimizes only "gets up".

    Symptom

    On 2026-08-09 the user ran a V0-lineage recovery policy on the real robot and stopped it: very violent, kicking on the floor, dangerous. The planned next rung (a heavier torque_headroom) was never started.

    Context

    The spec had pre-registered that R0/R1 products stay in simulation and that the real-robot precondition was the R3 smoothing rungs plus a bridge-slew check plus a hanging protocol; the robustness (DR) rungs had not run. In simulation the policy passed 100% with a get-up of about a second. Which ONNX, which gain profile and whether a torque/joint log existed were left "to be recorded later" and never were.

    Change

    The V0 ladder was stopped at its best product (R3.1, sim only) and a re-rooting proposal was put to the user. The spec's four-layer account: style (the reward pays for standing early and nothing pays for slowness - HumanUP's "Stage I" get-up, "fast but unsafe ... infeasible for real-world deployment"); impact (full-range absolute targets with no hard bound, raw |a| up to 4.77, action saturation 100%, a single-step change of 0.306 saturating hip_pitch); transfer (zero DR, friction pinned at 1.0, the learned leg bracing); link (the bridge's RL slew equals vel_limit, 0.2-0.66 rad per step, while the real pipeline has 1-2 steps of time-varying latency and acceptance ran at delay 0).

    Outcome

    The line was re-rooted twice (training-side rate limit, then the beta-anchored action space) and gained a hang protocol before the next real attempt; on 08-11 a beta-anchored policy produced the line's first real get-up.

    Mechanism

    A task reward that pays for standing early selects the fastest feasible get-up; with absolute full-range targets every large target jump is a torque impulse bounded only by the clip; zero DR and braced-leg solutions do not transfer; and a limiter set at the velocity limit does nothing at 50 Hz.

    Conflicts

    The four layers are the spec's reconstruction from simulation probes and the literature; the real run's policy file, gain profile and log were never recorded, so no layer was confirmed against hardware data.

    Applies when

    • a first hardware trial of a high-effort skill is being scheduled
    • sim success is high but the policy saturates actions or torques
    • pre-registered hardware preconditions are not all met
    “用户真机反馈:**非常猛、地上乱踢、危险**,叫停(R3.3 torque_headroom 加档已选型 weight −0.5→−1.5,未启动)。真机细节(哪个 onnx、什么档、有无 τ/q log)**待补记** … 任务从"能起来"变成 **"安全、慢、可迁移"** … **链路层**:桥层 slew RL 档 = vel_limit(10/20/33 rad/s ≈ 每拍 0.2~0.66 rad), 对 recovery 形同虚设;真机 1~2 拍时变延迟,验收默认 delay 0。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 真机叫停与换根判决(2026-08-09)
  • Choose the fork root by which candidate's shortfalls are recoverable, not by headline scorefork-root-recoverable-shortfall
    Mechanism understoodomnifork-selectionfork-selectionprocess

    When picking a checkpoint to fork from, rank candidates by whether their weaknesses are trainable-back, not by current headline metrics; prefer the candidate whose deficits the upcoming training directly pays for.

    Symptom

    Multiple candidate checkpoints for the omni-command ladder root, each best at something different: fric-3000 had the best tracking precision (vx 88-91%) and hardened plant robustness; s1e-500 had lower precision (vx 82%) but was the only candidate that could still walk backward.

    Context

    Root selection ran as a data probe, not a preference vote: 6 candidates x 8 out-of-distribution omni commands x 20 seeds = 960 cells (probe_omni_0808.json). s1e-500 @pw1.0 survived 20/20 in all eight conditions including backward at 67% tracking; the deep-trained fric lineage scored backward 0-3/20 despite better forward precision.

    Change

    Decision criterion made explicit: list what each candidate exclusively wins at, then ask which of those wins the loser could train back. fric-3000's exclusive wins (precision, plant robustness) are both retrainable - precision is directly optimized by the reward, plant hardening is a planned later pass. s1e-500's exclusive wins (backward plasticity 20/20 vs 2/20, disturbance margin 159/160 vs 125/160, push chirality symmetry 40/40 vs 17/40) had all been shown unrecoverable - push-level rungs failed twice, chirality never recovered even with mirror augmentation on. Root = s1e-500.

    Outcome

    s1e-500 carried the whole C ladder; its backward skill was preserved through C2/C4 gates (regress budget <=2/20 enforced), and the final C4 product passed a 260-cell battery at 20/20 everywhere.

    Mechanism

    Training can re-earn anything the objective directly pays for, but capabilities that earlier training destroyed and never restored (plasticity, symmetry, robustness margins) are empirically one-way doors. The information-bearing comparison is therefore recoverability of each candidate's deficit, which the team stated as "独占项的可恢复性正好相反 —— 这就是判据" (the exclusive items' recoverability is exactly opposite - that is the criterion).

    Applies when

    • selecting a resume/fork root among several checkpoints
    • one candidate is more precise but another retains a skill the rest lost
    • planning a task-extension ladder from an existing lineage
    “fric-3000 赢在精度(vx 88~91%…)与 plant 鲁棒性 → 两样都训得回来…;s1e-500 赢在可塑性(C1 20/20 vs 2/20)、抗扰余量(159/160 vs 125/160)、手性对称(推 ±6 N·s 40/40 vs 17/40)→ 三样都训不回来 … 独占项的可恢复性正好相反 —— 这就是判据。”
    train/C_LADDER_RUN.md § 0. 为什么根是 s1e-500(数据,不是偏好)
  • Every gate on a penalty is an exit - to stop paying a stance tax the policy parked 2 deg outside a 30 deg uprightness gate (lunging), and, from scratch, just under a height gate (crouching); a positive band and an always-on guard fixed bothpenalty-gate-is-an-escape-hatch
    Replicatedrecoveryreward-shapingreward-shapinggate-battery

    Never gate a penalty on a state the policy can leave by getting worse; use always-on guards for what must never happen and positive, gated bands for what you want, and count every gate on a penalty as one more escape route to check in the logs.

    Symptom

    V2.9 added stance_width_task = relu(0.34 m - foot spacing) x standing gates, weight -10. Three checkpoints scored 0% on acceptance: standing height reached, feet on the ground, angular rate low, but the torso leaned 32.5/32.3/31.9 deg in a fore-aft lunge. The training dashboard read "tax paid off, base_height at full value".

    Context

    The penalty was gated by uprightness (tilt < 30 deg) and standing height; its tax had no time gate (500 steps x -1.65) while the standing income sat behind a 3 s zero gate (~325 steps), so leaning just past 30 deg lost a little gated income and saved the whole tax. base_height has no upright gate, so the lunge still collected it. The first metric also measured full horizontal spacing, so a staggered lunge counted as "wide".

    Change

    Two laws written down: a penalty may only carry gates the policy cannot escape by getting worse (make it an always-on guard) or it becomes a positive band ("not earned" is not "escaped"); and width is measured laterally in the base yaw frame. From scratch (V3.1 P1) with the lateral metric but the same gates, the policy parked just under the height gate instead (base_height 1.176/1.5, h ~ 0.30 m against a 0.3264 gate), the beta curriculum never advanced in 1,700 iterations, and the run was stopped early. P1b flipped the penalty into a positive band +2.0 x clamp(lateral / 0.34) x standing gates; P1c added yaw_guard = -5 x relu(|hip_yaw| - 30 deg), always on, no gate, no exemption.

    Outcome

    P1b: lateral stance 0.364 m, all four categories 100%, MuJoCo mu 1.0/0.4 both 100%. P1c: all six acceptance criteria passed for the first time on the line (hip-yaw saturation 1.2%), the guard's tax converging to -0.016 (almost never touched).

    Mechanism

    A gated tax that is not paid is saved, so the policy moves to the cheapest state just outside the gate; a gated income that is not earned is simply lost, so a positive band has no exit. HoST's style penalties are ungated or binary - the same law seen from the other side.

    Applies when

    • adding a penalty multiplied by an uprightness, height, phase or contact gate
    • a policy settles just beyond a gate threshold
    • training metrics look paid-up while acceptance collapses
    “**罚项的门 = 策略的逃生门**。带直立门的负项可以靠"变得更差"(倾出门外) 全时免税;HoST 的 style 罚全部无门控/二值恰是同一律的反面实证。修律: 负项只许挂"变差逃不掉"的门(上限护栏 always-on),或改正向 band (收不到 ≠ 逃掉)。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §46 定案(血统内第四败 + 两条新律)
  • The walk phase machine structurally cannot express flight - rebuild the representation for duty < 0.5, teach flight with a mask tax, never a cliff bountyphase-machine-structural-limits
    Mechanism understoodrunreward-shapingreward-shapingcurriculumgate-battery

    When a new gait changes the contact pattern's structure, audit whether the phase/mask representation can express it and rebuild the representation if not; teach the new contact pattern with graded mask-mismatch pressure and count it in acceptance with artifact-proof definitions (minimum segment length), never with cliff bounties.

    Symptom

    Running requires both feet airborne, but the walk-era phase machine switches legs by the sign of sin(phase) - with duty < 0.5 the two swing windows must OVERLAP during flight, which a sign-switching representation cannot express at all.

    Context

    The run phase machine was re-architected rather than patched: per-leg phases (left = phi, right = phi+0.5 mod 1) with leg_phase < duty defining stance, aligned to the walk sin convention at duty=0.5 so the machines agree where their domains overlap. Flight is taught by the SAME mechanism that once cured foot-dragging, direction reversed: in the two planned flight windows the contact mask is (0,0) and feet_contact_number_duty charges -0.3 per foot still on the ground - a mild ~0.16/step tax, deliberately NOT a cliff: "悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文". The reference shape (half-sine bump over swing progress) is zero at window boundaries by construction, eliminating the clearing-window that the C4 probe measured to cost 13-19% on non-sinusoidal references. The shape self-check ("零代价选项是什么") was run on three behaviors: standing pays ref everywhere (known cmd=0 stepping risk, booked), walking pays only the flight-window tax, proper running collects full marks.

    Change

    New rewards.py run section (leg_phase_duty / stance_mask_duty / ref_run / clearance_run / contact_duty) with the walk versions untouched byte-for-byte; flight acceptance metric defined with a segment-length floor (>=40 ms to count) so numeric contact flicker cannot fake flight.

    Outcome

    Flight became expressible and taught by a calibrated mild pressure; the walk lineage's phase code stayed frozen as its own contract.

    Mechanism

    A phase representation defines which contact patterns exist in the reward's vocabulary; duty cycling below 0.5 introduces states (double-flight) outside a half-period sign convention's language, so no weight tuning can teach them. And rare desirable events taught by cliff-shaped bounties invite hacks (jumping in place); the graded mask tax prices the planned pattern without creating a jackpot.

    Applies when

    • extending a walking stack to running/jumping (duty < 0.5)
    • a desired contact pattern never appears despite reward increases
    • defining flight/contact acceptance metrics
    “duty<0.5 时摆动窗 (1−duty)T > T/2,两腿摆动窗在腾空段重叠 —— sin 符号切腿的机制结构上表达不了"双脚同时在空中"。… feet_contact_number_duty 对"还踩着地"持续 −0.3/脚 —— 与 walk 治拖地同一机制,方向相反。不腾空的税 ~0.16/步 … 梯度温和不构成悬崖(悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文)。”
    train/RUN_V0_SPEC.md § 5. 相位机设计
  • The best checkpoint to SHIP is not the best checkpoint to CONTINUE FROM - maturity is capital against adaptation shockroot-maturity-vs-product-quality
    Mechanism understoodomnifork-selectionfork-selectioncurriculumgate-battery

    Decide shipping points and fork roots separately: gates rank products, but a root candidate must prove itself by surviving a continuation under the next rung's shift (dual-arm if in doubt) - and prefer the more-trained point as root when product metrics conflict with maturity.

    Symptom

    A band re-audit found s1e-300 beat the incumbent root s1e-500 on nearly every quality gate (stepping 19/20 vs 13/20 with historically-best 26.9 mm swing, speed gate 14/20 vs 2/20, heading 26 vs 54 deg/20 s) - suggesting the root had been mis-picked and the younger point should take over.

    Context

    The dual-arm control settled it the other way: continuing the S2 PD rung from s1e-500 adapted smoothly (3/3 smoke throughout), while the b300 control arm (same config, from s1e-300) fell into a survival valley under the PD shock (+100 iters: 1/3 -> 0/3), never climbed out within budget, and its 800-iter product scored 13/20 survival - eliminated. Verdict: "幼年点自身指标再好也扛不住新 DR 适应冲击, 成熟度是本钱,s1e-500 根被数据背书" - a young point's own metrics, however good, do not survive new-DR adaptation shock; maturity is capital. The audit still yielded value: the band scan (200-1000, per-100) mapped the lineage's arc (200 dragging -> 300 peak -> 400+ decay -> 900+ drift blowout), and 300 remains the better PRODUCT answer for shipping-as-is questions.

    Change

    Selection doctrine split into two questions with different answers: best-product point (quality gates at the point itself) vs best-root point (survives adaptation shocks; more training age = more capital), each decided by its own evidence - and root claims settled by a dual-arm continuation test, not by point metrics.

    Outcome

    s1e-500 kept the root role with data behind it; the S2e ladder built on it passed rung after rung, while the b300 line was closed at the cost of one control arm.

    Mechanism

    Early checkpoints sit near sharp optima with less accumulated robustness structure; their headline metrics reflect the narrow training distribution, not resilience to distribution shifts. A continuation rung is itself a distribution shift, so the root property being selected for is shock tolerance - observable only by actually continuing, never by static gates.

    Applies when

    • a younger checkpoint outscores the current root on quality gates
    • choosing the base for a robustification or command ladder
    • a continuation run stalls in an early survival valley
    “b300 对照臂 … PD 冲击下存活谷(+100 起 1/3→0/3),预算尽未爬出,800 档 20-seed 存活 13/20 出局——幼年点自身指标再好也扛不住新 DR 适应冲击,成熟度是本钱,s1e-500 根被数据背书”
    train/README.md § omni_s2e_pd (b300 对照臂) / s1e 选点重审
  • Compare the achieved reward to the computed ignore-floor to tell "never learned" from "learned but unprofitable"ignore-floor-diagnosis
    Mechanism understoodomniattributionattributionreward-shaping

    For any skill that trains flat, compute the reward the null policy would earn on that term; achieved==floor means the behavior never paid out (find why: exploration, reward observability, or feasibility) - do not tune weights first.

    Symptom

    C4 sidewalk failed on both arms; the question was whether the policy had found sidewalk and rejected it as unprofitable, or never found it at all - two diagnoses with opposite fixes.

    Context

    The theoretical value of the sidewalk tracking term for a policy that completely ignores the command was computable from the command distribution: 0.189. Trained final values landed at 0.187 (arm A) and 0.204 (arm B) - sitting exactly on the ignore-floor - while the honest balance at the optimum actually favored sidewalking (net +0.88/step inside the side bucket). Later the same arithmetic closed the whole saga: for the true reward landscape, doing real sidewalk scored 0.153 vs 0.944 for ignoring - the policy's refusal "是理性最优,不是探索失败" (rational optimum, not exploration failure) under one hypothesis, and under the final measurement-corrected account the policy had "每一次都在 做理性选择" (made the rational choice every time).

    Change

    Diagnostic rule adopted: compute the ignore-floor for the new term; if the achieved value sits on it, the behavior was never expressed in useful volume (or the reward cannot distinguish it - check both); if the achieved value is above floor but the behavior is absent at deployment, the policy sampled it and priced it out - then the reward balance, not exploration, is the lever.

    Outcome

    Correctly identified that PPO had not merely under-valued sidewalk; each subsequent hypothesis (waveform sign, regularization cage, exploration form, reward kernel) was tested against this floor arithmetic, which kept the search honest through three reversals.

    Mechanism

    Every reward term has a computable value under the null behavior; the achieved-vs-floor gap is a one-number audit of whether the optimizer ever monetized the target behavior. It converts "training failed" into one of two mechanistically distinct states with different fixes.

    Applies when

    • a new skill's tracking reward plateaus early
    • deciding between exploration fixes and reward-weight fixes
    • post-mortem of a failed curriculum rung
    “track_lin_vel_y_exp 训练终值恰好坐在「完全无视指令」的底分上(A 0.187 / B 0.204,理论值 0.189),而终点 balance 明明有利(side 桶内净 +0.88/步)。不是学会了不划算,是根本没学到。”
    train/C_LADDER_RUN.md § 3e. 为什么首战 FAIL

Next page