Skip to content

Training Coach

The coach reads a training run and returns a diagnosis, in-whitelist proposals and an experiment plan — and every claim cites one of the cards below. This is that corpus: 22 methodology rules distilled from a real sim2real programme, and 171 episode cards behind them, each with the sentence in the war history it came from.

Doctrine

A report may cite any of these as doctrine-N.

  1. doctrine-1Contract freeze and fingerprint discipline

    The policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.

    Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).

    Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.

    contract-freeze-and-checkerlegacy-profile-pinningderived-asset-staleness-checkgain-profile-belongs-in-the-stamprecovery-two-policies-and-a-state-machinewalk-recovery-fsm-handoff

  2. doctrine-2Attribution by resolved training params - never eval-override knobs

    Capability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.

    Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).

    Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.

    kd-bandwidth-mu-law-attributiondeploy-knob-attribution-before-retrainingcycle-time-override-is-oodresolved-config-is-source-of-truth

  3. doctrine-3PASS gates become constraints; FAIL gates become objectives

    Once a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.

    Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.

    Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".

    fixed-acceptance-matrix-per-rungpreregistered-stop-criteria-per-rung

  4. doctrine-4One variable per ladder rung - counted against what the checkpoint saw

    A rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.

    Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.

    Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.

    resume-state-dr-auditorthogonal-batch-with-ablation-orderfill-the-missing-factorial-cell

  5. doctrine-5Pre-register risks, readings, and stop criteria before the ladder

    Before a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.

    Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).

    Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.

    preregister-risks-and-fork-readingspreregistered-stop-criteria-per-rungpreregistered-real-expectationsfeasibility-accounts-lock-design-point

  6. doctrine-6Plant parameters are measured, never invented

    Every plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.

    Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).

    Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.

    friction-measured-not-guessedarmature-n2-rotor-inertiatorque-limit-shape-by-measured-peakslatency-lerp-reverse-extrapolationfeasibility-accounts-lock-design-pointsim-api-friction-columnsget-up-feasibility-accounts-before-trainingsingle-support-gain-authority-probe

  7. doctrine-7Sim2sim gate before sim2real - under deployment conditions

    Every checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).

    Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).

    Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.

    sim2sim-gate-before-sim2realeval-plant-honesty-contact-paramspipeline-latency-is-plant-not-drbody-frame-velocity-api-audittorque-penalty-bought-by-leg-bracingtorque-disagreement-between-simulators-unresolved

  8. doctrine-8Observation honesty - the actor's inputs are a hardware contract

    The actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.

    Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).

    Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.

    observation-honesty-critic-onlyhistory-obs-needs-plant-variationreward-observability-limitdeploy-heading-loop-and-align-training

  9. doctrine-9Reward economics are audited in realized currency

    Reward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.

    Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).

    Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.

    realized-contribution-auditreward-cost-of-ignoring-auditgate-new-reward-terms-by-commandignore-floor-diagnosiscalibrate-threshold-between-healthy-and-sickcalibration-threshold-with-withdrawal-clauseinert-reward-term-auditseated-basin-dead-exp-kerneltail-torque-needs-hinge-on-computed-demand

  10. doctrine-10The zero-cost option must be the desired behavior

    For every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.

    Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).

    Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.

    penalize-the-slip-not-the-jointsaturation-cheating-zero-rate-costmoving-gate-42x-stand-taxtermination-closes-degenerate-basinpenalty-gate-is-an-escape-hatchsoft-limit-penalty-charges-nominal-poseunpriced-foot-attitude-is-a-free-variableenumerate-cheapest-cheats-before-trainingbinary-band-reward-fake-touchdown

  11. doctrine-11Measurement discipline: independent referees, signs, distributions

    A disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.

    Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).

    Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.

    independent-referee-for-metric-disputesbody-frame-velocity-api-auditsame-sign-response-is-yaw-biasmedian-hides-bimodal-distributionratio-metrics-need-absolute-checkheading-integral-not-body-ratesame-distribution-reward-comparisonmultiseed-sign-test-for-driftsingle-impulse-recovery-is-chaotic

  12. doctrine-12The deployment pipeline is plant

    Irreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.

    Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).

    Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.

    pipeline-latency-is-plant-not-drpower-scale-hurts-nonforward-axesdeploy-scaling-not-training-equivalentteleop-command-band-per-axislatency-dr-covers-measured-pipelinedeploy-rate-limiter-windupslew-anchor-is-an-integratorbeta-anchored-action-targetpower-derating-cuts-full-range-contract

  13. doctrine-13DR budget is finite; its distribution is the measured support

    Robustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.

    Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).

    Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.

    push-dr-conditional-budget-conservationdr-tail-plant-continuationconstant-value-dr-overfits-margintask-shaping-before-plant-hardeningcom-randomization-forces-leg-spreadcom-dr-rollback-on-symptomthin-dr-judged-by-channel-coveragefriction-priority-re-measured-after-plant-change

  14. doctrine-14Gates measure what hardware feels: posture, margins, stripped assists

    Acceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.

    Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).

    Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.

    task-metrics-vs-posture-metricsstand-gate-posture-not-survivalconstant-value-dr-overfits-marginsuspension-probe-removes-free-stabilizerhip-roll-sum-predicts-lateral-driftchirality-scored-separatelylow-speed-commands-reward-draggingend-state-confusion-matrixfrozen-acceptance-distribution-and-pinned-seedvideo-as-acceptance-recordepisode-length-bounds-what-a-gate-seesremoved-wall-returns-on-hardware

  15. doctrine-15Fork and root selection: recoverability, maturity, frozen rewards

    Choose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.

    Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).

    Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.

    fork-root-recoverable-shortfallroot-maturity-vs-product-qualityfine-tune-reward-change-falsified

  16. doctrine-16Curricula: verified engagement, lineage counters, disease-phase gating

    Automatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.

    Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).

    Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.

    auto-curriculum-engagement-checkcurriculum-counter-lineage-stepsgate-penalties-to-the-disease-phaseaggregate-metrics-mask-subgroup-failurebucket-share-is-not-a-gradient-levercurriculum-criterion-conditioned-on-lagging-categoryper-step-income-drives-speed-time-gatetime-gate-vs-wide-stance-retire-the-fix

  17. doctrine-17Probe before training: feasibility first, hypotheses in tables

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.

    Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).

    Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).

    open-loop-probe-before-reward-tuninghypothesis-table-code-auditperiod-doubling-evidence-racesuspended-test-isolates-actuator-authorityconfiguration-probe-wall-not-slopeprone-dead-end-is-foot-placementdof-vel-penalty-is-not-a-pacing-knobamplitude-cut-falsified-yaw-fix

  18. doctrine-18External advice is recomputed locally; values transfer as ratios

    Every external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.

    Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).

    Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.

    external-advice-audit-against-own-arithmetictransfer-ratios-not-absoluteslatency-dr-covers-measured-pipelinereference-structure-fk-amplitude-divisioncycle-average-tracking-for-gait-quantitiesadvisor-paraphrase-vs-paper

  19. doctrine-19Hardware sessions are scripted experiments, not tuning sessions

    Real-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.

    Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).

    Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.

    risk-ordered-real-deploymentbattery-bracketed-real-abknow-zero-command-behaviorpush-test-chirality-protocolno-field-tuning-protocolreversible-single-variable-field-experimentssim-veto-needs-real-confirmationfirst-real-get-up-violent-stage-one-policystaged-hang-mat-floor-for-get-uppower-cycle-preflighttwo-machine-config-disciplinefall-guard-becomes-a-statehardware-log-is-the-attribution-input

  20. doctrine-20Close questions in writing; restart when the debt is structural

    Audited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.

    Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).

    Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.

    frozen-verdicts-semantic-boundariesstale-verdicts-under-old-stackzero-offset-calibration-shifts-envelopeplant-swap-invariants-vs-shiftsfreeze-lineage-fix-structure-restartminimal-reward-table-with-provenancewrite-hardware-verdicts-back

  21. doctrine-21Name the quantity in the space it lives in

    A goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.

    Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).

    Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.

    joint-space-proxy-for-task-space-quantitycontact-detector-single-signal-lieszero-partial-credit-is-not-an-iteration-problemheading-integral-not-body-rate

  22. doctrine-22Continuation needs a live gradient; a release is chosen by a scan

    Continue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.

    Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).

    Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.

    converged-continuation-is-poisoncheckpoint-choice-is-a-full-gate-scanstance-decided-by-get-up-pathstop-stacking-roll-back-and-auditcurriculum-history-is-part-of-the-productcontinuation-budget-not-from-zeroroot-maturity-vs-product-quality

Experience cards

171 cards matching “curriculum-history-is-part-of-the-product”.

  • Push DR helped one lineage and hurt another at the same dose - robustness budget is conserved and gets borrowed, not createdpush-dr-conditional-budget-conservation
    Replicatedomnidr-tuningdomain-randomizationcurriculumattribution

    Before opening a disturbance-DR rung, measure whether the untrained policy already meets the spec; if training it anyway, expect the benefit to be conditional on the lineage's existing DR load, grade the intensity, and audit retained margins - budget spent elsewhere will be borrowed back.

    Symptom

    The push rung's outcome flipped with the lineage: direct +/-0.6 m/s push failed outright on first attempt (base walking collapsed - kd1.2 scan 0/3 from iter 3300, sim2sim self-falls with pushes OFF - no PASS point existed); staged +/-0.3 then gave the narrow-kd single-working-point lineage real gains (push survival 1/5 -> 4/5) while the SAME dose made the dual-working-point balanced-band lineage WORSE (20-seed survival 18 -> 12/20 plus across-the-board push regression).

    Context

    The four-ladder verdict ("四梯定案", s2e/s2f at both intensities) named the pattern: "push DR 收益条件性" - the benefit is conditional on how much robustness budget the lineage has already spent. The law candidate: "DR 总预算守恒, 平衡带鲁棒性从抗扰余量借" - total DR budget is conserved; a lineage already covering a wide plant band pays for push tolerance out of its disturbance margin. Both S2 ladders therefore closed at the friction rung, with the decisive numerator: untrained push tolerance already met the 4-6 N*s requirement, so the rung was not needed at all ("⑥ push 不训(收益条件性,免训 ±0.6 已达 标)"). The same accounting later justified the C-before-S2 ordering ("push/μ 两轮已实证 DR 预算有限且会被重分配") and trimmed the second S2 pass to three rungs.

    Change

    Push removed from the standing ladder; graded intensity retained as the method IF a lineage ever needs push training; "does the untrained policy already meet the disturbance spec" instituted as the first check before opening any disturbance rung.

    Outcome

    Two rungs (push, ground mu) deleted from the second S2 pass on measured grounds; the ladder's real yield was re-stated honestly as precision, not robustness (speed gate 0 -> 20/20, zero-command drift 0.98 -> 0.06 m, but push 159 -> 125/160).

    Mechanism

    A fixed-capacity policy allocates representation and margin across the training distribution; adding a disturbance axis to a lineage that already spans a wide plant family forces reallocation - the new tolerance is bought with existing margins. Lineages with narrow plant coverage have free budget, so the identical DR dose lands as gain. Benefit is a property of (dose x lineage state), never of the dose alone.

    Applies when

    • proposing push/perturbation training on a hardened lineage
    • the same DR rung helped one lineage and hurt another
    • accounting where a ladder's robustness gains actually came from
    “push DR 收益条件性 —— s2e⑥a (单工作点血统 kd 窄带) ±0.3 得抗推 1/5→4/5; s2f⑥ (双工作点平衡带血统) 同档反而 20-seed 存活 18→12/20 且抗推全面倒退。规律候选: DR 总预算守恒, 平衡带鲁棒性从抗扰余量借。两阶梯均以 ⑤ 摩擦级收官 … 抗推 4~6 N·s 免训已达标。”
    train/OMNI_V0_SPEC.md § 4. ⑥ push 四梯定案 (2026-08-07)
  • A plateau in a training curve was a population mix, not a half-learned skill - 60% standing at 0.372 m and 40% sitting at 0.19 m - and a category at hard zero stayed at zero through 3,000 more iterationszero-partial-credit-is-not-an-iteration-problem
    Mechanism understoodrecoveryattributionmeasurementattributioncurriculum

    Before buying iterations for a plateau, split the metric by category and check whether it is bimodal; a category at hard zero with no partial credit is missing a capability or a reachable state, and more iterations under an unchanged config will only polish the categories that already work.

    Symptom

    After R0.1 the training-side base_height sat near 0.29 m and the curve was still climbing at the iteration cap, which read as "train it longer".

    Context

    Candidate A (user decision) was a child-run from R0.1's last checkpoint with zero config change - the logged env.yaml files differ only in log_dir - for 3,000 more iterations (R0.2). The per-category acceptance split was already available: prone had scored 0/156 with no partial credit.

    Change

    Continue training unchanged, then read the result by category rather than by the pooled curve.

    Outcome

    supine 91.5 -> 96.4%, side 82.2 -> 87.9%, get-up 0.96 -> 0.84 s, pose error 0.91 -> 0.54 - all improvements to categories that already stood. Prone stayed 0/153; mid 55.9 -> 41.2% was within noise (n=34). Height by category was binary - standing groups 0.372/0.373 m, seated groups 0.187/0.194 m, nothing between - so the pooled 0.29 was 0.61 x 0.372 + 0.39 x 0.19 = 0.30 (measured 0.307): six in ten standing, four in ten sitting. The rendered prone episode was still kneel-sitting at t = 8 s.

    Mechanism

    The pooled mean of a binary outcome only moves when the mix moves; PPO kept polishing the subpopulation that already succeeded while the failing one produced no advantage signal to follow.

    Conflicts

    R0.2 recorded the missing capability as "prone lacks rolling over"; R0.3's end-state confusion matrix retracted that - prone had righted its torso in 159/159 episodes and was failing to stand from the W-sit. The lesson that iterations could not fix it holds; the named cause was wrong.

    Applies when

    • a training curve plateaus while acceptance shows one category at zero
    • deciding between "train longer" and "change something"
    • pooled training metrics are read as the typical episode
    “**分类别 h 中位把"平台 = 人口混合"钉死了**:数值是**二值**的 —— 站立组 0.372/0.373,坐姿组 0.187/0.194,**中间没有过渡态**。 … **这也是本仓此后读该指标的通用告诫:全体混合的期望会把 双峰分布平均成一个不存在的中间值,必须分类别看。** … **结论:A 不能过门,原因确定为 prone 缺"翻身"这一技能,不是迭代不够。**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §13 R0.2(recovery_r0_2,child-run 续训):A 走完了 —— 推不动 prone
  • The fallen-state reset was designed, not sampled from SO(3) - fixed category shares with jitter, a low drop that settles physically, equal left/right shares for mirror augmentation, and a numeric check before trainingfallen-pose-reset-distribution
    Observed oncerecoverytraining-runcurriculumdomain-randomizationprocess

    Build a fallen-start distribution from named, physically plausible categories with jitter and a settle phase, keep mirrored categories at equal probability, and check the realized shares and penetration numerically before spending a training run on it.

    Symptom

    A get-up policy can only learn from the fallen states its resets produce; uniformly random orientations produce ground-penetrating and limit-jammed states the robot can never be in.

    Context

    R0 reset_root_fallen: supine 30%, prone 30%, side_l 15%, side_r 15%, mid (random axis 50-125 deg) 10%, +/-15 deg jitter, full yaw, dropped from 0.28-0.40 m and left to settle under physics, joints uniform inside the soft limits with a 5% margin plus small random velocities. Random SO(3) was rejected (the advisor agreed). side_l and side_r must have equal probability because mirror augmentation turns a left fall into a right fall. The advisor had proposed supine and prone only for R0; the spec included side and mid because the feasibility accounts showed physical solutions for all of them, and wrote "narrow back to supine+prone" down as the first fallback. With no display on the training box the reset was checked numerically instead of by eye.

    Change

    Category mix as above; realized shares, settle height and penetration measured over 512 envs before the first run. A fallen-state bank (real falls, settled and stored) was pre-registered for R2.

    Outcome

    Realized shares 29.3/31.6/16.4/17.8% against the config, settle +0.262 m, final penetration 0/512 (a 0.10 m peak at the write instant, ankle links only, pushed out within 80 ms because the 0.28 m drop floor is shorter than a fully extended leg). R0's failure was a reward basin, not a reset artifact. The fallen-state bank stayed unbuilt through V3.1 (checklist item open); R0.3 later re-sliced the prone share into roll_l/roll_r bands, which is what forced the acceptance distribution to be frozen separately.

    Mechanism

    A category-structured, physically settled start distribution keeps training on states the robot can actually occupy, and equal mirrored shares keep mirror augmentation a pure doubling of data rather than a bias.

    Applies when

    • designing reset distributions for get-up, recovery or multi-contact skills
    • mirror/symmetry augmentation is on and the task has chiral start states
    • no viewport is available to inspect resets on the training machine
    “角度 jitter ±15°、yaw 全域、0.28~0.40 m 低空放下由物理沉降,关节软限位内 均匀(留 5% 余量)+ 小随机速度。**不用 random SO(3)**(会采出穿地/极限卡死 等现实不可能状态,顾问同判) … side_l/side_r **概率必须相等**(镜像增强的样本同分布前提)”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §3 R0 任务定义 / §9 核查单
  • Prone get-up sat at 0/159 until two gated hinge terms moved the seated feet - first sideways (561 -> 360 mm), then fore-aft (-168 -> +56 mm) - and success went to 158/159 with nothing else changedprone-dead-end-is-foot-placement
    Mechanism understoodrecoveryreward-shapingreward-shapingattributioncurriculum

    When a stuck state and a successful state differ geometrically, penalize the discriminating quantity with a gated hinge that is exactly zero in the state the policy actually reaches (measure it - not the nominal), then re-probe: flattening one axis can move the discriminant to another.

    Symptom

    Prone falls always righted and then sat with the feet splayed wide or tucked behind the hips, from where the policy never stood (0% for four generations).

    Context

    Four lines of evidence pointed at foot position: the configuration probe (ankles 215 mm apart stood 52.3%, 561 mm apart 0.0%); FK showing the action contract's nominal (a = 0) is itself a 465 mm straddle, so the action_rate and still terms were pulling toward the splits; biomechanics (feet tucked under the body cut peak hip-extension torque 148.8 -> 32.7 N*m, -78%); and HoST's foot-displacement term, which this reward table lacked. The earlier "not a reward hole" reading was corrected to "a gradient hole, not a level hole": at the dead point the heaviest term (upright) was saturated with zero gradient, still paid for not moving, and the one live gradient (base_height) pointed at the thigh-horizontal torque barrier. A prone ROM scan had already ruled out pushing up from prone.

    Change

    R0.4: feet_spread_excess = clamp(ankle distance - 0.215, 0, inf) x upright gate, weight -2.0, plus a height-decay factor added after measuring that the policy's real standing stance was 406 mm, not the 215 mm nominal (the plain version would have taxed every successful stand 0.38/s). R0.5: the same shape on the fore-aft axis, feet_fore_seated = |fore-aft offset - 0.05| x upright gate x height decay, target +50 mm (the measured natural offset of standing postures). One variable per rung.

    Outcome

    R0.4: seated ankle distance 561 -> 360 mm, supine/side exactly unchanged, prone 0 -> 1.9%, mid 45.9 -> 62.2%; a probe then showed the discriminant had moved to the fore-aft axis (standing starts +42 to +51 mm, the prone seat -168 mm). R0.5: supine 99.4, prone 99.4, side 100, mid 100%, re-falls 0%; both geometry terms collapsed to ~0 near iteration 13,100 as base_height rose, and the prone fore-aft offset went -168 -> +56 mm - the term's own target, closing the causal chain. The cost, unmeasured at the time: action jitter rose 33% (sum |da|^2 6.82 -> 9.06).

    Mechanism

    An upright-gated hinge is inert while the robot rolls and exactly zero in the achieved stance, so it adds gradient only inside the stuck basin; a seated robot with its feet behind or outside its COM must make a kinematically unfavourable transition to stand, and moving the feet under the body removes it.

    Applies when

    • a get-up or transition skill fails from one start category only
    • successful and failed episodes differ in a measurable geometric quantity
    • a shaping term might tax the posture successful episodes already use
    “`recovery_r0_5`,唯一变量 = 追加 `feet_fore_seated`(与 R0.4 同形状,只换测量轴)。 … 对照 R0.4 的 prone(3.1%,360 mm,**−168 mm**):前后偏移从 −168 走到 +56, 正是这一项的目标量,**判别量被消掉后成功率随之到顶** —— 因果链完整。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §20 R0.5(前后向脚位置项):成功率门全过 —— prone 0/159 → 158/159
  • Continuing a converged policy on a change that carried no new gradient drifted its transfer from 100/98% to 80/28% over 3,000 iterations while every Isaac gate stayed perfect - scan every checkpoint on the second simulator's friction axisconverged-continuation-is-poison
    Observed oncerecoverytraining-runfork-selectionsim2simcurriculum

    Before continuing a converged policy, check that the change creates a live gradient; if it does not, cap the budget at a few hundred iterations, and in every continuation scan each checkpoint on the second simulator's transfer axis (for example low friction) - trainer-side gates can stay perfect while transfer decays.

    Symptom

    V2.7-A (swap the flat_feet term for a compensated version, continue from v2_6c) finished with the line's best Isaac score (100%) and a MuJoCo transfer collapse: mu 1.0 98 -> 80%, mu 0.4 98 -> 28%; the stance it was meant to widen had not moved.

    Context

    The new term's calibration run showed a near-zero tax from the start: the policy already satisfied it, so the reward landscape offered nothing new. A checkpoint scan on MuJoCo mu {1.0, 0.4} located the damage: +100 iterations 100/98% (better than the baseline), then 86/54, 60/38, 80/28 - monotonic decay with training length, while entropy and action noise rose (7.77 -> 8.18, 0.588 -> 0.612): drift, not sharpening.

    Change

    Rule written in: with no new gradient, a continuation budget is short (at most a few hundred iterations) and the MuJoCo transfer axis enters every checkpoint scan. The next rung (V2.7b, a live stance-width gradient) was budgeted at 1,000 iterations with mu {1.0, 0.4} scans every 100 and a stop-on-signal rule.

    Outcome

    V2.7b kept transfer at the same depth (mu 1.0 98% / mu 0.4 92% at +1,000, where A had already rotted to 86/54) and at +3,000 (100/96%): a live gradient preserved transfer. V2.8 then broke that pattern (mu 0.4 2%): the gradient must also be compatible with the policy's existing form.

    Mechanism

    On a converged reward landscape PPO keeps updating without a signal to follow, and the random walk is pulled toward whatever the training plant rewards idiosyncratically - invisible in the trainer's own gates.

    Conflicts

    The drift mechanism is the spec's reading of one decay series plus one contrasting run; V2.8 is recorded as an exception to "live gradient keeps transfer".

    Applies when

    • fine-tuning a converged policy with a small reward change
    • a continuation run's trainer-side metrics improve while real or cross-sim results worsen
    • choosing which checkpoint of a continuation to ship
    “**checkpoint 扫定死因**(μ1.0/μ0.4):**29500(+100 iter)= 100/98%** (优于基线!)→ 30400 = 86/54 → 31400 = 60/38 → 32398 = 80/28 —— **迁移随续训长度单调衰减**。 … **教训入库:收敛均衡上的长续训是毒药 —— 无新梯度时 续训预算须短(≲数百 iter),且 MuJoCo 迁移轴必须进 checkpoint 扫描。**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §43 结果:V2.7-A 判 FAIL —— 换刀本身无罪,毒在续训预算
  • After installing the measured plant, re-run all generations paired on old and new plant - identity metrics carry verdicts, physics metrics re-baselineplant-swap-invariants-vs-shifts
    Mechanism understoodwalksim2sim-gatesim2simplant-calibrationgate-batteryattribution

    Treat every plant upgrade as an era boundary: re-run the retained policy set paired (same seeds/flags) on both plants, carry forward only verdicts whose metrics proved plant-invariant, re-baseline the rest - and mine the systematic shifts as measurements of the old plant's biases.

    Symptom

    With the plant finally fully measured (weighed masses 9.792 kg, bench-identified armature, in-situ friction), no historical sim number was comparable to new runs - "历史 sim 数字跨纪元不可比" - and it was unknown which historical verdicts still held.

    Context

    The era-2c cross-test ran all seven walk generations on the complete plant under one harness (14/14 survived), then re-ran the same 14 configurations on the OLD plant retrieved from git, same flags and seeds, with a self-check (one historical record reproduced digit-for-digit). The split was clean. Policy-identity metrics moved essentially zero across the plant swap - dominant frequency (v7's period-doubling 1.20 -> 1.20), knee amplitude (9.9 -> 10.0), foot distance (+/-2 mm), saturation (100 -> 100, 0 -> 0) - so all seven cross-generation verdicts (freeze signature, saturation-line closure, knee-collapse location, slip-penalty accounting, N2 lineage, v6's balance, v11's triad) were re-confirmed on the honest plant. Plant-physics metrics shifted systematically with ordering preserved: slip down 10-25% (measured friction makes ground-twisting costlier), landing vertical velocity down 15-50% ("旧 plant 高估落地 凶度"第二次独立证实), landing force mixed (mass up 2.2% vs friction braking the swing - two effects fighting). Exactly ONE behavior-level change: v8's low-speed period-doubling vanished (1.30 -> 2.50, lift normalizing) - confirming it had been machine-dependent bifurcation-edge behavior that armature+friction push off the knife edge, while v7's period-doubling stood untouched: saturation-freeze-driven, a policy property, not a numerical accident.

    Change

    Era re-baselining protocol: after any plant upgrade, one paired same-seed sweep of all retained generations on old and new plant; verdicts keyed to identity metrics carry over, thresholds re-read against the new-plant table, and the differences themselves become plant-physics findings.

    Outcome

    Seven verdicts survived with evidence rather than assumption; two causes of the period-doubling family were separated with plant-side proof; and the sim's landing-violence overestimate was independently confirmed a second time.

    Mechanism

    A policy's structural properties (frequencies, amplitudes, frozen joints) are functions of its weights and survive plant changes; contact-mediated quantities are joint properties of policy and plant and shift when the plant becomes honest. Pairing seeds across plants isolates the plant's contribution exactly, so the sweep both validates history and measures what the old plant had been lying about.

    Applies when

    • installing measured masses/armature/friction into the sim
    • historical thresholds are cited across a plant change
    • a hardware-only behavior might be bifurcation-edge sensitivity
    “策略身份指标逐位不动:主频(v7 1.20→1.20)、膝摆 … 这些是策略属性,plant 换代携带无损,历史定论因此全部成立。… 唯一行为级变化:v8@0.15 的倍周期消失(主频 1.30→2.50…)——印证当时"分岔边缘、机器相关"的判定:armature+摩擦把 v8 推离刀锋;v7 的倍周期纹丝不动(1.20→1.20),它是饱和冻结驱动的深层属性,不是数值巧合。”
    train/README.md § 纪元 2c 全代同机横测 (2026-08-04): 完全体 plant 上历史结论全部存活
  • Symmetrizing the config made the gait MORE asymmetric - the asymmetry lived in the policy weightsasymmetry-in-weights-not-config
    Mechanism understoodwalkattributionattributionreward-shapingcurriculum

    Localize a persistent asymmetry by intervening at the config layer first: if the symptom survives (or worsens), it is in the weights - fix it with symmetry-constrained training, not with trims or offsets.

    Symptom

    walk_v1 on hardware: straight-line command curved 149 deg in 15 s (9.9 deg/s) with 3.06 m lateral runout; turn gain +31% one way vs +129% the other (75% difference); knee asymmetry 4.4 deg in sim, 9.6 deg on the robot.

    Context

    The obvious suspect was the asymmetric default pose in the config. The decisive test: symmetrize standing_pose and run the SAME policy in sim - the asymmetry got LARGER (hip_pitch 6.8 -> 9.3 deg). Root cause therefore not in the config but baked into the policy weights: PPO without a symmetry constraint commonly converges one-sided, because splitting the work 50/50 and loading one side yield the same return, and the gradient falls randomly into one of the equivalent optima.

    Change

    Fix redirected from config trimming to retraining with mirror data augmentation (walk_v2 spec) - a weights-level fix for a weights-level disease.

    Outcome

    With augmentation (and the symmetric-default precondition), stand_v1 reached 0.0 deg asymmetry on all six joint pairs (from 4.4-7.7 deg), height fluctuation 7 mm -> 1 mm, mean |action| down 33%.

    Mechanism

    Reward-equivalent solution families (who carries the load) leave the symmetric solution unpreferred; SGD picks an arbitrary member and entrenches it. Config changes move the coordinate frame around the entrenched asymmetric function - they cannot move the function. The counterfactual test (change config, watch symptom) localizes the layer the disease lives in.

    Applies when

    • a robot veers or loads one side despite a symmetric-looking config
    • deciding between config trims and retraining for an asymmetry
    • mirrored-turn gains differ by tens of percent
    “根因不在配置里:把 standing_pose 对称化后在 sim 里跑同一策略,不对称反而变大(hip_pitch 6.8°→9.3°)—— 说明不对称烙在策略权重里。这是无对称约束的 PPO 的常见收敛结果(左右各担一半与一边多担的回报相同,梯度会随机落进其中一个)。”
    train/RETRAIN_v2.md § 1. 为什么是对称增强(证据)
  • Audit which joints your imitation term constrains - a task that needs deviation is fighting the referenceimitation-term-scope-audit
    Mechanism understoodomnireward-shapingreward-shapingcurriculum

    List which joints your imitation/deviation terms actually constrain and check the new skill's required motion against that list; for balance-coupled joints deliver references as feed-forward residuals, not absolute-position targets - and never assume "reference = 0" is neutral.

    Symptom

    Sidewalk would not learn despite a dedicated tracking reward; meanwhile the gait-shaping imitation term (joint_pos_ref) computed its error norm over ALL 12 joints while its reference covered only the 6 sagittal joints - roll/yaw reference was constantly 0.

    Context

    Two prior generations had shown the forward gait itself was taught by joint_pos_ref, not discovered by PPO (v6 halved the shaping and swing height collapsed 35 mm -> 4 mm). So the reference is load-bearing - but sidewalk requires hip_roll to deviate from nominal, and the all-joints norm punished exactly that deviation: "一边悬赏一边罚过程" (posting a bounty while punishing the process). A follow-up experiment (C4-redo3, free_roll=True releasing the 4 roll joints from the norm) raised the regularization headroom 6x -> 27x yet sidewalk stayed flat and released hip_roll wandered, killing other skills - net negative, withdrawn. A --roll-absolute probe showed the converse failure: pinning roll to a clock-driven absolute trajectory drove tilt 6.9 -> 13.7 deg. Conclusion recorded: absolute-position imitation cannot teach actions that must be superimposed on state feedback.

    Change

    The audit reframed the problem: neither punishing roll deviation nor freeing roll nor absolute roll tracking works; the reference for a balance-coupled joint must be delivered as feed-forward under the policy's residual control (see feedforward-for-phase-locked-skills).

    Outcome

    free_roll rung: joint_pos_ref term rose 0.887 -> 1.104 (release confirmed effective) but vy stayed flat; regularization hypothesis eliminated by experiment.

    Mechanism

    An imitation error norm defines a cage: joints inside it are pulled to the reference in absolute position, so any skill requiring systematic deviation is taxed per step; but joints carrying active balance cannot follow absolute references either, since their correct position depends on state. The scope and the delivery mechanism of the reference are therefore design decisions per joint, not defaults.

    Applies when

    • adding a skill that moves joints your reference sets to zero/nominal
    • an imitation or deviation penalty coexists with a new tracking reward
    • considering releasing joints from a shaping term mid-lineage
    “前进步态也不是 PPO 自己发现的,是 joint_pos_ref 教出来的(v6 砍半塑形 → 抬脚 35 mm 塌到 4 mm…)。而 ref_joint_offset 原本只写 6 个矢状面关节,roll/yaw 参考恒 0 —— 侧走既没被教,roll 一偏离 nominal 反被 joint_pos_ref 扣分。一边悬赏一边罚过程。”
    train/C_LADDER_RUN.md § 3e. 为什么首战 FAIL / 3i. 解锁笼子
  • When training fails repeatedly, inject the target behavior open-loop - stop tuning rewards for an unverified behavioropen-loop-probe-before-reward-tuning
    Mechanism understoodomniattributionattributionprocesscurriculum

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop on the real plant/sim first, and only resume training once you hold a measured, safe, sign-verified target trajectory.

    Symptom

    Three sidewalk training rounds failed; hypotheses multiplied (exploration failure / wrong reference waveform / insufficient authority) with no way to pick between them by running more training.

    Context

    Instead of a fourth reward guess, the team wrote probe_side_ref.py: the candidate reference is injected open-loop on top of a frozen policy's output (bypassing PPO entirely), directly measuring "what happens if the robot literally does this waveform" - separating all three hypotheses in one experiment (5 seeds x 8 s per condition, several waveform families and gains). The probe immediately eliminated the authority hypothesis (full-amplitude execution, 5/5 survival) and localized the problem to the waveform/measurement side. The closing principle was written down after the saga: without a verified target behavior, tuning rewards is "在黑暗里试钥匙" (trying keys in the dark).

    Change

    Standing method: before opening another training rung on a failing skill, build an open-loop (or task-space) generator of the intended behavior, measure whether the physical system can express it and what it looks like - then train toward a verified, quantified target.

    Outcome

    The probe chain produced the verified waveform (reversed-sign triangle, half gain), quantified safe amplitude (tilt 8.2 deg at band top, foot distance clear of the wall), exposed the metric bug when probe and training disagreed, and ultimately supplied the feed-forward that made C4 pass in +100 iters.

    Mechanism

    Training couples exploration, reward design, and feasibility into one opaque outcome; open-loop injection cuts the loop and tests feasibility and waveform alone. A behavior demonstrated open-loop converts the remaining failure into a pure credit-assignment/reward question - and its measured trajectory becomes the reference itself.

    Applies when

    • repeated training failures on one skill with multiple live hypotheses
    • uncertainty whether the platform can physically express the behavior
    • a reference trajectory's shape/sign/amplitude is guessed, not measured
    “三轮 FAIL 之后不再猜,写 train/probe_side_ref.py 把参考开环注入到策略输出之上(绕过 PPO),直接量「照这个波形做会怎样」,一次分开三个假说:甲 探索 / 乙 波形 / 丙 权限。”
    train/C_LADDER_RUN.md § 3f. C4 真因定谳(开环探针) / 3k. 建议的下一步
  • Mirror augmentation over an asymmetric default injects systematic error - symmetrize the default first and verify the transform bit-exactmirror-augmentation-needs-symmetric-default
    Mechanism understoodwalkobservation-designobservation-honestycurriculumprocess

    Before enabling any symmetry augmentation, make every constant inside the observation encoding exactly symmetric, and validate the mirror transform against forward kinematics to machine precision - an unverified augmentation is a new error source, not a regularizer.

    Symptom

    Mirror data augmentation was about to be added while both default poses (standing_pose, walk nominal_pose) were asymmetric - stale hand-tuned compensations from before a ground re-calibration, with hip_yaw differing 2.40 deg between sides and the foot soles actually tilted (pitch 2.88/1.35 deg, roll -2.47/+0.25 deg).

    Context

    The observation encodes joint_pos_rel = q - default. Under mirroring q_l -> -q_r, the relation (q-default)_l -> -(q-default)_r holds only if default_l = -default_r; with an asymmetric default, augmentation produces observation pairs that are NOT mirror images, i.e. "default 不对称时做镜像增强会引入系统性错误,比不做还糟" (worse than not doing it). The fix: adopt model geometric zero as standing default (MuJoCo FK verified: sole pitch/roll exactly 0, asymmetry 0.00 deg) and a symmetric crouch for walk (hip -0.25/knee -0.5/ankle -0.25 satisfying hip - knee + ankle = 0 to keep soles flat). The mirror transform itself was verified bit-exact before use: pseudovector vs polar-vector sign patterns (ang vel [-1,1,-1], gravity [1,-1,1], cmd [1,-1,-1]), joint swap-and-negate; FK check that left-foot pose under q equals the mirror of right-foot pose under mirror(q), measured error 0.00e+00.

    Change

    Defaults symmetrized first (with init heights recomputed by FK), stand policy retrained on the new default so both policies share one default; augmentation enabled only after the FK mirror test passed.

    Outcome

    stand_v1 achieved exact left/right pairing (l_knee -0.1013 / r_knee +0.1013), six-pair asymmetry 0.0 deg, height fluctuation 7 -> 1 mm, 33% less mean |action|.

    Mechanism

    Augmentation asserts an equivariance of the observation encoding; any asymmetric constant inside the encoding (the default) breaks the asserted symmetry, so the augmented data teaches a false invariance. Verifying the transform against FK geometry tests the assertion end to end, independent of the training stack.

    Applies when

    • adding mirror/symmetry augmentation to locomotion training
    • defaults or trims were hand-tuned per side at any point
    • observations are expressed relative to a default pose
    “观测里 joint_pos_rel = q − default。镜像下 q_l → −q_r,要让 (q−default)_l → −(q−default)_r 成立,必须 default_l = −default_r。default 不对称时做镜像增强会引入系统性错误,比不做还糟。… 位置误差与姿态矩阵误差实测均为 0.00e+00。”
    train/RETRAIN_v2.md § 2. 前提:default 姿态必须先对称化(不是可选项) / 3. 镜像变换
  • FK-verify a borrowed reference's structure, then size its amplitude by the reference's job - it pins phase, the policy adds liftreference-structure-fk-amplitude-division
    Mechanism understoodwalkreward-shapingreward-shapingcurriculumplant-calibration

    When borrowing a reference trajectory: verify its structural claim against your own kinematics (an invariant like flat-foot), assign it the phase-pinning job, and size amplitude low enough that the policy contributes the lift - moving toward a proven foreign value in halves, not jumps.

    Symptom

    walk_v4 had big knee swing (40-46 deg) but only 18-24 mm foot lift - amplitude without hip/knee/ankle phase coordination; later, walk_v5's real-robot swing ballooned to 73.6 deg (sim 55.7) with violent footfalls - amplitude over-driven by the reference.

    Context

    Structure first: Humanoid-Gym's 1:2:1 hip:knee:ankle reference was verified on the local model before adoption - the ratio exactly satisfies the locally derived flat-foot constraint hip - knee + ankle = 0, FK-tested at multiple amplitudes with sole pitch 0.00 deg throughout. Amplitude second, and here the first reasoning failed honestly: FK said shorter legs need LARGER reference scale (0.30 for 30 mm lift), and the FK was correct - but the premise was wrong ("FK 没错, 但前提错了"): it assumed foot lift must come from the reference. HighTorque Pi, same scale, uses 0.08 with a 0.02 m foot-height target - proof that lift is added by the policy ON TOP of the reference, whose actual job is pinning the phase relationship. Scale 0.30 made the reference the entire gait: over-constrained and over-driven. The correction went to 0.15, deliberately not Pi's 0.08: "一次只走一半, 留退路" (walk half the distance, keep a retreat).

    Change

    target_joint_pos_scale 0.30 -> 0.15 as one of v6-minimal's three changes, treating both the footfall force and the lateral kicking (yaw momentum scales with leg swing amplitude).

    Outcome

    v6 improved landing force 1.72x -> 1.55x, suspended tilt 45.9 -> 23.0 deg, turn-gain asymmetry 70% -> 19%; the later v6-halved-shaping experiment (35 mm -> 4 mm collapse) confirmed the reference still carries the gait's existence on this machine - the division of labor is real but machine-specific.

    Mechanism

    A joint-space reference plays two separable roles: encoding structure (phase relations that keep the foot flat) and injecting amplitude (energy). Structure transfers across robots and is checkable by FK against an invariant; amplitude is a negotiation with the policy, and over-assigning it to the reference removes the policy's freedom to modulate lift with state.

    Applies when

    • importing a reference gait / imitation target from another codebase
    • reference amplitude reasoning based on leg length alone
    • real swing amplitude far exceeds sim's under a strong reference
    “FK 没错, 但前提错了。我默认抬脚必须由参考轨迹产生。HighTorque Pi 同尺度机器人 … 用 0.08, 而它 target_feet_height = 0.02 m —— 说明抬脚是策略在参考之上加出来的, 参考只负责钉住髋/膝/踝的相位配合。我们取 0.30 等于让参考本身就是整个步态, 过约束 + 过驱动”
    train/WALK_V6_MINIMAL.md § ① target_joint_pos_scale 0.30 → 0.15
  • An outer heading P-loop at deploy cut drift 10x because its output stays inside the trained command band - then training was aligned to itdeploy-heading-loop-and-align-training
    Mechanism understoodwalkreal-deployreal-acceptancecurriculumattribution

    Fix drift-class problems first with an outer loop whose output provably stays inside the trained command band; when adopting it permanently, align the training command generator to the deployment's actual command mixture (feedback-driven AND constant), matching law, gain, and clip exactly.

    Symptom

    Persistent heading drift on straight-line walking (v6 net yaw 60.3 deg over 15 s) that reward-side fixes had only partially tamed.

    Context

    The deploy stack added --heading: an external P loop wz = clip(0.5 * wrap_to_pi(theta0 - theta), +/-0.6), recomputed each frame and fed into the policy's ordinary wz command slot. Measured: net yaw walk_v6 60.3 -> 5.8 deg, walk_v5 17.4 -> 4.2 deg. A run-level audit later corrected the mechanism story: training had heading_command=False since v1 - the policy had NEVER seen heading-error feedback, so the loop works purely because its output lands inside the trained command distribution wz ~ U(+/-0.6): "收益真实,当时的机理解释写错了" (the benefit is real; the mechanism explanation had been wrong). v8 then closed the loop properly: training-side heading command enabled with rel_heading_envs=0.5 - half the envs get heading-error-driven wz, half get explicit constant wz, because deployment feeds wz BOTH ways (straight-line = heading feedback, turning = constant command) and rel=1.0 would have made constant-wz turning out-of-distribution. The law, gain, and clip were aligned item-by-item between trainer and deploy tool.

    Change

    Deploy-side outer loop first (no retrain needed); then v8-D enabled the matching training-side heading command at rel=0.5 with identical gain (0.5) and clip (+/-0.6), contract unchanged (wz slot carries the computed value).

    Outcome

    Drift handled at deploy (5.8 deg) generations before training caught up; the alignment removed the residual train/deploy distribution mismatch, with the accepted cost booked (open-loop straight walking becomes more OOD for heading-envs - irrelevant since acceptance and deployment always run the loop).

    Mechanism

    A learned velocity-tracking policy is a valid inner loop for any outer controller whose commands stay within the trained command distribution - the policy needs no knowledge of the outer objective. Full alignment then requires training on the same mixture of command sources the deployment actually uses, in the observed proportions.

    Applies when

    • heading/position drift on a velocity-tracking policy
    • designing outer loops over learned locomotion controllers
    • training command distribution differs from how deployment feeds commands
    “审计更正(2026-08-02,run 级 env.yaml):训练侧自 v1 复盘起就是 heading_command=False … 策略从未见过航向误差反馈。--heading 是评估/部署侧外加的航向 P 环(wz=clip(0.5·err,±0.6), 落在训练分布 wz~U(±0.6) 内)。实测净偏航 walk_v6 60.3° → 5.8° … 收益真实,当时的机理解释写错了”
    train/WALK_V7_SPEC.md § 0. 本轮之前已经改掉 (航向闭环, 含审计更正)
  • The walk phase machine structurally cannot express flight - rebuild the representation for duty < 0.5, teach flight with a mask tax, never a cliff bountyphase-machine-structural-limits
    Mechanism understoodrunreward-shapingreward-shapingcurriculumgate-battery

    When a new gait changes the contact pattern's structure, audit whether the phase/mask representation can express it and rebuild the representation if not; teach the new contact pattern with graded mask-mismatch pressure and count it in acceptance with artifact-proof definitions (minimum segment length), never with cliff bounties.

    Symptom

    Running requires both feet airborne, but the walk-era phase machine switches legs by the sign of sin(phase) - with duty < 0.5 the two swing windows must OVERLAP during flight, which a sign-switching representation cannot express at all.

    Context

    The run phase machine was re-architected rather than patched: per-leg phases (left = phi, right = phi+0.5 mod 1) with leg_phase < duty defining stance, aligned to the walk sin convention at duty=0.5 so the machines agree where their domains overlap. Flight is taught by the SAME mechanism that once cured foot-dragging, direction reversed: in the two planned flight windows the contact mask is (0,0) and feet_contact_number_duty charges -0.3 per foot still on the ground - a mild ~0.16/step tax, deliberately NOT a cliff: "悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文". The reference shape (half-sine bump over swing progress) is zero at window boundaries by construction, eliminating the clearing-window that the C4 probe measured to cost 13-19% on non-sinusoidal references. The shape self-check ("零代价选项是什么") was run on three behaviors: standing pays ref everywhere (known cmd=0 stepping risk, booked), walking pays only the flight-window tax, proper running collects full marks.

    Change

    New rewards.py run section (leg_phase_duty / stance_mask_duty / ref_run / clearance_run / contact_duty) with the walk versions untouched byte-for-byte; flight acceptance metric defined with a segment-length floor (>=40 ms to count) so numeric contact flicker cannot fake flight.

    Outcome

    Flight became expressible and taught by a calibrated mild pressure; the walk lineage's phase code stayed frozen as its own contract.

    Mechanism

    A phase representation defines which contact patterns exist in the reward's vocabulary; duty cycling below 0.5 introduces states (double-flight) outside a half-period sign convention's language, so no weight tuning can teach them. And rare desirable events taught by cliff-shaped bounties invite hacks (jumping in place); the graded mask tax prices the planned pattern without creating a jackpot.

    Applies when

    • extending a walking stack to running/jumping (duty < 0.5)
    • a desired contact pattern never appears despite reward increases
    • defining flight/contact acceptance metrics
    “duty<0.5 时摆动窗 (1−duty)T > T/2,两腿摆动窗在腾空段重叠 —— sin 符号切腿的机制结构上表达不了"双脚同时在空中"。… feet_contact_number_duty 对"还踩着地"持续 −0.3/脚 —— 与 walk 治拖地同一机制,方向相反。不腾空的税 ~0.16/步 … 梯度温和不构成悬崖(悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文)。”
    train/RUN_V0_SPEC.md § 5. 相位机设计
  • Before resuming a checkpoint, diff the current cfg against what the checkpoint was trained withresume-state-dr-audit
    Replicatedomnitraining-runfork-selectiondomain-randomizationattributionprocess

    "One variable per rung" counts variables against what the checkpoint actually experienced: audit the checkpoint's logged training config and align every unintended difference before resuming.

    Symptom

    Two consecutive rungs (C1 back-mode, C2' forward-turn) failed from the same root with the same full-regression signature despite adding different new modes - so the mode was not the cause.

    Context

    Both runs resumed s1e-500 with the then-current cfg, which carried PD band (0.8,1.2) plus three DR events (base_com, joint_friction, push_robot) accumulated by later lineages. Verified on the training machine from the source of truth (the run's logged params/env.yaml): s1e-500's actual training state was PD +/-10% (0.9,1.1) and all three DR events None. Resuming it under the new cfg meant eating 4 new plant variables plus a new mode at once - the intended "1 variable" was actually 5. A worse variant (c1_redo from s2e_pd-1400) added push +/-0.3 to a root that had never seen it: near-total collapse within +100 iters.

    Change

    C2 aligned the cfg to the checkpoint's training state before resuming (PD back to (0.9,1.1), three DR events off) - making the new mode the only true variable. Permanent rule recorded: compare the checkpoint's training-time DR with the current cfg before any resume.

    Outcome

    C2 trained successfully from the same root that had "failed" twice (wz 20/20 with genuine sign-antisymmetric response by iter 700-800); the A/B falsification ("两个不同模式同签名崩") plus the env.yaml verification closed the attribution.

    Mechanism

    A resumed policy is instantly evaluated (and its value function trained) under whatever plant distribution the cfg specifies; every DR term the checkpoint never adapted to is a distribution shift applied on day one, compounding with the intended change. Single-variable discipline is therefore a property of (cfg diff) x (checkpoint history), not of the cfg diff alone.

    Applies when

    • resuming or forking any checkpoint under an evolved config
    • a resumed run degrades broadly within the first few hundred iterations
    • two different changes from the same root fail with the same signature
    “A/B 定谳:两个不同模式同签名崩 → 病因不是模式,是「从 s1e-500 续训」。… s1e-500 训练态 = kp/kd ±10% (0.9,1.1),base_com / joint_friction / push_robot 全 None;而 cfg 里带着 (0.8,1.2) + … 三个 DR —— 从它续训等于一次吃 4 个新 plant 变量 + 新模式 … 永久教训:续训前必须比对 checkpoint 的训练态 DR 与现行 cfg。单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」。”
    train/C_LADDER_RUN.md § 3b. 这不是重复实验 —— 前两次的病根已定位并修掉
  • After seven patch-generations, freeze the lineage as a regression baseline, fix the structural debts, and retrain from zerofreeze-lineage-fix-structure-restart
    Observed oncewalkprocessprocesscontract-freezecurriculum

    When successive rungs keep trading one symptom for another, ask whether the remaining problems are structural (contracts, latency, sampling, reward-table architecture); if so, freeze the lineage as regression baselines, pay the structural debts, and restart minimal - carrying forward laws and instruments, not weights and weights' patches.

    Symptom

    The v5-v11 walk lineage had accumulated interacting patches (reward terms, gates, clamps, per-joint scales) faster than it converged on the user's goal; v12's spec itself was superseded before training by an external review's verdict that the remaining problems were structural, not parametric.

    Context

    The 2026-08-05 status banner records the pivot: the walk profile was rolled back wholesale to v10b parameters, the v5-v11 lineage frozen "只作回归对照" (kept only as regression baselines), and four structural debts were named as prerequisites for a from-zero straight-walk baseline: the action-latency FIFO (fixed with its own test), the ONNX manifest contract, discrete command sampling, and a minimal reward table. The v12 spec - fully designed, partially implemented - was suspended: "本规格挂起,不再按此开训".

    Change

    Strategy switched from "one more patch generation" to freeze-fix-restart: lineage checkpoints retained as comparison anchors, infrastructure hardened first, then a clean retrain with a minimal reward table (this restart produced the s* generation that later became the real-robot SOTA line).

    Outcome

    A designed-and-ready training generation was deliberately not run - the review's structural findings outranked sunk design cost; the restart line inherited seven generations of laws (calibrations, gate batteries, falsified fixes) without inheriting their entangled reward table.

    Mechanism

    Patch lineages accumulate coupled terms whose interactions eventually cost more to reason about than a restart costs to train; the knowledge worth keeping is the laws and instruments (measured plant values, calibrated gates, falsified directions), not the entangled weights. A restart on hardened structure converts the lineage's lessons into a clean initial design instead of another delta.

    Applies when

    • repeated rungs shuffle symptoms without net progress
    • an external review flags infrastructure/contract debts
    • deciding between another patch generation and a clean retrain
    “同日外部评审定调换路线:冻结 v5~v11 血统(只作回归对照),修结构性问题(latency FIFO 已修 tests/test_action_latency.py、ONNX manifest 契约、离散命令采样、最小奖励表)后从零训直行基线。本规格挂起,不再按此开训。”
    train/WALK_V12_SPEC.md § ⚠️ 状态 (2026-08-05)
  • The first real-robot get-up was "very violent, kicking on the floor, dangerous" - a sim-perfect policy with no reason to be slow, unbounded absolute targets, no domain randomization and a rate limiter that filtered nothing; the task was restated as "safe, slow, transferable"first-real-get-up-violent-stage-one-policy
    Observed oncerecoveryreal-deployreal-acceptanceactuator-modelingattribution

    Do not put a get-up policy on hardware until its action is bounded (hard bound or state-anchored targets), smoothed, randomized and tested at the real pipeline's latency, and say explicitly that the task is "safe, slow and transferable" - a simulation-perfect policy optimizes only "gets up".

    Symptom

    On 2026-08-09 the user ran a V0-lineage recovery policy on the real robot and stopped it: very violent, kicking on the floor, dangerous. The planned next rung (a heavier torque_headroom) was never started.

    Context

    The spec had pre-registered that R0/R1 products stay in simulation and that the real-robot precondition was the R3 smoothing rungs plus a bridge-slew check plus a hanging protocol; the robustness (DR) rungs had not run. In simulation the policy passed 100% with a get-up of about a second. Which ONNX, which gain profile and whether a torque/joint log existed were left "to be recorded later" and never were.

    Change

    The V0 ladder was stopped at its best product (R3.1, sim only) and a re-rooting proposal was put to the user. The spec's four-layer account: style (the reward pays for standing early and nothing pays for slowness - HumanUP's "Stage I" get-up, "fast but unsafe ... infeasible for real-world deployment"); impact (full-range absolute targets with no hard bound, raw |a| up to 4.77, action saturation 100%, a single-step change of 0.306 saturating hip_pitch); transfer (zero DR, friction pinned at 1.0, the learned leg bracing); link (the bridge's RL slew equals vel_limit, 0.2-0.66 rad per step, while the real pipeline has 1-2 steps of time-varying latency and acceptance ran at delay 0).

    Outcome

    The line was re-rooted twice (training-side rate limit, then the beta-anchored action space) and gained a hang protocol before the next real attempt; on 08-11 a beta-anchored policy produced the line's first real get-up.

    Mechanism

    A task reward that pays for standing early selects the fastest feasible get-up; with absolute full-range targets every large target jump is a torque impulse bounded only by the clip; zero DR and braced-leg solutions do not transfer; and a limiter set at the velocity limit does nothing at 50 Hz.

    Conflicts

    The four layers are the spec's reconstruction from simulation probes and the literature; the real run's policy file, gain profile and log were never recorded, so no layer was confirmed against hardware data.

    Applies when

    • a first hardware trial of a high-effort skill is being scheduled
    • sim success is high but the policy saturates actions or torques
    • pre-registered hardware preconditions are not all met
    “用户真机反馈:**非常猛、地上乱踢、危险**,叫停(R3.3 torque_headroom 加档已选型 weight −0.5→−1.5,未启动)。真机细节(哪个 onnx、什么档、有无 τ/q log)**待补记** … 任务从"能起来"变成 **"安全、慢、可迁移"** … **链路层**:桥层 slew RL 档 = vel_limit(10/20/33 rad/s ≈ 每拍 0.2~0.66 rad), 对 recovery 形同虚设;真机 1~2 拍时变延迟,验收默认 delay 0。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 真机叫停与换根判决(2026-08-09)
  • The standing-pose reward had been pulling toward the narrow stance the whole line was fighting - a zero-training kinematic audit of the target vector found it, after first auditing the wrong nominalpose-target-geometric-audit
    Mechanism understoodrecoveryattributionreward-shapingattributionplant-calibration

    Before training on a pose target, audit it with forward kinematics - is it geometrically consistent (feet flat, intended stance, intended width) and is it the frame you think it is (action nominal vs standing default)? A posture term's target may itself be the attractor you are fighting.

    Symptom

    Several rungs aimed at widening the stance failed; the stance stayed narrow as if something kept pulling it back.

    Context

    The audit (MuJoCo forward kinematics, no training): every 5 deg of hip roll widens the stance ~5.5 cm (0.271 m at 5 deg, 0.383 m at 15 deg); at 47 deg of hip yaw a wide stance cannot be flat-footed (residual foot tilt ~0.7 x hip roll), which explained the stalled rungs. The first report also said the stand_pose nominal (hip roll 25, knee 60) has a 63 deg residual foot tilt - but that was the action frame's nominal (the limit-midpoint squat), which stand_pose never used, despite a docstring warning not to mix them. stand_pose's real target was DEFAULT_JOINT_POS: the contract's all-zero pose, legs parallel, ~0.22 m apart.

    Change

    The disease statement was corrected in writing: the narrow stance was not an accidental by-product of proxy traps but the target stand_pose had been actively rewarding. A stored "narrow the stance" knife was marked toxic. V2.8 moved the target to a flat 15-deg stance (sigma 3 -> 1.5, flat_feet margin 5 -> 20 deg).

    Outcome

    V2.8 still failed in-lineage (stance unchanged, feet nearly overlapping, mu 0.4 transfer 2%) - see stance-decided-by-get-up-path - and the from-scratch V3.1 removed both roll joints from stand_pose and put width into a task-space term, which is what finally produced a 0.355 m flat stance. The 63 deg finding was kept as a warning: an action nominal used as a standing target would be a ready-made pit.

    Mechanism

    A posture term with a sharp kernel around the wrong target is an active attractor; every other term fighting it pays twice.

    Applies when

    • a posture keeps returning despite penalties against it
    • a reward uses a default or nominal pose as its target
    • the contract has more than one "nominal" (action frame vs standing pose)
    “上文"stand_pose 的 nominal (hip25/knee60) 残倾 63°"**审计错了对象**:那是 **动作参考系 nominal**(限位中点蹲),stand_pose 从未指向它(函数 docstring 原文即警告"两者别混",还是混了 —— 记档)。 … **修正后的病根陈述:窄站距不是代理陷阱的意外副产物,而是 stand_pose 一直在主动奖励的目标本身**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 勘误(实现时抓到):审计混了两个 nominal —— 真病根比误诊的更直白
  • The real robot's right-leg kicking was over-trained-delay times loop gain - irreducible pipeline latency is plant, model it fully from day onepipeline-latency-is-plant-not-dr
    Mechanism understoodomniactuator-modelingactuator-modelingreal-acceptanceattributionsim2sim

    Measure the end-to-end action pipeline delay and build it into the nominal plant and every acceptance gate from day one; treat power/scale deratings that "fix" oscillation as gain-reduction crutches flagging an unmodeled delay, and expect higher-feedback-gain policies to be MORE delay-fragile.

    Symptom

    On hardware, s1c/s1d at action scale 1.0 always kicked wildly with the right leg (s1c only ran as SOTA at power 0.8; s1d only at 0.7) - while sim showed nothing under default evaluation.

    Context

    Sim reproduced the incident item by item once the real pipeline delay was injected: s1d@1.0 with --delay 1 fell at 10.2 s, --delay 2 at 5.2 s; s1c@1.0 stressed (r_hip_roll saturation 5 -> 16%; "右脚" = the policy's chirality makes the right leg its high-gain leg); and the combos that worked on hardware (s1c@0.8+delay2, s1d@0.7+delay2) all survived in sim. Mechanism: the real pipeline is ~1-2 ticks (BusWorker next-cycle pickup + CAN round trip) but S1.1 trained only to 1 tick - "超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定" (delay beyond training x full loop gain = oscillation; the power derating had been buying stability by compressing loop gain). s1d was MORE fragile than s1c because its yaw 3-layer stack had learned higher feedback gain - higher gain, lower delay tolerance. Three changes: latency DR widened to cover reality; acceptance gates and smoke runs moved permanently to --delay 2 ("门必须在真机条件下预测 真机"); and the doctrine written twice-paid: "不可约的管线属性(延迟、 限速)不是'随机化选项',是 plant 本体,第一天就该全额建模" - S1's nominal-then-robust staging falsified by hardware for the second time. The later s1e hardware run at power 1.0 (no kicking, normal force) closed the loop: "0.8 = 旧代拐杖" - the derating had been a crutch for the under-modeled delay, not a real requirement.

    Change

    Latency modeled as plant from day one of any lineage (measured 1-2 ticks covered, bridge-layer rate limits likewise modeled by default); every gate and smoke evaluation issued under --delay 2.

    Outcome

    Kicking reproduced, explained, and eliminated in the s1e generation at full scale and full power; the deploy-side crutches (0.7/0.8) retired for the new lineage.

    Mechanism

    Feedback oscillation onset is a product of loop gain and phase lag; a policy trained below the real delay learns gains that sit past the real stability margin, and any output derating masks it by scaling gain down. Since pipeline delay is deterministic hardware property - not an uncertainty - it belongs in the nominal plant, and every evaluation must include it or the gate predicts a robot that does not exist.

    Applies when

    • hardware oscillation/kicking that sim only reproduces with added delay
    • a policy only runs on hardware at reduced power/scale
    • defining what belongs in the nominal plant vs the DR list
    “真实链路延迟 ~1~2 拍 … S1.1 只训到 1 拍——超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定。s1d 比 s1c 更脆 = yaw 三层栈学出更高反馈增益,增益越高延迟容忍越低。… 教训入账:S1「先标称后鲁棒」第二次被真机证伪——不可约的管线属性(延迟、限速)不是"随机化选项",是 plant 本体,第一天就该全额建模。”
    train/OMNI_V0_SPEC.md § 3. S1.4(真机右脚乱踢事故强制)
  • Real-robot trials of a new skill were staged by risk - a hanging dry run with the robot posed by hand, then one short try per category on a mat with the hardest last, then the composed behaviour (switch + walking) last - with the user present and a log every timestaged-hang-mat-floor-for-get-up
    Replicatedrecoveryreal-deployreal-acceptanceprocess

    Stage a new skill's hardware trials by risk - hanging dry run posed by hand, one short try on a mat per start category with the hardest last, the composed behaviour last - with an operator ready to cut enable and a log for every try; relax a safety ban only for short, attended runs and say so in writing.

    Symptom

    A get-up policy acts violently near the ground by design, and the first unstaged real run of the line was stopped as dangerous.

    Context

    The hanging checklist written with the first stamped recovery product (v2_5, 2026-08-11), to be ticked item by item with the user present: both machines on the same commit and firmware torque limits checked; the robot hung from a single point about 0.1 m off the ground; a dry run with the robot posed by hand into supine and prone to watch that the target stream is gentle (the beta contract keeps targets within +/-0.25 rad of the measured pose, so enabling causes no homing fling); the first floor try is supine only, on a mat, once, with torque and joint logs; categories are added one at a time, prone last; any kicking or oscillation cuts enable immediately. For the switch (08-14) the runbook orders: hang with the standing policy as the locomotion side, then on a mat push the robot over and let it recover, and only last swap in the walking policy. An exemption was also written: edge-standing policies stay banned from long or unattended runs, but a short single A/B with the user present, hung or on a mat, is allowed. The one-leg line reused the same order (hang, then floor with a spotter, 60 s segments with a temperature check).

    Change

    Real trials as a checklist of stages, each gated on the previous one, with the composed behaviour last.

    Outcome

    The line's first real get-up (v2_6, 08-11) came through this protocol and was reported "fairly stable"; no further hardware outcomes of the switch are recorded in the spec.

    Mechanism

    Each stage exposes one new risk (commanded targets without contact, a single category with contact, harder categories, then the interaction of two policies), so a failure is attributable and cheap.

    Applies when

    • first hardware trial of a recovery, jumping or other high-impact skill
    • switching between two policies on hardware for the first time
    • a policy with a known posture defect needs a comparison run
    “吊挂空跑: 手动摆到 supine/prone 姿态, 看目标流是否温和 (β 帽 7.5 N·m, 目标永远贴着当前 q ±0.25 rad —— 使能瞬间无归位甩动, 这是 β 契约附带保证) … 落地首试: supine 一类, 垫子, 单次; τ/q --log 全程记录 … 逐类别扩展 (prone 最后), 每类先单次”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §40 吊挂执行单(真机首试;需用户在场,逐项打勾)
  • A quadratic clearance reward stalled for 3500 iters near target - switch to an indicator on accumulated heightindicator-reward-avoids-gradient-decay
    Mechanism understoodwalkreward-shapingreward-shaping

    When a shaped term plateaus near its target, check the gradient profile: replace vanishing-gradient forms with threshold/indicator forms for the final approach, and prefer delta-accumulation over absolute positions to immunize against frame offsets.

    Symptom

    The quadratic-error clearance term froze at -0.002 from iteration 2000 to 5500 - thousands of iterations with no progress on foot lift.

    Context

    Diagnosis: a quadratic penalty's gradient vanishes as the error approaches target, so exactly where the last millimeters must be earned the incentive fades to nothing. The replacement (Humanoid-Gym form): accumulate the swing-phase height climb per foot, reward a BINARY indicator |accumulated - target| < 0.01 masked to the planned swing window, reset on contact, weight +1.6 as a positive reward. Two properties: the indicator's incentive is constant until the threshold is crossed (no decay zone), and accumulating height DELTAS makes any constant sole-frame offset cancel automatically - which structurally sidesteps the earlier 0.0585 m zero-point bug ("顺带绕开我先前那个'忘了减 0.0585 导致惩罚恒为 0'的坑").

    Change

    Clearance reformulated from quadratic penalty on instantaneous height to indicator on per-swing accumulated climb (target 0.03 m by leg-length scaling, weight +1.6).

    Outcome

    Part of the v5 package under which lift finally moved (v5 29 mm, v6 34 mm vs the stalled 18-24 mm era); the offset-cancellation property removed one whole bug class from the term.

    Mechanism

    Policy-gradient learning follows the reward's local slope; quadratic shaping concentrates slope far from target and starves it near target, so convergence stalls precisely at the finish line. An indicator pays a constant bounty until the goal is met; formulating on deltas rather than absolutes removes sensitivity to reference- frame constants.

    Applies when

    • a reward term's value freezes short of target for thousands of iters
    • designing clearance/height/precision terms
    • reward code depends on absolute link positions
    “现行二次型在接近 target 时梯度趋零 —— 这正是 clearance 从 iter 2000 到 5500 卡在 −0.002 不动的原因。… 二值指示在跨过阈值前梯度恒定,没有衰减区 … 累积 delta 让 SOLE_OFFSET 自动抵消”
    train/WALK_V5_SPEC.md § 3. clearance 改峰值型(去掉二次型的梯度衰减)
  • The recovery line's real-robot verdicts live in three places that disagree - the spec's "fairly stable, first real stand-up" for v2_6, an undated runbook note that only v3_1p1c works, and a first run whose details were never recordedwrite-hardware-verdicts-back
    Observed oncerecoveryreal-deployreal-acceptanceprocessattribution

    A hardware verdict is a dated entry in the authoritative ledger - policy file and stamp, gain profile, floor, battery, tries, log file, what was seen - written back the same day; a note in a command file is a pointer, not a verdict, and a newer verdict that contradicts an older one must say so.

    Symptom

    Asked "which recovery policy works on the real robot", the sources give different answers, and none of them carries the conditions of the test.

    Context

    08-09: the first real run was stopped as violent and dangerous; which ONNX, which gain profile and whether a log existed were marked "to be recorded" and never were. 08-11: v2_6 was "fairly stable", the line's first real get-up, with splits after standing; v2_5b's result and both CSVs were "to be reported". 08-14: v3_1p1c was stamped and pushed, with "the real first test still needs the user present"; the spec records no hardware result for it. The operator runbook (undated) puts above the v2_6 and v2_5b floor commands the note that none of the recovery policies below work, only recovery_v3_1p1c - a verdict never written back into the spec, with no date, floor, battery, number of tries or log attached.

    Change

    None recorded in the sources; this card records the gap.

    Outcome

    The line's authoritative record ends with v3_1p1c as the product awaiting its first real test, while the operator's note implies it is the only one that works and that v2_6 (recorded as a success) does not.

    Mechanism

    Verdicts given at the robot travel by word of mouth and command-file comments; without a record carrying the conditions, a later reader cannot tell a changed verdict from a changed floor, battery or stack.

    Conflicts

    §43 (2026-08-11) records v2_6 as the first successful real get-up ("fairly stable"); the undated runbook says every recovery policy except recovery_v3_1p1c does not work; §49 (2026-08-14) says v3_1p1c's first real test was still pending. The runbook's claim has no date and was never written back to the spec, so it cannot be ordered against §43.

    Applies when

    • choosing which policy to deploy from an operator's notes
    • a hardware session ends without a written result
    • two documents disagree about what worked on the robot
    “下面的recovery都不行 只有recovery_v3_1p1c.onnx”
    RL系统/FOLLOW THIS copy 2.md § #### Recovery Policy (operator runbook, undated)
  • Four consecutive fixes were each continued from the previous fix's degraded state until the user stopped the ladder - "change parameters, don't stack errors" - rolled back to the last good checkpoint and audited the target geometry firststop-stacking-roll-back-and-audit
    Observed oncerecoveryprocessprocessfork-selectionattribution

    When successive rungs each start from the previous rung's output and the target symptom does not move, stop, roll back to the last good checkpoint and re-derive the next change from an audit; keep the measurements, discard the stacked remedies.

    Symptom

    After the real-robot splits, the V2.7 ladder tried to widen the stance: a term swap (A), a new stance knife (b), more iterations (甲), a doubled weight (乙). Stance barely moved while hip yaw ratcheted 46.7 -> 49.9 -> 52.5 deg toward its 60 deg limit.

    Context

    Each rung started from the previous rung's output. The user ruled on 2026-08-11 that things had gone wrong from V2.7-A: go back to v2_6 and rethink which parameters to change instead of stacking errors.

    Change

    乙 was killed at start and not counted; the product baseline rolled back to v2_6c model_29399; every measurement and law learned on the ladder was kept ("the data is real; what stacked was the treatment"). Before any new training, a zero-training kinematic audit of the stance targets was run.

    Outcome

    The audit found the stand_pose target itself rewarding the narrow stance (pose-target-geometric-audit) and showed geometrically why the yawed stance could not be widened with flat feet - so 乙 was proven unnecessary without running it. The next in-lineage attempts still failed, which is what established that the stance is set by the get-up path.

    Mechanism

    A rung continued from a degraded state inherits its compensations, so each new fix answers the previous fix's side effects; the yaw ratchet was the visible trace of that stacking.

    Applies when

    • three or more corrective rungs in a row without progress on the target metric
    • a side-effect metric ratchets in one direction across rungs
    • a new rung is being planned from the latest (not the best) checkpoint
    “用户裁:"从 V2.7-A 开始就出问题了,应该回到 2.6 再思考如何改变参数而不是 错误叠加。"认账:A 的补丁 → b 的新刀 → 甲的加时 → 乙的加权,每级都从上级 的**退化状态**续(yaw 46.7→52.5° 的棘轮就是叠加痕迹)。 … 数据是真的,叠加的是处置。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 方法论裁定(2026-08-11,用户):V2.7 全阶梯叫停,回滚 v2_6
  • The action-delay was implemented as lerp - beyond one step it extrapolated BACKWARD, so a whole lineage trained on a fictitious actuatorlatency-lerp-reverse-extrapolation
    Mechanism understoodinfraactuator-modelingactuator-modelingplant-calibrationsim2sim

    Unit-test plant-model code (delays, filters, randomizers) against hand-computed truth across its FULL configured range, not just the nominal case - a delay must be a queue, and any interpolation used outside [0,1] is a silent plant corruption that training will faithfully absorb.

    Symptom

    Every walk/stand model up to v11 had been trained on a silently wrong plant: the action-latency implementation lerp(cur, prev, lag) is only an interpolation for lag <= 1 - at lag 3 it computes 3*prev - 2*cur, a REVERSE extrapolation. With the configured (0, 0.06) s at 50 Hz (lag in [0,3]), about 2/3 of environments were adapting to actuator dynamics that do not exist.

    Context

    Listed as evidence item #1 for the full restart: "全部旧模型训在错误 plant 上" and the head suspect for the real robot's wild kicking. The fix replaced it with a true FIFO delay line (commit 7f13793) plus its own regression test (tests/test_action_latency.py) - but every exported ONNX predated the fix, which is part of why the lineage was frozen rather than patched.

    Change

    Delay implementation rewritten as an honest FIFO with unit tests; the restart baseline trained on the corrected plant from day one.

    Outcome

    A generation-scale training investment was revealed to have a corrupt plant underneath; the class of bug (plausible-looking math that silently changes meaning outside its valid range) got a permanent test.

    Mechanism

    lerp(a, b, w) leaves the segment for w > 1; used as a delay it fabricates high-gain inverted dynamics precisely in the largest-delay draws, so the policy learns compensation for an actuator that cannot exist - and DR then trains robustness to the artifact rather than to reality. No training metric can catch this: the sim is self-consistent, just wrong.

    Applies when

    • implementing or auditing action delay / filtering in a trainer
    • a lineage behaves as if compensating dynamics nobody modeled
    • deciding whether old checkpoints are salvageable after a plant bug
    “动作延迟旧实现 lerp(cur, prev, lag) 在 lag>1 时是反向外插(w=3 → 3·prev−2·cur),配置 [0,0.06]s@50Hz 即 lag∈[0,3],约 2/3 env 在适应不存在的执行器动态。7f13793 已换真 FIFO,但所有 ONNX 均训于修复之前 —— 真机"乱踢"的头号嫌疑。”
    train/OMNI_V0_SPEC.md § 0. 为什么从零 (1)
  • Brief the operator on the lineage's measured zero-command and untrained-axis behavior before handing over the joystickknow-zero-command-behavior
    Replicatedomnireal-deployreal-acceptanceprocess

    Before any teleop/demo, measure and write down the policy's zero-command behavior and per-axis competence, label untrained axes explicitly as not-bugs, and set the floor/procedure to accommodate the known drift.

    Symptom

    A teleop session was about to start on a policy that does not stand still at zero command and has never been trained on lateral commands - behaviors an unbriefed operator would report as bugs or emergencies.

    Context

    Three measured facts were written into the teleop instructions ("都有 实测依据, 不是猜"): (1) A/D (lateral) keys will get essentially no response - probe-measured sidewalk tracking ~3%, an untrained axis: "这正是 C4 要解决的事, 不是 bug"; (2) no keypress = cmd 0, and this lineage does not stand still at zero command - a three-generation lineage property: paces in place, drifts right ~5 cm/s, net rotation -30 deg/20 s; sim survival is 20/20 (it will not fall) but it walks away slowly, so leave floor margin especially on the right; (3) S (backward) WILL respond - probe-measured 20/20 survival, 67% tracking untrained, which is also why this root was chosen for the C ladder. Plus a keybinding dry-run while suspended before touching down.

    Change

    Operator briefing became part of the deployment artifact: expected response per key, expected idle behavior with magnitudes and directions, and the distinction between untrained (expected, not a bug) and abnormal.

    Outcome

    The session proceeded with correct interpretations available in advance; the known zero-command wander was handled by floor margin and start-with-command procedure rather than misdiagnosed on the spot.

    Mechanism

    A learned policy's off-nominal behaviors (idle drift, untrained axes) are lineage properties, stable and measurable in sim beforehand; operator surprise converts known properties into false incident reports and unsafe reactions. A briefing transfers the measured behavior model to the person holding the controller.

    Applies when

    • handing a learned policy to an operator or demo audience
    • the policy idles in a non-stationary way at zero command
    • some command axes are untrained in the current lineage
    “A/D 基本不会有反应 —— s1e 从未训过非零 vy, 选根探针实测侧走跟踪率 ~3% … 这正是 C4 要解决的事, 不是 bug。… 不按键 = cmd 0, 而 s1e 在零指令下不站定 —— 血统属性, 三代实录: 原地踏步 + 右漂 ~5 cm/s + 净旋 −30°/20s。”
    train/REAL_RUN_S2.md § 附: WSAD 遥控 上机前必须知道的三条
  • A training-side rate limit anchored on the last commanded target is an integrator inside the balance loop - two unrelated lineages converged to the same 34-43% re-fall rate, a soft penalty could not fix it, and the bandwidth arithmetic said safety and standing could not coexistslew-anchor-is-an-integrator
    Mechanism understoodrecoverytraining-runactuator-modelingcontract-freezeattribution

    If you constrain actions in training, anchor the constraint on the measured state, not on the previous command - a limiter with memory adds lag inside the balance loop; and when two different lineages converge to the same failure rate, treat the cause as structural and stop adding soft penalties.

    Symptom

    With the rate limit moved into training (V1), policies either could not stand up or stood up and kept falling again: the stand oscillated, fell and climbed back, 34-43% of the time.

    Context

    V1.0-B (user decision: explore new postures from scratch, hard constraint in training): target <- prev + clip(target - prev, +/-S*dt) at the TIGHT rates, anchored on the last issued target like the bridge. Four runs: v1_0 from scratch 0.8% - every category righted under the limit, then knelt (the limit also damped the exploration that had escaped the seated basin in V0); v1_0c continued from R3.1 99.8% get-up but 36-43% re-fall and knee jitter 0.604 rad/s; v1_0p with a pull-assist curriculum (56 N -> 0, fully withdrawn) 39.1% unassisted, re-fall 34-42%; v1_0w with a windup-gap penalty 25.4%, re-fall 18-43%. Removing the limiter from v1_0p gave 0.0%: the policy had co-adapted with it.

    Change

    The rate-limit route was declared dead after four runs and the line was re-rooted on a beta-anchored action space (V2, user approval required).

    Outcome

    V2.0's first acceptance at full authority already showed re-falls of 0-1% (V1: 34-43%), the structural bet paying off before any tuning.

    Mechanism

    Anchored on the previous command, a saturated policy becomes a rate controller - one more integrator in the loop - and active balance through that lag oscillates, while a kneeling sit needs no active control and is stable. The arithmetic: kp 30 needs 0.4 rad of error for 12 N*m; at 4 rad/s that takes 0.1 s, half the pendulum time constant sqrt(0.38/9.8) ~ 0.2 s; keeping standing bandwidth needs S of at least ~8 rad/s, within 20% of the 10 rad/s velocity limit - no bound at all. Anchoring on the measured angle (q + beta*a) makes the full kp*beta authority available in one step, with no memory, and caps the impact at the same time.

    Applies when

    • adding rate limits, slew limits or target filters to a policy's action path
    • a policy stands but oscillates and re-falls after a constraint was added
    • different lineages or curricula land on the same failure signature
    “v1_0p(拉力课程):会站(prone 79.9%),再摔 34~42% —— **两条完全不同 血统、不同学习路径,收敛到同一失败率**。 … 动作饱和时它退化为 速率控制 = 环内多一个积分器;主动站姿平衡穿过该滞后必振荡(v1_0 的跪坐不需 主动控制,所以稳)。对照:**HoST 的 β 锚在当前实测 q,无记忆、无积分器**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §31 结构病定案:slew 的目标锚 = 控制环里的积分器
  • Record where every failed episode ends - an end-state confusion matrix showed all failures finishing seated and overturned a "cannot roll over" diagnosis that per-category success rates hide by constructionend-state-confusion-matrix
    Mechanism understoodrecoverysim-evalmeasurementgate-batteryattribution

    For any multi-category acceptance, report where each failed episode ends, not only which category it started in; it costs a few lines and no extra simulation, and it separates "cannot reach the goal" from "reaches the wrong basin".

    Symptom

    Prone scored 0% for three generations; the working diagnosis was "prone lacks the roll-over skill", and R0.3 spent a run adding prone-to-side roll-arc start states. It bought nothing: the 45-deg roll band itself only moved from 24.2% to 26.6% after 3,000 iterations.

    Context

    Acceptance reported success per starting category. A final-state table (lying prone / on the side / supine / seated / standing for every failed episode) was added to accept_recovery.py at R0.3.

    Change

    The confusion matrix became a permanent part of the acceptance output, and the prone diagnosis was rewritten from it.

    Outcome

    The prone, side and supine columns were all zero - every failure ended seated - and prone had righted its torso in 159/159 episodes (tilt under 30 deg in 100%). The missing ability was standing up from one specific seated configuration, not rolling over, which redirected the next rungs to foot placement and to a configuration probe.

    Mechanism

    Per-category success rates collapse "reached the wrong basin" and "never reached anything" into the same zero; the end state separates them.

    Applies when

    • a category sits at 0% and the diagnosis rests on its label
    • recovery, manipulation or navigation tasks with distinct terminal states
    • an intervention aimed at the presumed cause shows no effect
    “**① 末态混淆矩阵 —— 固化(已在 `accept_recovery.py`)。** 它给出的 "趴/侧躺/仰躺三列全 0、所有失败都终于坐姿"是本线最改变决策的一个事实, 而**逐类成功率按构造看不见它**。 … 成本十来行、零额外仿真。 … **② prone 病因更正(旧诊断作废)。** 旧:"缺翻身"。新:**prone 159/159 全部 把躯干翻正**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §15 R0.3 判决 + 三件事的判断
  • A nonzero response with the same sign for + and - commands is bias, not abilitysame-sign-response-is-yaw-bias
    Replicatedomnisim-evalmeasurementgate-batteryattribution

    Before crediting any directional skill, test both command signs: response must flip sign with the command; a same-signed pair is a bias to subtract, not an ability to report.

    Symptom

    Root-selection probe showed nonzero wz "tracking percentages" on turn commands, tempting the read that candidates could partially turn.

    Context

    During C-ladder root selection, s1e-500's measured yaw rate was +0.084 rad/s for cmd +0.3 and +0.093 rad/s for cmd -0.3 - same sign both ways. The same check on the C2 baseline gave wz+0.20 -> -0.13 and wz-0.20 -> +0.12 (again same sign), while the alternative root s2e_pd-1400 gave +0.16 / -0.16 - opposite signs, i.e. a genuine 16% command response.

    Change

    Reading corrected and written into the execution sheet: percentages on directional commands are meaningless unless the +cmd and -cmd responses have opposite signs; all three candidates were re-classified as "cannot turn, cannot sidewalk - C2/C3/C4 learn from zero". Acceptance criteria thereafter required "tracking >=50% AND left/right opposite-signed".

    Outcome

    Prevented crediting turn/sidewalk ability that did not exist; the antisymmetry clause became a standing part of every turn and sidewalk PASS condition (C2, C4, C4-redo levels all carry "且左右反号").

    Mechanism

    A constant yaw (or lateral) bias projects onto any command's sign convention and shows up as fake fractional tracking; only sign-antisymmetry under command reversal distinguishes a feedback response to the command from an open-loop offset.

    Applies when

    • evaluating turn/sidewalk/any signed-command tracking percentages
    • a candidate shows partial tracking on an axis it was never trained on
    • writing PASS criteria for a new directional skill
    “C2/C3 那些非零的 wz 百分比不是转向能力 —— 转向+ 与 转向− 的实测同号(s1e:cmd +0.3 → +0.084,cmd −0.3 → +0.093 rad/s),那是恒定偏航偏置。… 三个候选都不会转、都不会侧走。”
    train/C_LADDER_RUN.md § 0. 读数纠正(重要,别引错)
  • Guessed joint friction was 2.5x low and damping 5x high - measure, then DR around nominalfriction-measured-not-guessed
    Mechanism understoodwalkplant-calibrationplant-calibrationdomain-randomization

    Measure frictionloss and damping separately (they need different rigs), put the measured value at DR center, and express DR as an additive band around that nominal - a DR range around a guessed value can exclude the real robot entirely.

    Symptom

    Old MJCF friction values were invented, not measured; when finally measured, every guessed value was wrong by a large factor in some direction.

    Context

    Joint friction split into Coulomb (frictionloss, tau_c) and viscous (damping b). Measured with the robot hung from a crane (吊机测) while armature was measured no-load; the two measurements are deliberately separated. Old MJCF: frictionloss 0.05, damping 0.1, DR joint_friction range [0, 0.1] "凭空拍的" (made up out of thin air).

    Change

    Replace guessed values with measured ones - frictionloss: RS06 0.15 / RS02 0.12 / RS00 0.13 N*m (old 0.05, i.e. 2.5x too low); damping: 0.02 N*m*s/rad on all three motor types (old 0.1, i.e. 5x too high). DR reshaped from an absolute made-up range [0, 0.1] to an additive band around measured nominal: joint_friction_add [-0.05, +0.10].

    Outcome

    "摩擦定稿(与 armature 一起, plant 参数第一次全部来自实测)" - friction frozen as part of the first fully-measured plant; DR now brackets a measured truth instead of spanning an invented interval.

    Mechanism

    Coulomb friction and viscous damping have opposite behavioral signatures (constant-torque threshold vs velocity-proportional drag); guessing both wrong in opposite directions gives a plant that is simultaneously too easy to start moving and too hard to move fast. DR centered on a wrong nominal makes the policy robust to a family of plants that does not contain the real one.

    Applies when

    • plant friction/damping values have no measurement provenance
    • DR ranges are absolute intervals rather than bands around a nominal
    • policy is over- or under-damped on hardware relative to sim
    “测关节摩擦。吊机测,而armature应该空机测试。… frictionloss τ_c (N·m) │ 0.15 │ 0.12 │ 0.13 │ 0.05(低 2.5×) … damping b (N·m·s/rad) │ 0.02 │ 0.02 │ 0.02 │ 0.1(高 5×) … DR │ joint_friction_add: [−0.05, +0.10] 叠标称 │ 旧 [0, 0.1] 凭空拍的”
    Experience.md § 摩擦定稿表 (lines 12-25)
  • Choose the fork root by which candidate's shortfalls are recoverable, not by headline scorefork-root-recoverable-shortfall
    Mechanism understoodomnifork-selectionfork-selectionprocess

    When picking a checkpoint to fork from, rank candidates by whether their weaknesses are trainable-back, not by current headline metrics; prefer the candidate whose deficits the upcoming training directly pays for.

    Symptom

    Multiple candidate checkpoints for the omni-command ladder root, each best at something different: fric-3000 had the best tracking precision (vx 88-91%) and hardened plant robustness; s1e-500 had lower precision (vx 82%) but was the only candidate that could still walk backward.

    Context

    Root selection ran as a data probe, not a preference vote: 6 candidates x 8 out-of-distribution omni commands x 20 seeds = 960 cells (probe_omni_0808.json). s1e-500 @pw1.0 survived 20/20 in all eight conditions including backward at 67% tracking; the deep-trained fric lineage scored backward 0-3/20 despite better forward precision.

    Change

    Decision criterion made explicit: list what each candidate exclusively wins at, then ask which of those wins the loser could train back. fric-3000's exclusive wins (precision, plant robustness) are both retrainable - precision is directly optimized by the reward, plant hardening is a planned later pass. s1e-500's exclusive wins (backward plasticity 20/20 vs 2/20, disturbance margin 159/160 vs 125/160, push chirality symmetry 40/40 vs 17/40) had all been shown unrecoverable - push-level rungs failed twice, chirality never recovered even with mirror augmentation on. Root = s1e-500.

    Outcome

    s1e-500 carried the whole C ladder; its backward skill was preserved through C2/C4 gates (regress budget <=2/20 enforced), and the final C4 product passed a 260-cell battery at 20/20 everywhere.

    Mechanism

    Training can re-earn anything the objective directly pays for, but capabilities that earlier training destroyed and never restored (plasticity, symmetry, robustness margins) are empirically one-way doors. The information-bearing comparison is therefore recoverability of each candidate's deficit, which the team stated as "独占项的可恢复性正好相反 —— 这就是判据" (the exclusive items' recoverability is exactly opposite - that is the criterion).

    Applies when

    • selecting a resume/fork root among several checkpoints
    • one candidate is more precise but another retains a skill the rest lost
    • planning a task-extension ladder from an existing lineage
    “fric-3000 赢在精度(vx 88~91%…)与 plant 鲁棒性 → 两样都训得回来…;s1e-500 赢在可塑性(C1 20/20 vs 2/20)、抗扰余量(159/160 vs 125/160)、手性对称(推 ±6 N·s 40/40 vs 17/40)→ 三样都训不回来 … 独占项的可恢复性正好相反 —— 这就是判据。”
    train/C_LADDER_RUN.md § 0. 为什么根是 s1e-500(数据,不是偏好)
  • The advisor's "runtime five-stage state machine + per-stage reference poses + RL residual" appeared in none of the three papers it cited - reading the originals changed the plan and downgraded two widely repeated industry claimsadvisor-paraphrase-vs-paper
    Replicatedrecoveryprocessprocessattribution

    Read the primary source behind any piece of advice before adopting its architecture; record where the paraphrase and the original differ, downgrade claims the originals do not support to speculation, and adopt what the verified sources actually share.

    Symptom

    After the violent first real run, an advisor proposed re-architecting recovery as a runtime staged state machine with reference poses and an RL residual, citing HoST, HumanUP and StableMimic.

    Context

    The three papers were read in full on 2026-08-09 and tabulated (deployment form, what the "stages" really are, hard constraints, references). HoST: one end-to-end policy, height-gated rewards in training, action anchored as q + beta*a with a beta curriculum. HumanUP: two training stages with the same observation/action; the vendor's three-stage state machine is the baseline it beats (41.7% vs 78.3%); Stage II tracks an 8x slowed Stage I trajectory (4x too violent, 10x does not converge). StableMimic: a learned soft gate. On 08-10 more sources were checked the same way: the Agility page does not say Digit's self-righting was learned in simulation (only step recovery is stated as RL), so the claim was downgraded to speculation; HoST's support for the 12-DoF armless Mini Pi exists in its code repository, not in the paper text.

    Change

    Adopted only what the originals share: a hard action bound or anchored action space, strong smoothing including a second-difference term, a slowed hidden reference, heavy DR and real fallen states. The staged runtime state machine was not adopted.

    Outcome

    The next re-rooting candidates came straight from the verified material, and the beta-anchored action space (HoST, with the Mini Pi configuration as the nearest real-robot precedent) became V2, the lineage that later stood up on hardware.

    Mechanism

    A paraphrase compresses a paper into the advisor's own architecture; only the original shows what was actually deployed, what was a baseline, and which numbers came with which ablation.

    Applies when

    • an advisor, agent or summary proposes an architecture with citations
    • an industry claim ("X learned it in sim") is about to justify a design
    • several papers are cited for one combined recipe
    “⇒ **顾问的核心形态"runtime 五阶段状态机 + 每阶段参考姿态 + RL residual"在三篇引文 里均不存在**,其中 HumanUP 还点名 state machine 是局限。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 三篇引文精读判决(2026-08-09 全文核对;顾问转述与原文有出入)
  • Every gate on a penalty is an exit - to stop paying a stance tax the policy parked 2 deg outside a 30 deg uprightness gate (lunging), and, from scratch, just under a height gate (crouching); a positive band and an always-on guard fixed bothpenalty-gate-is-an-escape-hatch
    Replicatedrecoveryreward-shapingreward-shapinggate-battery

    Never gate a penalty on a state the policy can leave by getting worse; use always-on guards for what must never happen and positive, gated bands for what you want, and count every gate on a penalty as one more escape route to check in the logs.

    Symptom

    V2.9 added stance_width_task = relu(0.34 m - foot spacing) x standing gates, weight -10. Three checkpoints scored 0% on acceptance: standing height reached, feet on the ground, angular rate low, but the torso leaned 32.5/32.3/31.9 deg in a fore-aft lunge. The training dashboard read "tax paid off, base_height at full value".

    Context

    The penalty was gated by uprightness (tilt < 30 deg) and standing height; its tax had no time gate (500 steps x -1.65) while the standing income sat behind a 3 s zero gate (~325 steps), so leaning just past 30 deg lost a little gated income and saved the whole tax. base_height has no upright gate, so the lunge still collected it. The first metric also measured full horizontal spacing, so a staggered lunge counted as "wide".

    Change

    Two laws written down: a penalty may only carry gates the policy cannot escape by getting worse (make it an always-on guard) or it becomes a positive band ("not earned" is not "escaped"); and width is measured laterally in the base yaw frame. From scratch (V3.1 P1) with the lateral metric but the same gates, the policy parked just under the height gate instead (base_height 1.176/1.5, h ~ 0.30 m against a 0.3264 gate), the beta curriculum never advanced in 1,700 iterations, and the run was stopped early. P1b flipped the penalty into a positive band +2.0 x clamp(lateral / 0.34) x standing gates; P1c added yaw_guard = -5 x relu(|hip_yaw| - 30 deg), always on, no gate, no exemption.

    Outcome

    P1b: lateral stance 0.364 m, all four categories 100%, MuJoCo mu 1.0/0.4 both 100%. P1c: all six acceptance criteria passed for the first time on the line (hip-yaw saturation 1.2%), the guard's tax converging to -0.016 (almost never touched).

    Mechanism

    A gated tax that is not paid is saved, so the policy moves to the cheapest state just outside the gate; a gated income that is not earned is simply lost, so a positive band has no exit. HoST's style penalties are ungated or binary - the same law seen from the other side.

    Applies when

    • adding a penalty multiplied by an uprightness, height, phase or contact gate
    • a policy settles just beyond a gate threshold
    • training metrics look paid-up while acceptance collapses
    “**罚项的门 = 策略的逃生门**。带直立门的负项可以靠"变得更差"(倾出门外) 全时免税;HoST 的 style 罚全部无门控/二值恰是同一律的反面实证。修律: 负项只许挂"变差逃不掉"的门(上限护栏 always-on),或改正向 band (收不到 ≠ 逃掉)。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §46 定案(血统内第四败 + 两条新律)
  • A get-up policy righted itself and sat - three terms paid the seated pose 84% of the return, and the only shaping term that could tell sitting from standing was an exp kernel outputting 5e-5seated-basin-dead-exp-kernel
    Mechanism understoodrecoveryreward-shapingreward-shapingattribution

    When a policy parks in a degenerate posture, tabulate what each reward term pays that posture against the target (watch contact terms that reward touching rather than bearing load) and evaluate every exp kernel at the error actually observed; a kernel narrower than the real error is switched off, and widening it is a one-variable repair that adds nothing new.

    Symptom

    R0 converged by iteration 700 and gained 1.7% over the next 2,300; success 0.0% in all four fall categories. Three of four success conditions passed (tilt median 1.0 deg, both feet in contact 99.6%, angular rate low); height passed 0.6% (median 0.204 m against 0.326). The robot knelt in a W-sit: hip yaw +/-47 deg, knees folded to 92% of the hard limit, shins flat, pelvis on the ground, torso vertical.

    Context

    Minimal reward table: upright (1-g_z)/2 +2.0, base_height linear progress +1.5, stand_pose exp(-||q-q_stand||^2/std^2) x upright gate +1.0 with std 1.0, still +0.5 and feet_on_ground +0.5 both x the upright gate, plus regularizers. The upright gate is a hinge that opens below 30 deg of tilt. The pre-registered fallbacks were then checked against the measured state: tightening the tilt gate was falsified (tilt was already 1.0 deg); a success bonus contradicted the spec's own no-cliff-bounty rule; narrowing the categories was useless (all four converged to the same pose); raising init noise was too weak for a basin this deep. Only a half-rise intermediate state addressed it, and a cheaper repair existed.

    Change

    R0.1 (user decision, single variable): stand_pose std 1.0 -> 3.0. Not a new term and not a bounty - repairing a declared term that was numerically dead. The two runs' logged env.yaml differ in log_dir and std only.

    Outcome

    R0.1 58.2% overall (R0 0.0%): supine 91.5%, side 82.2%, mid 55.9%, prone 0/156; knees fully straight; ||q-q_stand||^2 9.99 -> 0.91 and the stand_pose term 4.6e-5 -> 0.90; get-up ~1 s, no re-falls, the curve still rising at the 3,000-iteration cap. Prone stayed at zero and needed a different fix (see prone-dead-end-is-foot-placement).

    Mechanism

    Sitting earned upright 1.98/2.0, still 0.43/0.5 and feet_on_ground 0.46/0.5 - 3.0 of a 3.57 per-second return - because feet_on_ground asked for contact, not load. The only terms separating sitting from standing were base_height (+0.70/s for standing) and stand_pose, whose kernel at the real 9.99 rad^2 error (75% of it in the two knees) was exp(-9.99) = 4.6e-5 with a gradient near 1e-4. Standing up meant risking 3.0/s to gain 0.70/s while unfolding knees at 92% of their limit under load. With std 3 the same term is exp(-9.99/9) = 0.33 - a live gradient, three quarters of it on the folded knees.

    Applies when

    • a policy converges early to an upright but low, seated or kneeling pose
    • a posture-matching exp term reads ~0 in the training logs
    • contact-based rewards saturate while the task metric does not move
    “**关键:`feet_on_ground` 只问"触地"不问"承重", 跪坐时双脚确实贴地,照样满分。** 三项 3.0/s = 总回报 3.57/s 的 84%。 … **exp(−9.99) = 4.6e-5** —— 权重 1.0 的项实际输出 5e-5、梯度 ~1e-4, **不是"还没学会",是数值上根本不存在**。 … **R0.1 决定(用户 2026-08-09 定,单变量)**:`stand_pose` 的 `std` **1.0 → 3.0**。 不是加新奖励、不是悬崖悬赏,而是**修复一个已声明但数值失效的项**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §11 R0 首跑(recovery_r0, 2026-08-09):FAIL —— 翻正了但坐着
  • Fix a too-deep nominal pose before adding any penalties - the default stance defines the basin training starts innominal-posture-before-penalties
    Mechanism understoodwalkreward-shapingreward-shapingplant-calibration

    Before tuning penalties on a degenerate gait, audit the nominal pose and height targets against morphology and published ratios; if the default stance encodes the degenerate behavior, fix it first - and recompute dependent quantities (init height) by FK, not by hand.

    Symptom

    Policy lived in a crouched shuffle; nominal knee angle was 0.5 rad (28.6 deg) - deeper than published configs (Unitree G1 0.3 rad / 17.2 deg, Booster T1 0.4 rad) - so the policy's starting point and its action-space center both sat inside the crouch basin.

    Context

    Initially ranked "secondary" in the local diagnosis, this was promoted to co-first priority by the cross-check against published reward tables, which states that with nominal knee flexion above ~0.4 rad, fixing the posture must precede adding any penalties ("改这个之前别加任何 惩罚都是白费"). Companion base-height items: walk profile had weakened base_height_l2 to -5.0 (base class -10, field standard -10 to -20, "the second most common cause of death"), and the height target must be the STANDING height (0.384), not the crouch height.

    Change

    Nominal knee 0.5 -> 0.3 rad with init_base_height recomputed by MuJoCo FK (0.3739 -> 0.3802); base_height_l2 restored to -10 with standing height target; both bundled as first-priority alongside the clearance term.

    Outcome

    Part of the v5/v6 package that lifted swing height to 34 mm and tracking to 87%; the crouch basin stopped being the default answer.

    Mechanism

    The nominal pose is the fixed point every regularizer pulls toward and the point where action=0 lands; if that point is itself the degenerate posture, every penalty fights the geometry. Correcting the attractor is prior to shaping the gradient field around it.

    Applies when

    • policy converges to a crouched or collapsed posture
    • nominal joint angles were chosen for stability rather than gait
    • base-height reward targets or weights were locally weakened
    “研究明确说"nominal 膝屈超过 ~0.4 rad 必须先改,改这个之前别加任何惩罚"。我们是 0.50,超标。… base_height_l2 在 walk profile 里被减到 −5.0(基类是 −10)。研究说这是"第二常见死因"且应 −10 ~ −20。改回 −10。目标高度用站立高 0.384 是对的(研究要求 target 必须是*站立*高度而非蹲姿)。”
    train/WALK_DIAGNOSIS.md § 修正 ①(升级优先级) / 修正 ④
  • A single-signal contact detector lied in both directions - foot height flagged 40% false flight on a walking gait, contact force alone flagged false flight during low-friction slip - so flight became force < 5 N AND sole height > 5 mmcontact-detector-single-signal-lies
    Replicatedrunsim-evalmeasurementgate-battery

    Define contact and flight events from two independent signals (force and geometry) in conjunction, validate the detector on a behaviour known not to contain the event before using it as a gate, and match the trainer's threshold when comparing across simulators.

    Symptom

    The run line needed a flight-fraction gate. The first MuJoCo version, "sole higher than 2 mm", measured 40% flight on a walking policy that never flies. Five weeks later the one-leg gate, using contact force alone, reported support-foot "flight" segments at mu 0.4 for a policy that was not hopping.

    Context

    During a walking step the toe lifts or the heel strikes with the foot pitched, so the ankle-roll origin rises a few mm while part of the sole still touches - height alone calls that flight. Switching to contact force < 5 N (the same threshold Isaac's contact reward uses) zeroed the false flight on walking. In the one-leg re-test every force-only "flight" segment had a measured sole height of 0.0 mm: the normal force chattered while slip corrections played out on low friction.

    Change

    Run line: flight = contact force < 5 N ("lift-off must be judged by contact force"). One-leg line (2026-09-16): flight = force < 5 N AND sole height > 5 mm, recorded as the same measurement lesson in the opposite direction; the gate's behavioural meaning was unchanged.

    Outcome

    With force-based detection the walking policy read 0.0 flight and the run policy's zero flight was confirmed by two plants; with the conjunctive definition the one-leg support-foot gate stopped reporting false hops.

    Mechanism

    Each signal has its own failure: geometry moves without leaving the ground (foot pitch), and forces drop without leaving the ground (slip chatter); requiring both removes both families of false positives.

    Applies when

    • writing a flight, lift-off, hop or slip detector for a gate
    • a gate reports an event the video does not show
    • reusing a detector on a different gait or floor friction
    “首版用**足底高度>2mm** 判离地, 在 omni_s1e 走路策略上测出 40% 假腾空 … 改用**接触力 <5N**(与 Isaac feet_contact_number 同源阈值)后 空检归零 (walk 策略 flight_frac 0.0)。**课文: 离地判定必须用接触力, 高度判 会把脚的俯仰当腾空**”
    train/README.md § run R1 立项 (2026-08-09): Mac 侧新工具 + 一次空检抓获
  • A walking policy's tilt cutoff is a legal state for a recovery policy - the default 45 deg fall guard had to be raised for recovery tests and is disabled once the switch owns falls, so the abort chain becomes the recovery timeout, the operator's cut, and the firmware torque limitsfall-guard-becomes-a-state
    Observed oncerecoveryreal-deployreal-acceptancehardwareprocess

    When a new skill makes a safety cutoff's trigger a legal state, replace the cutoff with a bound of the skill's own (a timeout ending in a safe stop) instead of just switching it off; keep the operator's cut and the firmware limits as independent layers, and write every flag change into the run sheet.

    Symptom

    deploy_policy's default protection stops the robot beyond 45 deg of tilt. A recovery policy starts lying at roughly 90-97 deg, so under the default it is refused on the spot - a flag the first hanging checklist forgot.

    Context

    The layers in the sources: deploy_policy's tilt cutoff (default 45 deg, a line in the safety chain); for standalone recovery tests the cutoff was raised (110 deg in the spec's A/B sheet; 181 deg, effectively off, in some runbook commands); with --recovery-policy the cutoff is disabled because a fall is now a state, not an exception, and RECOVERY lasting over 15 s ends in a safe stop (the runbook calls it the line where the spotter steps in). Independent of the policy: firmware torque limits checked at start (12/17/11 N*m, set_torque --check), the operator cutting enable at any kicking or oscillation, and in the one-leg teleop a space-bar stop that puts the foot down.

    Change

    The flag was added to the run sheets, and the FSM replaced the removed cutoff with its own bound (the timeout).

    Outcome

    The spec records the flag omission and its fix; it does not record the FSM's timeout being exercised on hardware.

    Mechanism

    A safety cutoff encodes one policy's notion of "abnormal"; a new skill whose normal operation lies beyond it either cannot run or runs with the cutoff off, and only a replacement bound keeps the chain closed.

    Applies when

    • deploying recovery, fall-damage or acrobatic skills behind existing safety checks
    • a run sheet disables a protection flag
    • listing the abort chain for a hardware session
    “--max-tilt-deg(默认 45°,安全链第 13 行写的那个)。recovery 的合法状态覆盖整个倾角域,把它抬到 181 = 实效关闭 … RECOVERY 超时 15s 会自动安全停(看护介入线)”
    RL系统/FOLLOW THIS copy 2.md § FSM 吊挂首测 ② 落地测 / #### Recovery Policy (operator runbook, undated)
  • A stand gate judged by survival passes a robot that wanders a meter - judge posture insteadstand-gate-posture-not-survival
    Mechanism understoodomnireal-acceptancegate-batteryreal-acceptance

    For every gate, ask what behavior the metric is a proxy for and validate its ordering against real observations; replace metrics whose ordering disagrees with reality, and demote them explicitly rather than silently.

    Symptom

    Real robot "standing" drifted 0.5-1.4 m across the floor while the sim stand gate scored a clean 20/20 - because the gate's metric was episode survival, which wandering does not violate.

    Context

    C2's stand condition was originally written as "survival regression <=2/20". A real-robot counter-example on 2026-08-08 forced the re-judgment: wandering robots survive. Cross-checking candidate sim metrics against real-robot feel showed max-tilt median ordering agreed with hands-on ranking, while displacement ordering was actually OPPOSITE to real impressions - so displacement was demoted to a reference quantity, not a gate.

    Change

    Stand PASS criterion rewritten from survival to posture: "stand tilt max median <= root baseline +1.5 deg"; displacement kept only as reference. Applied to all subsequent rungs (C4 and redo levels inherit it).

    Outcome

    Later rungs gated stand on tilt (e.g. C4 product: 7.7 deg vs parent 7.1 deg, +0.6 deg PASS); the wandering failure mode became visible to the battery instead of hidden by survival.

    Mechanism

    A gate metric is a proxy for an intended behavior; survival is a proxy for "did not fall", not "stood still". Metric choice must be validated against ground truth (real-robot feel/measurement), and a proxy whose ordering disagrees with reality on real data is worse than no metric - it steers selection backwards.

    Applies when

    • writing PASS conditions for stand/idle/hold behaviors
    • a gate passes policies that visibly misbehave on hardware
    • choosing between candidate metrics for an acceptance battery
    “stand 必须用位姿判,不能用存活判(2026-08-08 真机反证改判):站着乱走 0.5~1.4 m 时存活照样 20/20 —— C2 的 stand 条件原写「存活退化 ≤2/20」,选错了指标。改为: stand 倾角 max 中位 ≤ 根基线 +1.5°(倾角与真机手感排序一致;位移排序与真机相反,降为参考量)”
    train/C_LADDER_RUN.md § 3b. PASS 条件 ⚠️ stand 必须用位姿判
  • Teleop fed the sidewalk axis a command beyond its training band - feet clipped; give each axis its own speed settingteleop-command-band-per-axis
    Mechanism understoodomnireal-deployreal-acceptancehardwareattribution

    Give every command axis its own teleop scale, clamped to that axis's training band, and reproduce any hardware incident in sim with the exact deployed command values before touching training.

    Symptom

    Robot stepped on its own foot when sidewalking left under teleop - and only when going left.

    Context

    The teleop tool used one speed setting for all axes: --teleop-speed 0.20 applied to A/D sent cmd_vy = 0.20, above the training band's top (0.08-0.18) where foot-spacing margin is thinnest. Sim reproduction of the incident (product policy, pw0.8, 5 seeds x 20 s, true collision threshold = single foot width 104 mm): at vy 0.20 the minimum foot distance was 111-115 mm - 7-11 mm from self-collision - vs 147 mm at vy 0.10. Left was 4x more dangerous than right (25% vs 6% of time inside the 160 mm soft wall at vy 0.10), matching the left-only symptom; the margin did not degrade over time (pressing more just lengthened exposure).

    Change

    deploy_policy gained --teleop-side (default 0.10), separating the lateral speed from the forward speed so each axis's teleop command sits inside its own trained band.

    Outcome

    Command now inside the band with 43 mm margin at default; the incident became a quantified, reproduced, closed account rather than a mystery.

    Mechanism

    The policy's competence envelope is the training command distribution per axis; teleop mappings that share one scalar across axes silently command out-of-band inputs on the weakest axis. Asymmetric risk (left vs right) came from the policy's own chirality bias, so a symmetric command produced an asymmetric hazard.

    Applies when

    • wiring a joystick/teleop layer over a learned policy
    • a hardware incident occurs on one command direction only
    • training bands differ across command axes
    “A/D 一直与 W/S 共用速度档,所以按 A 下发的是 vy = 0.20 —— 既超训练带(0.08~0.18)上沿 … 0.20(遥控实际值)| 111~115 mm | 7~11 mm … 且左比右危险 4 倍 … 处置:deploy_policy 新增 --teleop-side(默认 0.10),侧移与前进档分开。”
    train/C_LADDER_RUN.md § 3p. 一 向左走踩到自己 → --teleop-speed 0.20 同时喂给了 vy
  • The shipped checkpoint was chosen by scanning checkpoints on the full gate - neighbours 100 iterations apart failed 1 and 38 cells, late checkpoints degraded - never by taking the last one, and training stopped on signals, not on a schedulecheckpoint-choice-is-a-full-gate-scan
    Replicatedonelegsim-evalfork-selectiongate-battery

    Choose a release checkpoint by running the full acceptance battery over a band of checkpoints (including the transfer axis), stop training on measured signals rather than a fixed iteration count, and expect adjacent checkpoints to differ sharply.

    Symptom

    Gate results moved sharply and non-monotonically between checkpoints of the same run, and the last checkpoint was often not the best.

    Context

    One-leg V0r1: the 2,000 neighbourhood was best; from 2,500 on the nominal gates degraded (late overtraining); 2,000 itself had one real micro-hop (17.7 mm over 5 frames); 2,300 was all green and shipped. V0r2: failed cells per checkpoint 2,000:19, 2,100:38, 2,200:1, 2,300:3, 2,400:27, 2,500:12, 3,000:18 - 2,200 shipped. The recovery line learned the same from the other side: stopping v2_6 early at a scheduled point left a policy whose re-fall rate had spiked to 9-22% before consolidation healed it ("stop on signals, not on the schedule"), and a continuation's transfer decayed checkpoint by checkpoint while Isaac stayed perfect.

    Change

    The acceptance rule "scan checkpoints, do not look only at the last one" is written into the one-leg gates (called the S1 discipline); release candidates are chosen from the scan.

    Outcome

    Both one-leg releases were mid-run checkpoints (2,300 and 2,200) chosen by the full 40-cell battery.

    Mechanism

    PPO keeps changing the policy after the gates saturate; with no gradient toward the gate's conditions, later checkpoints wander, so gate quality is a noisy function of iteration.

    Applies when

    • picking which checkpoint of a run to export and stamp
    • a run is stopped at a fixed iteration budget
    • final-checkpoint results are worse than mid-run smoke tests
    “Isaac 侧 S1 纪律: 验收扫 checkpoint,不是只看最后一个。 … 扫描判决: 2000 邻域最优——2500+ 标称面退化(⑤③② 散挂, 晚期过训), 2000 有一例真微跳(L s100 μ1.2, 17.7mm/5帧), 2300 全绿。”
    git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §6 验收门 / §8 核查单 5 与 7
  • Score left and right separately - averages hide chirality breaking that mirror augmentation does not preventchirality-scored-separately
    Replicatedomnigate-batterygate-batteryreal-acceptance

    Report every mirrored skill as two numbers with an explicit gap budget; never accept an average, and never assume augmentation guarantees symmetry - measure it per lineage and treat breakage as hard to reverse.

    Symptom

    Policies developed quantified left/right asymmetry (e.g. C2-700 turned right at 82% but left at 67% - a 15 pp gap; push tolerance 40/40 symmetric on the root vs 17/40 on a deep-trained descendant), and averaged metrics would have reported healthy midpoints.

    Context

    The repo had policy-level symmetry-breaking evidence strong enough to make separate scoring a battery rule: "左右必须分开打分 … 平均 vy 跟踪会把它掩盖". Notably, chirality broke and never recovered even though mirror augmentation (command-level mirror_prob 0.5) was on the whole time - augmentation reduced but did not prevent asymmetry, and once broken it stayed broken through subsequent rungs. PASS conditions therefore carried explicit symmetry budgets (left/right tracking gap <=10 pp), and sim's predicted asymmetry (700: right faster than left) was flagged for direct real-robot timing confirmation.

    Change

    Battery rule: every directional skill reports left and right (CW/CCW) as separate rows with a max-gap budget; mirror augmentation treated as mitigation, not proof of symmetry.

    Outcome

    The 700-vs-A800 asymmetry gap (15 pp vs 7 pp) became a first-class selection criterion; C4 product shipped with a measured 5 pp gap.

    Mechanism

    Averaging over mirrored conditions cancels antisymmetric error exactly where it matters; and symmetry lost during training is a lineage injury (like plasticity loss) that later rungs do not spontaneously heal, so it must be gated, not assumed.

    Applies when

    • evaluating turn/sidewalk/push-recovery or any mirrored skill
    • relying on mirror/symmetry augmentation
    • selecting between checkpoints with similar average scores
    “左右必须分开打分(left/right lateral、CW/CCW turn 各自一行)—— 本仓已有 policy-level symmetry breaking 的量化证据,平均 vy 跟踪会把它掩盖。”
    train/C_LADDER_RUN.md § 5. 固定验收矩阵 (左右分开打分)
  • Three hardware accounts locked the run design point - and the knee's real speed ceiling is tau_limit/kd, not the firmware limitfeasibility-accounts-lock-design-point
    Mechanism understoodrunplant-calibrationplant-calibrationactuator-modelinghardwareprocess

    Before opening a dynamic-gait training line, compute the full account set - tau_limit/kd effective speed ceilings, joint ROM under the intended reference geometry, and thermal RMS at the duty cycle - and let the accounts lock the design point; move only to pre-registered in-table alternates, re-running the accounts first.

    Symptom

    The run line was believed to require a firmware raise of the RS06 speed limit (10 rad/s) as a hard precondition, and the feasibility script's motor-envelope scan had marked 80/100 mm foot-lift cells "physically feasible".

    Context

    Three added accounts re-decided everything. (1) Damping tax: in MIT mode tau = kp*(q_des-q) - kd*qd, so sustained rotation is capped at tau_limit/kd = 12/1.5 = 8 rad/s - below the firmware's 10; at peak speeds 6.7-7.9 rad/s the damping term alone eats 10.1-11.9 N*m (84-99% of the torque limit). "提固件 limit_spd 越不过这道税 —— 它是 kd 与限扭的比,不是固件旋钮." (2) Joint ROM: the feasibility script had checked motor envelopes but NOT joint range - the ankle-pitch ROM caps 1:2:1 leg-shortening lift at 62 mm (soft) / 77 mm (hard), so the 80/100 mm "feasible" cells were voided; also firmware-independent. (3) Ankle thermal: duty 0.40 puts ankle RMS at 87% of continuous rating (0.35 -> 93%); long-period big-stride cells hit both ankle torque peak and heat. Verdict: firmware raise DEQUEUED (50 mm design point needs knee 6.7-7.3 < the 8 rad/s effective ceiling < firmware 10); vel_limit stays 10 so sim == robot. The three accounts uniquely lock the design point - 50 mm lift / T 0.60 s / duty 0.40 - "三笔账 唯一锁定,不是调参空间", with pre-registered alternates allowed only inside the table and only after re-running the accounts.

    Change

    Design point frozen from accounts; hardware precondition reversed by arithmetic rather than by test; reference amplitude (0.84 rad = FK inverse of 50 mm) derived, per-joint action scales sized to the required travel (knee 0.9, hip_pitch 0.6, ankle deliberately NOT amplified - hard limit is adjacent).

    Outcome

    A firmware work item left the critical path; an infeasible region of the design space was closed before any training; the remaining risk (knee tracking lag from the damping tax) was pre-registered with its own criterion and in-table fallback (duty 0.35) - "这不是'奖励没调好', 是 plant 账".

    Mechanism

    PD actuators in MIT mode pay kd*velocity out of the same torque budget that tracks position, so the effective speed ceiling is a ratio of configuration constants, invisible to firmware settings; and feasibility is the intersection of ALL constraint families (torque envelope, joint ROM, thermal RMS) - a scan that omits one family certifies impossible cells.

    Applies when

    • planning running/jumping or any high-rate gait on PD actuators
    • a firmware or hardware upgrade is assumed as a training precondition
    • a feasibility scan covers motor limits but not ROM or heat
    “膝的有效速度顶 = τ_limit/kd = 12/1.5 = 8 rad/s,不是固件的 10。… 提固件 limit_spd 越不过这道税 —— 它是 kd 与限扭的比,不是固件旋钮。… 可行性脚本只查了电机包络没查关节 ROM —— 其 80/100mm 的"物理可行"格作废。… 判决:RS06 提固件对 run v0 不是前置,出队”
    train/RUN_V0_SPEC.md § 1. 硬件账判决 / 2. 步态设计点

Next page