Skip to content

Training Coach

The coach reads a training run and returns a diagnosis, in-whitelist proposals and an experiment plan — and every claim cites one of the cards below. This is that corpus: 22 methodology rules distilled from a real sim2real programme, and 171 episode cards behind them, each with the sentence in the war history it came from.

Doctrine

A report may cite any of these as doctrine-N.

  1. doctrine-1Contract freeze and fingerprint discipline

    The policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.

    Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).

    Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.

    contract-freeze-and-checkerlegacy-profile-pinningderived-asset-staleness-checkgain-profile-belongs-in-the-stamprecovery-two-policies-and-a-state-machinewalk-recovery-fsm-handoff

  2. doctrine-2Attribution by resolved training params - never eval-override knobs

    Capability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.

    Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).

    Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.

    kd-bandwidth-mu-law-attributiondeploy-knob-attribution-before-retrainingcycle-time-override-is-oodresolved-config-is-source-of-truth

  3. doctrine-3PASS gates become constraints; FAIL gates become objectives

    Once a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.

    Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.

    Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".

    fixed-acceptance-matrix-per-rungpreregistered-stop-criteria-per-rung

  4. doctrine-4One variable per ladder rung - counted against what the checkpoint saw

    A rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.

    Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.

    Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.

    resume-state-dr-auditorthogonal-batch-with-ablation-orderfill-the-missing-factorial-cell

  5. doctrine-5Pre-register risks, readings, and stop criteria before the ladder

    Before a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.

    Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).

    Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.

    preregister-risks-and-fork-readingspreregistered-stop-criteria-per-rungpreregistered-real-expectationsfeasibility-accounts-lock-design-point

  6. doctrine-6Plant parameters are measured, never invented

    Every plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.

    Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).

    Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.

    friction-measured-not-guessedarmature-n2-rotor-inertiatorque-limit-shape-by-measured-peakslatency-lerp-reverse-extrapolationfeasibility-accounts-lock-design-pointsim-api-friction-columnsget-up-feasibility-accounts-before-trainingsingle-support-gain-authority-probe

  7. doctrine-7Sim2sim gate before sim2real - under deployment conditions

    Every checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).

    Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).

    Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.

    sim2sim-gate-before-sim2realeval-plant-honesty-contact-paramspipeline-latency-is-plant-not-drbody-frame-velocity-api-audittorque-penalty-bought-by-leg-bracingtorque-disagreement-between-simulators-unresolved

  8. doctrine-8Observation honesty - the actor's inputs are a hardware contract

    The actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.

    Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).

    Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.

    observation-honesty-critic-onlyhistory-obs-needs-plant-variationreward-observability-limitdeploy-heading-loop-and-align-training

  9. doctrine-9Reward economics are audited in realized currency

    Reward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.

    Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).

    Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.

    realized-contribution-auditreward-cost-of-ignoring-auditgate-new-reward-terms-by-commandignore-floor-diagnosiscalibrate-threshold-between-healthy-and-sickcalibration-threshold-with-withdrawal-clauseinert-reward-term-auditseated-basin-dead-exp-kerneltail-torque-needs-hinge-on-computed-demand

  10. doctrine-10The zero-cost option must be the desired behavior

    For every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.

    Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).

    Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.

    penalize-the-slip-not-the-jointsaturation-cheating-zero-rate-costmoving-gate-42x-stand-taxtermination-closes-degenerate-basinpenalty-gate-is-an-escape-hatchsoft-limit-penalty-charges-nominal-poseunpriced-foot-attitude-is-a-free-variableenumerate-cheapest-cheats-before-trainingbinary-band-reward-fake-touchdown

  11. doctrine-11Measurement discipline: independent referees, signs, distributions

    A disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.

    Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).

    Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.

    independent-referee-for-metric-disputesbody-frame-velocity-api-auditsame-sign-response-is-yaw-biasmedian-hides-bimodal-distributionratio-metrics-need-absolute-checkheading-integral-not-body-ratesame-distribution-reward-comparisonmultiseed-sign-test-for-driftsingle-impulse-recovery-is-chaotic

  12. doctrine-12The deployment pipeline is plant

    Irreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.

    Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).

    Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.

    pipeline-latency-is-plant-not-drpower-scale-hurts-nonforward-axesdeploy-scaling-not-training-equivalentteleop-command-band-per-axislatency-dr-covers-measured-pipelinedeploy-rate-limiter-windupslew-anchor-is-an-integratorbeta-anchored-action-targetpower-derating-cuts-full-range-contract

  13. doctrine-13DR budget is finite; its distribution is the measured support

    Robustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.

    Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).

    Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.

    push-dr-conditional-budget-conservationdr-tail-plant-continuationconstant-value-dr-overfits-margintask-shaping-before-plant-hardeningcom-randomization-forces-leg-spreadcom-dr-rollback-on-symptomthin-dr-judged-by-channel-coveragefriction-priority-re-measured-after-plant-change

  14. doctrine-14Gates measure what hardware feels: posture, margins, stripped assists

    Acceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.

    Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).

    Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.

    task-metrics-vs-posture-metricsstand-gate-posture-not-survivalconstant-value-dr-overfits-marginsuspension-probe-removes-free-stabilizerhip-roll-sum-predicts-lateral-driftchirality-scored-separatelylow-speed-commands-reward-draggingend-state-confusion-matrixfrozen-acceptance-distribution-and-pinned-seedvideo-as-acceptance-recordepisode-length-bounds-what-a-gate-seesremoved-wall-returns-on-hardware

  15. doctrine-15Fork and root selection: recoverability, maturity, frozen rewards

    Choose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.

    Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).

    Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.

    fork-root-recoverable-shortfallroot-maturity-vs-product-qualityfine-tune-reward-change-falsified

  16. doctrine-16Curricula: verified engagement, lineage counters, disease-phase gating

    Automatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.

    Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).

    Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.

    auto-curriculum-engagement-checkcurriculum-counter-lineage-stepsgate-penalties-to-the-disease-phaseaggregate-metrics-mask-subgroup-failurebucket-share-is-not-a-gradient-levercurriculum-criterion-conditioned-on-lagging-categoryper-step-income-drives-speed-time-gatetime-gate-vs-wide-stance-retire-the-fix

  17. doctrine-17Probe before training: feasibility first, hypotheses in tables

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.

    Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).

    Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).

    open-loop-probe-before-reward-tuninghypothesis-table-code-auditperiod-doubling-evidence-racesuspended-test-isolates-actuator-authorityconfiguration-probe-wall-not-slopeprone-dead-end-is-foot-placementdof-vel-penalty-is-not-a-pacing-knobamplitude-cut-falsified-yaw-fix

  18. doctrine-18External advice is recomputed locally; values transfer as ratios

    Every external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.

    Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).

    Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.

    external-advice-audit-against-own-arithmetictransfer-ratios-not-absoluteslatency-dr-covers-measured-pipelinereference-structure-fk-amplitude-divisioncycle-average-tracking-for-gait-quantitiesadvisor-paraphrase-vs-paper

  19. doctrine-19Hardware sessions are scripted experiments, not tuning sessions

    Real-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.

    Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).

    Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.

    risk-ordered-real-deploymentbattery-bracketed-real-abknow-zero-command-behaviorpush-test-chirality-protocolno-field-tuning-protocolreversible-single-variable-field-experimentssim-veto-needs-real-confirmationfirst-real-get-up-violent-stage-one-policystaged-hang-mat-floor-for-get-uppower-cycle-preflighttwo-machine-config-disciplinefall-guard-becomes-a-statehardware-log-is-the-attribution-input

  20. doctrine-20Close questions in writing; restart when the debt is structural

    Audited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.

    Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).

    Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.

    frozen-verdicts-semantic-boundariesstale-verdicts-under-old-stackzero-offset-calibration-shifts-envelopeplant-swap-invariants-vs-shiftsfreeze-lineage-fix-structure-restartminimal-reward-table-with-provenancewrite-hardware-verdicts-back

  21. doctrine-21Name the quantity in the space it lives in

    A goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.

    Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).

    Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.

    joint-space-proxy-for-task-space-quantitycontact-detector-single-signal-lieszero-partial-credit-is-not-an-iteration-problemheading-integral-not-body-rate

  22. doctrine-22Continuation needs a live gradient; a release is chosen by a scan

    Continue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.

    Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).

    Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.

    converged-continuation-is-poisoncheckpoint-choice-is-a-full-gate-scanstance-decided-by-get-up-pathstop-stacking-roll-back-and-auditcurriculum-history-is-part-of-the-productcontinuation-budget-not-from-zeroroot-maturity-vs-product-quality

Experience cards

53 cards matching “feasibility-accounts-lock-design-point”.

  • Three hardware accounts locked the run design point - and the knee's real speed ceiling is tau_limit/kd, not the firmware limitfeasibility-accounts-lock-design-point
    Mechanism understoodrunplant-calibrationplant-calibrationactuator-modelinghardwareprocess

    Before opening a dynamic-gait training line, compute the full account set - tau_limit/kd effective speed ceilings, joint ROM under the intended reference geometry, and thermal RMS at the duty cycle - and let the accounts lock the design point; move only to pre-registered in-table alternates, re-running the accounts first.

    Symptom

    The run line was believed to require a firmware raise of the RS06 speed limit (10 rad/s) as a hard precondition, and the feasibility script's motor-envelope scan had marked 80/100 mm foot-lift cells "physically feasible".

    Context

    Three added accounts re-decided everything. (1) Damping tax: in MIT mode tau = kp*(q_des-q) - kd*qd, so sustained rotation is capped at tau_limit/kd = 12/1.5 = 8 rad/s - below the firmware's 10; at peak speeds 6.7-7.9 rad/s the damping term alone eats 10.1-11.9 N*m (84-99% of the torque limit). "提固件 limit_spd 越不过这道税 —— 它是 kd 与限扭的比,不是固件旋钮." (2) Joint ROM: the feasibility script had checked motor envelopes but NOT joint range - the ankle-pitch ROM caps 1:2:1 leg-shortening lift at 62 mm (soft) / 77 mm (hard), so the 80/100 mm "feasible" cells were voided; also firmware-independent. (3) Ankle thermal: duty 0.40 puts ankle RMS at 87% of continuous rating (0.35 -> 93%); long-period big-stride cells hit both ankle torque peak and heat. Verdict: firmware raise DEQUEUED (50 mm design point needs knee 6.7-7.3 < the 8 rad/s effective ceiling < firmware 10); vel_limit stays 10 so sim == robot. The three accounts uniquely lock the design point - 50 mm lift / T 0.60 s / duty 0.40 - "三笔账 唯一锁定,不是调参空间", with pre-registered alternates allowed only inside the table and only after re-running the accounts.

    Change

    Design point frozen from accounts; hardware precondition reversed by arithmetic rather than by test; reference amplitude (0.84 rad = FK inverse of 50 mm) derived, per-joint action scales sized to the required travel (knee 0.9, hip_pitch 0.6, ankle deliberately NOT amplified - hard limit is adjacent).

    Outcome

    A firmware work item left the critical path; an infeasible region of the design space was closed before any training; the remaining risk (knee tracking lag from the damping tax) was pre-registered with its own criterion and in-table fallback (duty 0.35) - "这不是'奖励没调好', 是 plant 账".

    Mechanism

    PD actuators in MIT mode pay kd*velocity out of the same torque budget that tracks position, so the effective speed ceiling is a ratio of configuration constants, invisible to firmware settings; and feasibility is the intersection of ALL constraint families (torque envelope, joint ROM, thermal RMS) - a scan that omits one family certifies impossible cells.

    Applies when

    • planning running/jumping or any high-rate gait on PD actuators
    • a firmware or hardware upgrade is assumed as a training precondition
    • a feasibility scan covers motor limits but not ROM or heat
    “膝的有效速度顶 = τ_limit/kd = 12/1.5 = 8 rad/s,不是固件的 10。… 提固件 limit_spd 越不过这道税 —— 它是 kd 与限扭的比,不是固件旋钮。… 可行性脚本只查了电机包络没查关节 ROM —— 其 80/100mm 的"物理可行"格作废。… 判决:RS06 提固件对 run v0 不是前置,出队”
    train/RUN_V0_SPEC.md § 1. 硬件账判决 / 2. 步态设计点
  • Prove an armless get-up exists before training it - a connected static domain, 25% torque on the cheapest path, an 8 mm hand-over gap, a static roll-over - and write down what each scan cannot representget-up-feasibility-accounts-before-training
    Mechanism understoodrecoveryplant-calibrationplant-calibrationhardwareprocess

    Before training a get-up or any multi-contact skill, compute the quasi-static accounts - connectivity of the static domain, torque along the cheapest path, hand-over gaps, COM shift available for rolling - and state which configurations each scan cannot represent; when a policy gets stuck in one of those, extend the scan before blaming the reward.

    Symptom

    A torso-and-legs robot has no arms to push off the ground; whether it can get up from the floor at all was unknown when the line opened.

    Context

    recovery_feasibility.py ran three accounts before any training (the run line's "hard accounts first" discipline): a sagittal quasi-static scan (0.05 rad grid, 44,520 configurations, MuJoCo FK, flat-foot assumption). (1) The static standing domain (COM over the feet, torques in limit) has 29,586 cells, flood-fill connected with no islands, from a 0.097 m deepest squat to the 0.384 m stand. (2) The minimum-torque path peaks at 25% of the limits (knee 2.9/12, ankle 3.7/17 N*m) - a 4x margin. (3) All 508 ground-contact configurations have the contact behind the COM; the smallest gap to pure foot support is 8 mm. A roll-over account: swinging both straight legs to one side shifts the COM 96 mm against a 62 mm torso half-width - 1.6x, so rolling needs no momentum. Three conclusions were written down for later attribution: the legs are 80% of the mass (swinging them moves the COM), prone has no flat-foot hand-over face (merge into a supine/side sit first), and supine needs no sit-up (hip flexion is limited to 75 deg).

    Change

    The accounts gated opening the line and were cited in every later argument about what the robot can physically do.

    Outcome

    They held where they applied: in V1.0 every fall category was righted under a hard rate limit, which the spec records as the quasi-static roll-over account verified by training, and the 25% torque path was the basis for pursuing a slow get-up. They also misled once: account (3) is sagittal, and on 08-09 the spec corrected its scope - it cannot represent the splayed W-sit where the policy actually stalled. A follow-up prone hip-ROM scan (471,625 cells) found 3,912 two-foot-contact cells and none with both soles within 25 deg of level (best 40.2 deg): a flat-footed push-up from prone is infeasible on this robot, so the fix became where the feet go after sitting up.

    Mechanism

    A get-up needs a connected path through statically feasible configurations and enough torque along it; quasi-static accounts bound both cheaply, and momentum can only make the real problem easier. A reduced-dimensional scan, though, only speaks for the configurations it can express.

    Conflicts

    In R0.1-R0.2 the spec read account (3)'s "prone has no front hand-over" as "prone lacks the roll-over skill"; R0.3's confusion matrix showed prone had righted its torso 159/159, and the spec then restricted account (3) to the sagittal configurations it models.

    Applies when

    • opening a get-up, recovery or climbing skill on a new robot
    • a robot lacks arms or other obvious contact options
    • a policy stalls in a configuration a feasibility scan never modelled
    “本机 **torso + legs、无手臂可撑地**,开训前先证明存在不依赖手臂的物理解 … 双腿同侧直腿摆最大横移 **96 mm = 1.6×** … **静态摆腿即可翻身,无需动量**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §1 机械可行性判决(2026-08-09,recovery_feasibility.py 三笔账)
  • The smoke watcher panicked at the wrong operating point and its score goes blind when gates saturate - treat it as a survival sentinel, not a judgesmoke-watcher-operating-point
    Replicatedomnisim-evalmeasurementprocessgate-battery

    Configure every automated evaluator at the lineage's declared deployment operating point, give it graded metrics that cannot saturate, and until then scope its authority to catastrophe-detection - never let a mis-configured watcher stop or rank a rung on its own.

    Symptom

    Two watcher misfires in one ladder: (1) during the friction rung the watcher (evaluating at kd 1.0) reported panic-level 1/3 survival from iter 3300 - falsified by the official kd 1.2 scan, because the lineage's design operating point was kd 1.2 and the watcher lacked the --kd-scale passthrough; (2) during the PD rung the watcher's early-stop score froze at iter 1050 despite ongoing drift improvements, because with all eight gates passing (constant 0/3 failures) the score has no gradient left - "八门全过恒 0/3 时 score 对漂移改善盲, s1f 课文三现".

    Context

    Both are the same category: the in-training smoke loop is an instrument with its own configuration (operating point, score design), and its verdicts are only as aligned as that configuration. The booked doctrine: "冒烟只当存活哨兵" - until the watcher evaluates at the deployment operating point with graded metrics, its role is detecting catastrophes, not ranking checkpoints; ranking belongs to the full battery at the design operating point (and the watcher's scoring was separately patched to weight survival 3x so recoveries during hard phases are not early-stopped away).

    Change

    Watcher debt booked (--kd-scale passthrough); score saturation acknowledged with graded columns planned; selection authority kept with 20-seed batteries at the declared operating point.

    Outcome

    A false panic did not abort a rung that was actually passing at its design point; a frozen score did not hide real drift gains; the instrument's authority was scoped to what its configuration can actually see.

    Mechanism

    An evaluator is itself configured (gain profile, delay, metrics); evaluating a policy away from its design operating point measures a counterfactual robot, and bounded scores saturate once binary gates pass, losing all sensitivity. Instruments need the same operating-point discipline as deployments and graded outputs to retain gradient.

    Applies when

    • an automated smoke loop contradicts the official battery
    • early-stop scores freeze while graded metrics still improve
    • lineages with non-default deployment gain/delay profiles
    “watcher (kd1.0 口径) 3300 起 1/3 恐慌被 kd1.2 正式扫描证伪为考纲外假象 —— 工作点评测口径教训: watch_ckpt 缺 --kd-scale 透传 (待补), 冒烟只当存活哨兵。… watcher score 饱和误停 @1050(八门全过恒 0/3 时 score 对漂移改善盲, s1f 课文三现)”
    train/README.md § omni_s2e_fric (watcher 恐慌被证伪) / omni_s2e_pd (500 臂)
  • The best checkpoint to SHIP is not the best checkpoint to CONTINUE FROM - maturity is capital against adaptation shockroot-maturity-vs-product-quality
    Mechanism understoodomnifork-selectionfork-selectioncurriculumgate-battery

    Decide shipping points and fork roots separately: gates rank products, but a root candidate must prove itself by surviving a continuation under the next rung's shift (dual-arm if in doubt) - and prefer the more-trained point as root when product metrics conflict with maturity.

    Symptom

    A band re-audit found s1e-300 beat the incumbent root s1e-500 on nearly every quality gate (stepping 19/20 vs 13/20 with historically-best 26.9 mm swing, speed gate 14/20 vs 2/20, heading 26 vs 54 deg/20 s) - suggesting the root had been mis-picked and the younger point should take over.

    Context

    The dual-arm control settled it the other way: continuing the S2 PD rung from s1e-500 adapted smoothly (3/3 smoke throughout), while the b300 control arm (same config, from s1e-300) fell into a survival valley under the PD shock (+100 iters: 1/3 -> 0/3), never climbed out within budget, and its 800-iter product scored 13/20 survival - eliminated. Verdict: "幼年点自身指标再好也扛不住新 DR 适应冲击, 成熟度是本钱,s1e-500 根被数据背书" - a young point's own metrics, however good, do not survive new-DR adaptation shock; maturity is capital. The audit still yielded value: the band scan (200-1000, per-100) mapped the lineage's arc (200 dragging -> 300 peak -> 400+ decay -> 900+ drift blowout), and 300 remains the better PRODUCT answer for shipping-as-is questions.

    Change

    Selection doctrine split into two questions with different answers: best-product point (quality gates at the point itself) vs best-root point (survives adaptation shocks; more training age = more capital), each decided by its own evidence - and root claims settled by a dual-arm continuation test, not by point metrics.

    Outcome

    s1e-500 kept the root role with data behind it; the S2e ladder built on it passed rung after rung, while the b300 line was closed at the cost of one control arm.

    Mechanism

    Early checkpoints sit near sharp optima with less accumulated robustness structure; their headline metrics reflect the narrow training distribution, not resilience to distribution shifts. A continuation rung is itself a distribution shift, so the root property being selected for is shock tolerance - observable only by actually continuing, never by static gates.

    Applies when

    • a younger checkpoint outscores the current root on quality gates
    • choosing the base for a robustification or command ladder
    • a continuation run stalls in an early survival valley
    “b300 对照臂 … PD 冲击下存活谷(+100 起 1/3→0/3),预算尽未爬出,800 档 20-seed 存活 13/20 出局——幼年点自身指标再好也扛不住新 DR 适应冲击,成熟度是本钱,s1e-500 根被数据背书”
    train/README.md § omni_s2e_pd (b300 对照臂) / s1e 选点重审
  • Size each joint's action authority to its measured working range - lock a channel to zero only when its job is provably elsewhereper-joint-action-scale-lockdown
    Observed oncewalkobservation-designobservation-honestyreward-shapingcontract-freeze

    Set per-joint action scales from measured target ranges and momentum decompositions: full authority for working channels, working-range authority for balance channels, zero for channels whose contribution is measured negligible - and prefer this structural quieting over perpetual reward penalties, keeping the contract dimensions intact.

    Symptom

    "Walks crooked" - roll and yaw channels wandered; hip_yaw peak-to-peak reached 22.3 deg in v11 while contributing essentially nothing to locomotion; a uniform action scale of 0.5 gave every joint the same authority regardless of its actual job.

    Context

    The v12 design replaced the scalar action scale with per-joint scales justified by measurements: pitch-class 0.5 (the gait's entire working channel - untouched); hip_yaw 0.0 - lossless because it carries only 1.4% of yaw momentum (v6 decomposition) and straight-line targets use merely +/-0.006-0.03 ("乱动纯属浪费"); roll 0.2 - NOT zero, because lateral balance and weight transfer are roll's unique job (locking it would degenerate into foot-edge rocking, "比现在更歪"), and 0.2 covers the measured working range +/-0.06-0.19 while "0.5 的另一半全是 '歪歪扭扭'的来源". Costs were accepted consciously: turning demoted to an observation row with a fallback (yaw 0 -> 0.1 in v13). Contract preserved: the 12-dim action interface unchanged, yaw values simply neutralized. The structural lockdown also RETIRED the reward-side hip_yaw_quiet penalty - "A 的 yaw=0 结构性取代,不再付奖励塑形成本".

    Change

    action_scale_joint pitch 0.5 / roll 0.2 / yaw 0.0 wired through robot.yaml -> policy_io (verified: all-ones action gives hip_yaw target exactly 0) -> Isaac action term, guarded by the contract checker ("它就是抓这种双侧不一致的").

    Outcome

    Designed and verified on the shared side before the lineage freeze; stands as the pattern for authority sizing: structure replaces reward shaping wherever a channel should simply not act.

    Mechanism

    Action scale is a per-channel authority budget; uniform budgets give noise channels the same voice as working channels, and reward-side quieting then pays a permanent shaping tax for what a zero scale provides for free. But zeroing is only lossless when decomposition proves the channel's contribution negligible AND no unique function (balance) lives there.

    Conflicts

    Wired and verified on the config/deploy side but never trained - the 2026-08-05 reset suspended v12 before the Isaac-side run.

    Applies when

    • some joints wander without contributing to the task
    • a quieting penalty (deviation/L1) taxes every step forever
    • deciding action-space authority for a new task or robot
    “yaw=0 是无损的:实测它只贡献 1.4% 偏航动量、直行目标只 ±0.006~0.03,乱动纯属浪费。… roll 不能为 0:横向平衡/重心换脚是它的独有职责,锁死会退化成脚缘摇摆(比现在更歪)。0.2 的依据:各代实测 roll 目标只用 ±0.06~0.19,0.5 的另一半全是"歪歪扭扭"的来源。”
    train/WALK_V12_SPEC.md § 3. A —— 逐关节动作幅度(用户"只动 pitch"的安全版)
  • The deploy-side walk/recovery switch - into recovery at tilt > 65 deg held 0.3 s, back at tilt < 15 deg with angular rate < 1 rad/s and straight knees held 1 s, a 15 s timeout, last action cleared both ways - and the handoff steps the design required are only partly implementedwalk-recovery-fsm-handoff
    Observed oncerecoveryreal-deployreal-acceptancecontract-freezeprocess

    Specify a deploy-time controller switch as hysteretic, time-filtered predicates the robot can measure (proxy what it cannot, e.g. straight knees for height), a timeout that ends in a safe stop, and a complete handoff (history, clock, last action, command ramp) - then test that the code performs every handoff step, because the design document is not the implementation.

    Symptom

    With a recovery policy and a locomotion policy as separate networks, the robot needs a switch: when is it "fallen", when is it "up", and what state must be reset so the next policy does not act on the previous one's history.

    Context

    The 08-09 design: enter recovery when fallen (tilt > 55 deg or height < 0.60 x 0.384 m) for 150 ms, leave for a stand-hold when upright (tilt < 12 deg, height > 0.85 x 0.384 m, feet steady, |omega| < 0.8) for 400 ms - wide entry, strict exit, hysteresis - then a mandatory handoff trio before walking resumes (reset the walking policy's observation history, restart its phase clock at 0, clear its previous action and latency buffer) and a command ramp instead of a jump. The 08-14 implementation in deploy_policy (--recovery-policy): each policy under its own manifest contract (walk: nominal + scale; recovery: beta-anchored), the same rl_default gains, the tilt cutoff disabled; RECOVERY at tilt > 65 deg for 0.3 s; LOCO again at tilt < 15 deg and |omega| < 1 and knees straight (< 0.35 rad) for 1 s - the robot's computer has no height estimate, and straight knees stand in for height so a V3.0-style upright kneel cannot pass as standing; RECOVERY longer than 15 s ends in a safe stop; last_action cleared on both switches; command forced to 0 during RECOVERY; power scaling applies to LOCO only.

    Change

    An open account was written down with it: the recovery end state (0.355 m stance, hip yaw -/+27 deg) is outside the walking policy's training start distribution, so the first test must use stand / zero command as LOCO, and the long-term fix is to widen the walking policy's initial states rather than bend recovery's stance to suit walking.

    Outcome

    The spec records only a successful compile (py_compile); the hang test was left to be done on site. The operator runbook carries the three-step procedure (hang with stand as LOCO, mat and push, then omni walk as LOCO) but no outcome.

    Mechanism

    Two policies trained separately each assume their own history, clock and last action; a switch that carries any of them across feeds the next policy a state it never saw - the design called this "the walking policy seeing a ghost history".

    Conflicts

    The 08-09 design requires resetting history, clock and previous action plus a command ramp; the 08-14 implementation records last_action clearing and a zero command during recovery; the one-leg spec of 2026-09-14 lists the handoff hygiene as specified but not implemented - reset_history() is called by nothing (a 215-dim policy would carry four frames of pre-fall history), the phase clock is not zeroed, and there is no command ramp back to LOCO. No hardware run of the FSM is recorded in either source.

    Applies when

    • switching between separately trained policies on hardware
    • a policy with history or phase observations is re-enabled mid-run
    • the robot lacks a sensor the switching criterion was designed around
    “**判据**:进 RECOVERY = 倾角 >65°(`--fall-tilt-deg`)持续 0.3 s;回 LOCO = §5 真机可测子集:倾角 <15° ∧ |ω|<1 ∧ **膝直 <0.35 rad(NX 无高度观测, 高度门用膝直代理 —— 防 V3.0 型"跪坐但直立"误判)** 持续 1 s (`--recover-hold`);RECOVERY 单次 >15 s(`--recovery-max-time`)安全停。 … **切换卫生**:两向切换 last_action 清零;RECOVERY 态 cmd 强制 0”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §50 FSM 双策略调度(2026-08-14,用户令):deploy_policy --recovery-policy
  • Single-impulse push recovery is a binary chaotic quantity - cross-machine floating-point divergence can flip the outcomesingle-impulse-recovery-is-chaotic
    Mechanism understoodomnisim-evalmeasurementgate-batteryattribution

    Never gate or compare single-event recovery outcomes across machines or domains: evaluate disturbances as survival distributions over phases and seeds, compare longitudinally on one machine, and treat any single-point cliff as unconfirmed until it survives the statistical protocol.

    Symptom

    Mac evaluation found a hard "0.8 N*s cliff" (0/3 survival) that the training machine flatly contradicted: the identical protocol (0.8 impulse at 8 s, cmd 0.2) survived 3/3 there, and a 0.6/0.8/2/4 cross sweep survived everything.

    Context

    The verdict became a named lesson ("跨机混沌课文"): whether one specific push at one specific phase is survived depends on a trajectory that diverges across machines from floating-point differences alone - "单次冲量恢复是二值混沌量, 跨机浮点发散可翻结局". The boundary was drawn precisely: the 20-seed statistical gates DO agree across machines (established precedent), but that agreement cannot be extrapolated to single-point recovery tests. Protocol amended: disturbance evaluation uses multiple push phases (8/10/12 s), >=10 seeds, and only same-machine longitudinal comparisons; the Mac-side recommendation built on the unreproducible cliff was not adopted, while its directionally-consistent small-impulse data was kept.

    Change

    Push evaluation redefined from single-event pass/fail to multi-phase multi-seed statistics, with cross-machine comparison banned for event-level results and allowed for distribution-level ones.

    Outcome

    A false hardware-relevant "cliff" was prevented from steering the ladder (the s2e push rung decisions were made on same-machine statistics); the chaos lesson was cited again when real push tests were restricted to qualitative cross-domain use.

    Mechanism

    Perturbation recovery near the viability boundary has sensitive dependence on initial conditions; different BLAS/GPU reduction orders yield different trajectories from identical configs, so a binary outcome at one phase is machine-specific noise. Averaging over phases and seeds restores a quantity whose expectation is machine-stable.

    Applies when

    • a push/disturbance result differs between machines or sim and real
    • designing push-recovery acceptance tests
    • a sharp pass/fail cliff appears in a chaotic-regime evaluation
    “训练机上 Mac 原协议 (0.8 @8s cmd0.2) 3/3 全活 … 与 Mac 的 +0.8 0/3 直接矛盾。定性: 单次冲量恢复是二值混沌量, 跨机浮点发散可翻结局;统计门 (20-seed 八门) 跨机吻合的先例不能外推到单点恢复测试。协议改判: 抗推评测多相位 (push 时刻 8/10/12s) + ≥10 seed + 只做同机纵向比”
    train/README.md § s2e 支线终章 (跨机混沌课文)
  • Before training a one-leg stand, the accounts and a probe showed the default gains could not hold it at all - kp 20 needs 0.39 rad of error to carry the static roll moment, more than the whole adduction range - so per-joint gains came first, and thermal limits set the session lengthsingle-support-gain-authority-probe
    Mechanism understoodonelegplant-calibrationactuator-modelingplant-calibrationhardware

    Before training a posture that loads one joint statically, compute the steady tracking error load/kp and the series stiffness against m*g*h, and prove with a simple hand-written controller that the posture can be held under the deployment gains - change the gains first if it cannot; then size session length from the thermal account.

    Symptom

    The one-leg line (standing on one foot, the other folded back, no hopping) had to decide whether the existing gain profile could hold single support before any reward was designed.

    Context

    Hardware accounts (9.792 kg, COM 0.234 m high, 170 x 80 mm feet, legs 80% of the mass): moving the COM over one foot needs 107 mm of shift and the 20 deg hip-roll adduction range gives 131 mm - geometrically enough. The static frontal moment is 7.8-9 N*m, within RS02's 17 N*m - torque is enough. But at kp 20 carrying 7.8 N*m needs 0.39 rad of tracking error, more than the entire adduction range, and the real robot had already shown it: commanded +0.17, actual -0.04 (0.21 rad droop) under load, 0.0008 rad hanging - load, not the motor. A probe (probe_oneleg.py) then showed open-loop PD cannot hold single support on physics grounds, so the criterion became "an equilibrium exists and a hand-written 4-gain COM feedback can hold it": single-support roll stiffness is hip and ankle in series and must exceed m*g*h_com = 22.5 N*m/rad; ankle kp 12 in series with hip kp 80 gives only 10.4 (open loop 16/16 fell), ankle 60 with hip 80 gives 34.3 (52% margin).

    Change

    A per-joint gain profile (rl_oneleg: hip_roll kp 80, ankle_roll kp 60, the rest as rl_default) - which needed per-joint gain support in robot.yaml, the bridge, deploy and the trainer's actuator groups - decided before training. Thermal account: single support makes hip_roll the dominant heat load (about 7.8 N*m against a 7 N*m continuous rating), so acceptance and demos run in segments of at most 60 s with a temperature check.

    Outcome

    Under rl_oneleg the hand-written feedback held six cells cleanly for 6 s (hip_roll steady torque 2.1-3.4 N*m, half the thermal budget); under rl_default the same feedback on the same cells fell 0/4. The trained V0 policy then passed its 40-cell acceptance.

    Mechanism

    With PD position control, the steady error needed to carry a static load is load/kp; when that error exceeds the joint's range the posture is unreachable whatever the policy does, and series compliance between joints lowers the effective stiffness below the gravity stiffness that single support demands.

    Applies when

    • single-support, crouched or one-arm-load postures on PD actuators
    • a joint "droops" under load on hardware but tracks well when hanging
    • deciding whether a new skill needs its own gain profile
    “但 kp=20 时撑住 7.8 N·m 需要 **0.39 rad 跟踪误差 > 整个内收行程**。真机已实测: 命令 +0.17 实际 −0.04(droop 0.21 rad),悬挂时 0.0008 rad——是负载不是电机。 … 单支撑滚转是 hip/ankle **串联**刚度,必须 > m·g·h_com = 22.5 N·m/rad;ankle kp12 串 hip80 只有 10.4(开环 16/16 全摔),60 串 80 = 34.3(裕 52%) … **rl_default 同反馈同格 0/4 全摔**(增益档必要性对照)”
    git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §1-1 单脚站: 几何可行,卡点是 hip_roll 增益权限 / §2 A 线增益 / §5 probe 定谳
  • Push DR helped one lineage and hurt another at the same dose - robustness budget is conserved and gets borrowed, not createdpush-dr-conditional-budget-conservation
    Replicatedomnidr-tuningdomain-randomizationcurriculumattribution

    Before opening a disturbance-DR rung, measure whether the untrained policy already meets the spec; if training it anyway, expect the benefit to be conditional on the lineage's existing DR load, grade the intensity, and audit retained margins - budget spent elsewhere will be borrowed back.

    Symptom

    The push rung's outcome flipped with the lineage: direct +/-0.6 m/s push failed outright on first attempt (base walking collapsed - kd1.2 scan 0/3 from iter 3300, sim2sim self-falls with pushes OFF - no PASS point existed); staged +/-0.3 then gave the narrow-kd single-working-point lineage real gains (push survival 1/5 -> 4/5) while the SAME dose made the dual-working-point balanced-band lineage WORSE (20-seed survival 18 -> 12/20 plus across-the-board push regression).

    Context

    The four-ladder verdict ("四梯定案", s2e/s2f at both intensities) named the pattern: "push DR 收益条件性" - the benefit is conditional on how much robustness budget the lineage has already spent. The law candidate: "DR 总预算守恒, 平衡带鲁棒性从抗扰余量借" - total DR budget is conserved; a lineage already covering a wide plant band pays for push tolerance out of its disturbance margin. Both S2 ladders therefore closed at the friction rung, with the decisive numerator: untrained push tolerance already met the 4-6 N*s requirement, so the rung was not needed at all ("⑥ push 不训(收益条件性,免训 ±0.6 已达 标)"). The same accounting later justified the C-before-S2 ordering ("push/μ 两轮已实证 DR 预算有限且会被重分配") and trimmed the second S2 pass to three rungs.

    Change

    Push removed from the standing ladder; graded intensity retained as the method IF a lineage ever needs push training; "does the untrained policy already meet the disturbance spec" instituted as the first check before opening any disturbance rung.

    Outcome

    Two rungs (push, ground mu) deleted from the second S2 pass on measured grounds; the ladder's real yield was re-stated honestly as precision, not robustness (speed gate 0 -> 20/20, zero-command drift 0.98 -> 0.06 m, but push 159 -> 125/160).

    Mechanism

    A fixed-capacity policy allocates representation and margin across the training distribution; adding a disturbance axis to a lineage that already spans a wide plant family forces reallocation - the new tolerance is bought with existing margins. Lineages with narrow plant coverage have free budget, so the identical DR dose lands as gain. Benefit is a property of (dose x lineage state), never of the dose alone.

    Applies when

    • proposing push/perturbation training on a hardened lineage
    • the same DR rung helped one lineage and hurt another
    • accounting where a ladder's robustness gains actually came from
    “push DR 收益条件性 —— s2e⑥a (单工作点血统 kd 窄带) ±0.3 得抗推 1/5→4/5; s2f⑥ (双工作点平衡带血统) 同档反而 20-seed 存活 18→12/20 且抗推全面倒退。规律候选: DR 总预算守恒, 平衡带鲁棒性从抗扰余量借。两阶梯均以 ⑤ 摩擦级收官 … 抗推 4~6 N·s 免训已达标。”
    train/OMNI_V0_SPEC.md § 4. ⑥ push 四梯定案 (2026-08-07)
  • Compare the achieved reward to the computed ignore-floor to tell "never learned" from "learned but unprofitable"ignore-floor-diagnosis
    Mechanism understoodomniattributionattributionreward-shaping

    For any skill that trains flat, compute the reward the null policy would earn on that term; achieved==floor means the behavior never paid out (find why: exploration, reward observability, or feasibility) - do not tune weights first.

    Symptom

    C4 sidewalk failed on both arms; the question was whether the policy had found sidewalk and rejected it as unprofitable, or never found it at all - two diagnoses with opposite fixes.

    Context

    The theoretical value of the sidewalk tracking term for a policy that completely ignores the command was computable from the command distribution: 0.189. Trained final values landed at 0.187 (arm A) and 0.204 (arm B) - sitting exactly on the ignore-floor - while the honest balance at the optimum actually favored sidewalking (net +0.88/step inside the side bucket). Later the same arithmetic closed the whole saga: for the true reward landscape, doing real sidewalk scored 0.153 vs 0.944 for ignoring - the policy's refusal "是理性最优,不是探索失败" (rational optimum, not exploration failure) under one hypothesis, and under the final measurement-corrected account the policy had "每一次都在 做理性选择" (made the rational choice every time).

    Change

    Diagnostic rule adopted: compute the ignore-floor for the new term; if the achieved value sits on it, the behavior was never expressed in useful volume (or the reward cannot distinguish it - check both); if the achieved value is above floor but the behavior is absent at deployment, the policy sampled it and priced it out - then the reward balance, not exploration, is the lever.

    Outcome

    Correctly identified that PPO had not merely under-valued sidewalk; each subsequent hypothesis (waveform sign, regularization cage, exploration form, reward kernel) was tested against this floor arithmetic, which kept the search honest through three reversals.

    Mechanism

    Every reward term has a computable value under the null behavior; the achieved-vs-floor gap is a one-number audit of whether the optimizer ever monetized the target behavior. It converts "training failed" into one of two mechanistically distinct states with different fixes.

    Applies when

    • a new skill's tracking reward plateaus early
    • deciding between exploration fixes and reward-weight fixes
    • post-mortem of a failed curriculum rung
    “track_lin_vel_y_exp 训练终值恰好坐在「完全无视指令」的底分上(A 0.187 / B 0.204,理论值 0.189),而终点 balance 明明有利(side 桶内净 +0.88/步)。不是学会了不划算,是根本没学到。”
    train/C_LADDER_RUN.md § 3e. 为什么首战 FAIL
  • Drop the frozen policy into chosen configurations - a squat 2.7 cm lower than the stuck pose stood 52% of the time, the stuck W-sit 0%, and the interpolation between them showed a wall, not a slopeconfiguration-probe-wall-not-slope
    Mechanism understoodrecoveryattributionattributionmeasurementcurriculum

    When a policy is stuck, probe the frozen policy from a grid of hand-placed start configurations, including interpolations between the stuck state and a nearby state it escapes from; one read-only experiment separates height, torque, sampling and configuration and tells you whether to prevent entry or train the exit.

    Symptom

    After R0.3 the policy stood from 62% of starts and never from the W-sit it fell into; height, torque, missing samples and reward were all plausible suspects.

    Context

    A read-only probe placed the R0.3 policy directly into specified configurations. Squats (hip, knee, ankle) = (-.65,-1.3,-.65) stood 100%, (-1.0,-2.0,-1.0) 89.8%, (-1.2,-2.4,-1.2) at 0.176 m 52.3%; the measured W-sit at 0.203 m 0.0%; the account-(3) hand-over state (146 deg tilt) 34.4%; linear interpolations from the W-sit toward the squat at 25/50/75% stood 0.0/0.0/3.1%. The squat family's quasi-static torque is 16% of the limits, and the W-sit was visited ~9 s per episode in training. FK showed the squat family (-a,-2a,-a) keeps the torso vertical, the feet flat and the COM over the feet all the way from 0.146 m to 0.384 m.

    Change

    Height, torque and sampling were eliminated in one experiment; the next rungs targeted entering the W-sit (foot placement) instead of escaping it, and seeding the dead point itself was ruled out because it was already visited every episode.

    Outcome

    Pure configuration: the W-sit (hips externally rotated +/-47 deg, knees folded 110 deg, shins flat, feet beside the body) is a different place from the sagittal squat (feet flat under the COM). The policy's standing skill was bound to a narrow sagittal family, and the wall was confirmed by the interpolation. The foot-placement rungs that followed took prone from 0/159 to 158/159.

    Mechanism

    A learned skill covers the neighbourhood of the states it succeeded from; a start state outside that neighbourhood fails regardless of height or torque, and an interpolation that stays at zero until close to a working state shows the boundary is sharp.

    Applies when

    • a policy stalls in a specific posture and several causes are plausible
    • deciding between reverse-curriculum seeding and entry-prevention shaping
    • a feasibility account says a path exists but the policy does not take it
    “**决定性对比:比死点矮 2.7 cm 的蹲姿站立 52.3%,死点 0.0%。** 所以不是高度、 不是力矩(蹲姿族准静态力矩膝 1.96/12、踝 1.24/17,只占 16%)、也不是训练采样 (死点每局被访问 ~9 s)。**是纯位形问题** … 插值实验进一步显示这**不是坡是墙** —— 走到 75% 仍只有 3.1%”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §15 死点位形实验(只读探针,同一个 R0.3 策略放进指定位形)
  • Fix a too-deep nominal pose before adding any penalties - the default stance defines the basin training starts innominal-posture-before-penalties
    Mechanism understoodwalkreward-shapingreward-shapingplant-calibration

    Before tuning penalties on a degenerate gait, audit the nominal pose and height targets against morphology and published ratios; if the default stance encodes the degenerate behavior, fix it first - and recompute dependent quantities (init height) by FK, not by hand.

    Symptom

    Policy lived in a crouched shuffle; nominal knee angle was 0.5 rad (28.6 deg) - deeper than published configs (Unitree G1 0.3 rad / 17.2 deg, Booster T1 0.4 rad) - so the policy's starting point and its action-space center both sat inside the crouch basin.

    Context

    Initially ranked "secondary" in the local diagnosis, this was promoted to co-first priority by the cross-check against published reward tables, which states that with nominal knee flexion above ~0.4 rad, fixing the posture must precede adding any penalties ("改这个之前别加任何 惩罚都是白费"). Companion base-height items: walk profile had weakened base_height_l2 to -5.0 (base class -10, field standard -10 to -20, "the second most common cause of death"), and the height target must be the STANDING height (0.384), not the crouch height.

    Change

    Nominal knee 0.5 -> 0.3 rad with init_base_height recomputed by MuJoCo FK (0.3739 -> 0.3802); base_height_l2 restored to -10 with standing height target; both bundled as first-priority alongside the clearance term.

    Outcome

    Part of the v5/v6 package that lifted swing height to 34 mm and tracking to 87%; the crouch basin stopped being the default answer.

    Mechanism

    The nominal pose is the fixed point every regularizer pulls toward and the point where action=0 lands; if that point is itself the degenerate posture, every penalty fights the geometry. Correcting the attractor is prior to shaping the gradient field around it.

    Applies when

    • policy converges to a crouched or collapsed posture
    • nominal joint angles were chosen for stability rather than gait
    • base-height reward targets or weights were locally weakened
    “研究明确说"nominal 膝屈超过 ~0.4 rad 必须先改,改这个之前别加任何惩罚"。我们是 0.50,超标。… base_height_l2 在 walk profile 里被减到 −5.0(基类是 −10)。研究说这是"第二常见死因"且应 −10 ~ −20。改回 −10。目标高度用站立高 0.384 是对的(研究要求 target 必须是*站立*高度而非蹲姿)。”
    train/WALK_DIAGNOSIS.md § 修正 ①(升级优先级) / 修正 ④
  • When training fails repeatedly, inject the target behavior open-loop - stop tuning rewards for an unverified behavioropen-loop-probe-before-reward-tuning
    Mechanism understoodomniattributionattributionprocesscurriculum

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop on the real plant/sim first, and only resume training once you hold a measured, safe, sign-verified target trajectory.

    Symptom

    Three sidewalk training rounds failed; hypotheses multiplied (exploration failure / wrong reference waveform / insufficient authority) with no way to pick between them by running more training.

    Context

    Instead of a fourth reward guess, the team wrote probe_side_ref.py: the candidate reference is injected open-loop on top of a frozen policy's output (bypassing PPO entirely), directly measuring "what happens if the robot literally does this waveform" - separating all three hypotheses in one experiment (5 seeds x 8 s per condition, several waveform families and gains). The probe immediately eliminated the authority hypothesis (full-amplitude execution, 5/5 survival) and localized the problem to the waveform/measurement side. The closing principle was written down after the saga: without a verified target behavior, tuning rewards is "在黑暗里试钥匙" (trying keys in the dark).

    Change

    Standing method: before opening another training rung on a failing skill, build an open-loop (or task-space) generator of the intended behavior, measure whether the physical system can express it and what it looks like - then train toward a verified, quantified target.

    Outcome

    The probe chain produced the verified waveform (reversed-sign triangle, half gain), quantified safe amplitude (tilt 8.2 deg at band top, foot distance clear of the wall), exposed the metric bug when probe and training disagreed, and ultimately supplied the feed-forward that made C4 pass in +100 iters.

    Mechanism

    Training couples exploration, reward design, and feasibility into one opaque outcome; open-loop injection cuts the loop and tests feasibility and waveform alone. A behavior demonstrated open-loop converts the remaining failure into a pure credit-assignment/reward question - and its measured trajectory becomes the reference itself.

    Applies when

    • repeated training failures on one skill with multiple live hypotheses
    • uncertainty whether the platform can physically express the behavior
    • a reference trajectory's shape/sign/amplitude is guessed, not measured
    “三轮 FAIL 之后不再猜,写 train/probe_side_ref.py 把参考开环注入到策略输出之上(绕过 PPO),直接量「照这个波形做会怎样」,一次分开三个假说:甲 探索 / 乙 波形 / 丙 权限。”
    train/C_LADDER_RUN.md § 3f. C4 真因定谳(开环探针) / 3k. 建议的下一步
  • Mirror augmentation over an asymmetric default injects systematic error - symmetrize the default first and verify the transform bit-exactmirror-augmentation-needs-symmetric-default
    Mechanism understoodwalkobservation-designobservation-honestycurriculumprocess

    Before enabling any symmetry augmentation, make every constant inside the observation encoding exactly symmetric, and validate the mirror transform against forward kinematics to machine precision - an unverified augmentation is a new error source, not a regularizer.

    Symptom

    Mirror data augmentation was about to be added while both default poses (standing_pose, walk nominal_pose) were asymmetric - stale hand-tuned compensations from before a ground re-calibration, with hip_yaw differing 2.40 deg between sides and the foot soles actually tilted (pitch 2.88/1.35 deg, roll -2.47/+0.25 deg).

    Context

    The observation encodes joint_pos_rel = q - default. Under mirroring q_l -> -q_r, the relation (q-default)_l -> -(q-default)_r holds only if default_l = -default_r; with an asymmetric default, augmentation produces observation pairs that are NOT mirror images, i.e. "default 不对称时做镜像增强会引入系统性错误,比不做还糟" (worse than not doing it). The fix: adopt model geometric zero as standing default (MuJoCo FK verified: sole pitch/roll exactly 0, asymmetry 0.00 deg) and a symmetric crouch for walk (hip -0.25/knee -0.5/ankle -0.25 satisfying hip - knee + ankle = 0 to keep soles flat). The mirror transform itself was verified bit-exact before use: pseudovector vs polar-vector sign patterns (ang vel [-1,1,-1], gravity [1,-1,1], cmd [1,-1,-1]), joint swap-and-negate; FK check that left-foot pose under q equals the mirror of right-foot pose under mirror(q), measured error 0.00e+00.

    Change

    Defaults symmetrized first (with init heights recomputed by FK), stand policy retrained on the new default so both policies share one default; augmentation enabled only after the FK mirror test passed.

    Outcome

    stand_v1 achieved exact left/right pairing (l_knee -0.1013 / r_knee +0.1013), six-pair asymmetry 0.0 deg, height fluctuation 7 -> 1 mm, 33% less mean |action|.

    Mechanism

    Augmentation asserts an equivariance of the observation encoding; any asymmetric constant inside the encoding (the default) breaks the asserted symmetry, so the augmented data teaches a false invariance. Verifying the transform against FK geometry tests the assertion end to end, independent of the training stack.

    Applies when

    • adding mirror/symmetry augmentation to locomotion training
    • defaults or trims were hand-tuned per side at any point
    • observations are expressed relative to a default pose
    “观测里 joint_pos_rel = q − default。镜像下 q_l → −q_r,要让 (q−default)_l → −(q−default)_r 成立,必须 default_l = −default_r。default 不对称时做镜像增强会引入系统性错误,比不做还糟。… 位置误差与姿态矩阵误差实测均为 0.00e+00。”
    train/RETRAIN_v2.md § 2. 前提:default 姿态必须先对称化(不是可选项) / 3. 镜像变换
  • Knee swing collapsed because it directly trades against the slip penalty - price the conflict explicitly and clamp what reward cannot holdknee-swing-vs-slip-pricing
    Mechanism understoodwalkreward-shapingreward-shapingattribution

    When a behavior collapses as another metric improves, look for the term pair trading them and set their price ratio deliberately (with escalation and reverse tripwires pre-registered); where the policy actively spends action budget to undo your target, stop paying more reward and clamp the target space structurally.

    Symptom

    Knee peak-to-peak swing collapsed across generations - v5 33 deg, v10 26-30, v10b 7-8, v11 6.5-8.6 - and rolling the clock back did not recover it, acquitting the clock; the collapse tracked the gated slip penalty instead: "屈膝与不打滑在当前奖励里直接对抗" - v10b's excellent 93 deg slip was purchased with knee amplitude.

    Context

    Reward-side flexion fixes had failed three times: raising reference amplitude backfired twice (v9/v11), and v11's deep-squat default was actively fought by the policy - it spent 0.68 of action budget pulling the squat straight ("被策略花 0.68 动作拉直反杀"). v12's design accepted the conflict as real and attacked on two tracks: (1) ECONOMICS - a direct knee_swing_amplitude reward (+0.3, target 0.55 rad, capped at 0.6/step = 55% of tracking), explicitly opposed to the slip penalty by design ("显式对立——这正是设计:v12 就是这场对抗的定价实验"), with an escalation ladder (K +0.3 -> +0.5, then slip -0.5 -> -0.3, one layer at a time) and a reverse tripwire (slip telemetry back at v10 levels -> slip weight to -0.8, accept ~20 deg knee compromise); (2) STRUCTURE - knee target bounds [0.2, 0.9] rad so full straightening is physically impossible (straightest 11.5 deg) and the 0.68 fighting budget is released. A bonus falsifiable prediction was attached: phase-lock strength tracks amplitude (v9_probe 48 deg locked 2.5 Hz; v11 low-amplitude 1.36 Hz unlocked), so if K works, hardware phase-lock should return - one change, two verdicts.

    Change

    knee_swing_amplitude reward + knee target clamp + pre-registered escalation/reverse levers; the failed reward-side-only approach retired.

    Outcome

    The lineage was frozen before v12 trained (strategic reset), but the diagnosis stands as the walk line's clearest example of two reward terms trading a behavior between them, with the pricing experiment and structural clamp fully designed and calibrated.

    Mechanism

    When two terms price opposite aspects of one motion (swing amplitude creates yaw momentum that becomes slip), the optimizer settles wherever the price ratio puts it - patching one side moves the equilibrium, not the conflict. Explicit pricing makes the trade a designed quantity; structural clamps remove the regions where the policy spends budget fighting the designer.

    Conflicts

    The pricing experiment (K vs slip) was designed and calibrated but never trained - the 2026-08-05 reset suspended v12; the collapse attribution table and the 0.68-action counterattack are measured, the remedy's效果 is untested.

    Applies when

    • one gait quality degrades in lockstep with another's improvement
    • the policy visibly fights a default pose or reference
    • repeated reward-side fixes for the same behavior have failed
    “膝摆塌在 v10→v10b,头号嫌疑是门控滑移罚(四代实测膝 p2p:v5 33° / v10 26~30° / v10b 7~8° / v11 6.5~8.6°;退时钟没救回 → 非时钟)——"屈膝"与"不打滑"在当前奖励里直接对抗 … 奖励侧修屈膝已三败 … v11 深蹲 default 被策略花 0.68 动作拉直反杀”
    train/WALK_V12_SPEC.md § 0. 定位 / 2. K —— 膝摆经济(与滑移罚的对偶)
  • Add a single-point-suspension test to acceptance - the ground is a free stabilizer that hides divergencesuspension-probe-removes-free-stabilizer
    Mechanism understoodwalkgate-batterygate-batteryreal-acceptancesim2sim

    Include at least one acceptance condition that strips the environment's free stabilization (suspension, or equivalent) - the sim-passing policy that fails on hardware is often failing a condition the battery never posed.

    Symptom

    walk_v5 looked healthy in every on-ground sim test yet diverged on the real robot - the acceptance battery had never measured a condition that would have revealed it.

    Context

    The battery gained a single-point-suspension probe (robot hung, feet free): measure torso tilt while the policy runs without ground contact. v5 scored 45.9 deg mean tilt suspended - wildly unstable - which the file calls "最灵敏的失稳探针(拿掉地面这个免费稳定器)": ground reaction forces passively stabilize a marginal policy, so on-ground metrics saturate long before the policy's internal balance is actually sound. v6 halved it (23.0 deg, target <10 deg) - progress visible on a scale where on-ground numbers showed nothing.

    Change

    Suspended-tilt added as a standing acceptance row; run under the honest contact parameters battery (accept_v2 with measured condim 4 / torsional friction 0.035), under which v5 correctly FAILS in agreement with the real robot.

    Outcome

    The sim battery's verdict on v5 flipped from pass to fail, matching hardware; suspended tilt became the discriminating metric between v5 and v6 (45.9 vs 23.0 deg) when ground metrics differed little.

    Mechanism

    Contact with the ground closes a stabilizing feedback loop the policy gets for free; removing it exposes the policy's own attitude control authority. A metric measured only in the assisted condition cannot rank policies by the unassisted quantity that hardware will actually demand during perturbations and flight phases.

    Applies when

    • sim acceptance passes but hardware diverges
    • designing an acceptance battery for a legged robot
    • two candidates tie on ground metrics
    “单点吊那条是最灵敏的失稳探针(拿掉地面这个"免费稳定器"), v5 在地上一切正常却在真机发散, 就是因为验收从没测过这个工况。”
    train/WALK_V6_MINIMAL.md § 5. 验收
  • The get-up kept getting faster because standing earlier paid more every step - lowering torque authority barely slowed it, and only zeroing the standing income for the first 3 s moved the pace into the design bandper-step-income-drives-speed-time-gate
    Mechanism understoodrecoveryreward-shapingreward-shapingcurriculum

    When a skill is too fast, find the term that pays for finishing early and gate that income by time; keep the "get into position" term ungated so the policy does not learn to wait, use a ramp instead of a cliff, and confirm with a paired same-level experiment that the drift is motivational before changing it.

    Symptom

    The user judged the get-up too fast (Isaac medians about 0.7-1.6 s) and suspected path dependence: the policy seemed to get faster the longer it trained.

    Context

    Lowering the beta authority 0.40 -> 0.30 cut impact but moved supine only 1.70 -> 2.00 s: coordination-limited, not torque-limited. A paired experiment inside one beta level (checkpoint 15,600 vs 18,499, +2,900 iterations, same ruler) measured the drift: get-up medians -7 to -10%. The spec concluded the motive, not the path, was the cause - per-step standing income pays for every early step, and any lineage (even one from scratch) races toward the fastest solution inside its constraints.

    Change

    V2.5: the standing income (base_height, stand_pose, still, feet_on_ground) multiplied by w(t) = clamp(t/3 s, 0, 1); upright deliberately NOT gated, so righting and sitting up early still pay and the policy is not taught to lie flat and wait; a ramp, not a step. V2.5b: zero before t0 = 3 s, then a 1 s ramp.

    Outcome

    V2.5: Isaac 100%, get-up +17-43% slower, MuJoCo 100/100/98/100% (the best cross-simulator reading yet), still short of the 3.5-4.5 s design band - a linear ramp only discounts early income. V2.5b: MuJoCo supine 2.04 -> 4.10 s and prone 3.18 -> 4.04 s, inside the band; the Isaac pace barely moved (a lineage habit on a gradient-free plateau). Later the zero gate proved harmful when trained from scratch (curriculum-history-is-part-of-the-product) and in the V3.1 lineage (time-gate-vs-wide-stance-retire-the-fix).

    Mechanism

    Constraints on authority or velocity change how the fastest solution looks; the time structure of the task income decides how fast the fastest solution is.

    Applies when

    • a policy is faster or more aggressive than wanted and constraints do not slow it
    • progress-style rewards pay every step spent at the goal
    • performance drifts faster with more training at fixed settings
    “**V2.5 机制(唯一)**:站立收入(base_height/stand_pose/still/feet_on_ground) 乘时间斜坡 w(t)=clamp(t/T_gate,0,1),T_gate=3.0 s;**upright 刻意不门控** (翻正/坐直早期照常拿钱,防"躺平等门开" … **用户假设量化 证实:逐步计酬动机在 β 包络内持续压缩时间,约束挡不住动机 —— V2.5 动机层 修法为正解。**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §39 V2.5 预注册 / §39 补 配对实验
  • A constant-value plant rung passed every binary gate with record scores - and shipped 60% thinner posture margins that hardware exposedconstant-value-dr-overfits-margin
    Mechanism understoodomnidr-tuningdomain-randomizationreal-acceptancegate-battery

    Randomize deployment-critical axes over a narrow band spanning the measured real support - never a single value, never a fictitious tail - and report graded margin quantities (tilt margin) next to binary gates, because saturated gates rank thin-margin and thick-margin policies identically.

    Symptom

    s2_lag1 (trained at constant 1-frame latency) posted the strongest sim gate sheet in history (20/20 everywhere) yet was unstable on hardware, while s1e (trained across the full 0-3 frame band) was the every-run-stable SOTA at the same power.

    Context

    The sim autopsy (new --delay-jitter harness modeling the BusWorker's time-varying phase drift): 18 runs across constant and time-varying delays ALL survived - time variation alone does not kill - but the tilt-margin ordering reproduced hardware exactly: s1e 7.7-9.1 deg (thickest) < s2_lag1 10.7-15.0 < s1c 16.2-18.2. Attribution: constant-value training permits precise specialization to that one value; s1e's band diversity forced cross-value robustness - "恒定 1 帧训练 vs s1e 的 0~3 帧全带——分布多样性逼出跨值鲁棒,恒定值允许精确 特化" - so the constant-rung policy's margins were ~60% thinner, fine in sim's clean world, pushed over the line by real-world disturbances. Tool lesson booked: "存活门二值饱和后掩盖裕度差" - binary survival gates saturate and hide margin differences; graded margin columns (tilt-max) belong in the report. The synthesis with the opposite failure (wide tails cause drag-glide): the proposed resolution was a NARROW uniform band (0.02, 0.04) covering exactly the real 1-2 ticks - diversity inside the measured support, no tail, no single point. s1e's root selection later leaned on the same property: its full-band latency training "预装" the delay rungs and delivered "全工况稳定裕度" that survived power derating.

    Change

    DR-on-an-axis design refined to a three-way distinction: no wide fictitious tails (drag), no single constant values (thin margins), but a narrow band spanning the measured real support; acceptance reports gained graded margin columns alongside binary gates.

    Outcome

    The tilt-margin column entered the standard report; the s1e root (band-trained) carried the C ladder while the constant-value branch was archived with its three contributions credited.

    Mechanism

    Robustness margins are shaped by the diversity of the training distribution, not just its support: a point-mass distribution lets the optimizer trade margin for on-point performance, while a band forces solutions that keep margin across the band - and binary survival metrics cannot see the difference until the margin is spent on hardware.

    Conflicts

    The narrow-band (0.02,0.04) resolution was a pending recommendation ("裁决建议(待用户)") at the time of writing; the lineage instead moved root to s1e whose full-band training predated the staged ladder - the deterministic-staging card and this card record the two failure modes the final design must avoid simultaneously.

    Applies when

    • a rung trained at a fixed plant value aces sim but wobbles on hardware
    • binary acceptance gates are all saturated across candidates
    • choosing between constant, banded, and wide DR on one axis
    “18 跑全活,时变性单独不足以击杀;但 tilt_max 裕度排序完整复现真机:s1e 7.7~9.1°(最厚)< s2_lag1 10.7~15.0 … 恒定 1 帧训练 vs s1e 的 0~3 帧全带——分布多样性逼出跨值鲁棒,恒定值允许精确特化 … 存活门二值饱和后掩盖裕度差(s2_lag1 sim 门 20/20 史上最强却真机不稳)”
    train/README.md § s2_lag1 真机不稳 × s1e 稳的 sim 对拍(2026-08-07,时变延迟实验)
  • Reward fixes come in causal chains - foot height, then landing impact, then foot spacingreward-chain-foot-height-landing-spacing
    Observed oncewalkreward-shapingreward-shaping

    Plan reward shaping as a chain, not a point fix: when you patch a degenerate gait behavior, pre-register which adjacent behavior the optimizer will exploit next and watch for it.

    Symptom

    Three problems appeared strictly in sequence: (1) swing feet lifted too low; (2) after fixing that, feet slammed down - "实际比视频里暴力得多" (far more violent in person than on video); (3) after fixing that, feet drifted too close together and collided.

    Context

    Each reward fix removed one degenerate optimum and exposed the next. The fix for foot spacing (COM lateral randomization +/-5 cm to force leg spread) itself caused base side-to-side sway, requiring a further foot-to-centerline distance penalty. Lucen had just solved its own foot-height problem (19mm -> 40mm swing height) and logged landing impact and foot spacing as the predicted next two problems.

    Change

    Chain of additions - (1) penalty when swing foot below 5 cm; (2) landing vertical-velocity penalty at touchdown; (3) COM lateral randomization +/-5 cm, then foot-centerline distance penalty to cancel the induced sway.

    Outcome

    Reference robot progressed through each stage; each individual fix worked and predictably surfaced the successor problem. For Lucen the chain served as a pre-registered roadmap of what breaks next.

    Mechanism

    Locomotion rewards are coupled through contact dynamics: raising swing height adds potential energy that must go somewhere at touchdown (impact); penalizing impact and forcing robustness to COM shifts changes lateral support strategy (spacing/sway). The optimizer always exploits the cheapest unpenalized channel, so fixing one channel routes the exploit to its neighbor.

    Applies when

    • adding a foot-height / clearance reward
    • feet slam or landing impact grows after a clearance fix
    • feet converge toward the centerline or self-collide
    • any single-reward fix to a coupled gait behavior
    “抬脚太低 → 加惩罚:摆动足低于 5 cm 就扣分 / 加完之后砸脚 → 抬起来了但落地极猛,"实际比视频里暴力得多" → 加落地速度惩罚 … / 两脚太近甚至互撞 → 先试质心横向随机化 ±5 cm … 有效但引发新问题——基座开始左右摇摆 → 再加足-中心线距离惩罚 … 这三条是串联的:每个修复都会暴露下一个问题。”
    Experience.md § 三个问题的解法链 (lines 72-77)
  • Design the next run to complete the 2x2 - either outcome then convicts or acquits a factor cleanlyfill-the-missing-factorial-cell
    Mechanism understoodwalkprocessprocessattributionreward-shaping

    When two config factors are jointly suspected, lay out the factorial of existing evidence, spend one run on the missing cell with both interpretations and an early-abort tripwire written in advance - and treat either outcome as a verdict, not a disappointment.

    Symptom

    Hip joints froze at the action clamp in v7, but the history could not say whether the culprit was the raised action_rate (-0.2) or the halved reference amplitude (scale 0.15): existing versions covered only three corners of the (rate x scale) space - v5 (-0.03, 0.30) healthy, v6 (-0.03, 0.15) healthy, v7 (-0.2, 0.15) frozen.

    Context

    v9 was designed explicitly as the missing cell (-0.2, 0.30), with the readings pre-registered: v9 not frozen -> the real anti-freeze force was always the reference amplitude and -0.2 may stay; v9 frozen -> -0.2 is convicted beyond appeal (freezes at both amplitudes) and the next version goes straight to a structural fix. "两个结局都是干净的信息" - both endings are clean information.

    Change

    One training run allocated purely to complete the factorial, with freeze tripwires (joint_pos_ref telemetry <0.1 at iter 1000-1500 -> abort, do not run to 6000) so a conviction costs the minimum compute.

    Outcome

    v9 froze - the rate weight was convicted at both amplitudes ("−0.2 铁案定罪"), and v10 moved to the structural saturation fix with the weight question closed instead of re-litigated.

    Mechanism

    Three corners of a 2x2 leave the two factors confounded in the failure corner; the fourth observation makes each factor's marginal effect identifiable. Pre-registering both readings turns the run into a guaranteed-informative experiment regardless of outcome.

    Applies when

    • two config changes are confounded in a failure
    • version history already covers some corners of a factor grid
    • deciding what single experiment buys the most attribution
    “这恰好补齐一个 2×2 实验矩阵的缺格 … v9 不冻 → 真正的抗冻结主力一直是参考摆幅,−0.2 可以留;v9 仍冻 → −0.2 铁案定罪(两种摆幅下都冻),v10 直接上结构修复 … 两个结局都是干净的信息。”
    train/WALK_V9_SPEC.md § 0. 设计原则 (2×2 实验矩阵)
  • Ungated phase shaping made standing 42x more expensive than stepping - and the stepping was cooking the hip motorsmoving-gate-42x-stand-tax
    Mechanism understoodomnireward-shapingreward-shapinghardwarereal-acceptanceprocess

    Gate every phase/clock-driven shaping term on the command that justifies motion, price the cmd=0 case explicitly during design - and when a reward flaw is sim-only, book it with trigger conditions instead of operating immediately on a working lineage.

    Symptom

    At cmd = 0 the sim policy never stood still - it stepped in place and crept 0.98 m per 20 s; on the real robot the same lineage's stepping made hip_roll motors run 20 degC hotter than every other joint (43-48 degC vs 25-28).

    Context

    The arithmetic closed it: the three phase-shaping terms (joint_pos_ref 1.6, feet_contact_number 1.2, feet_clearance_swing 1.6) are driven by the gait clock with NO command gating, so standing at cmd=0 forfeits 3.32/step of shaping while honest standing earns only 0.078 of tracking - stepping wins 42x. The heat chain: perpetual stepping = perpetual single support = one hip_roll stalled at ~4.3 N*m (25% of torque limit) carrying the torso's frontal-plane moment - two hip_rolls = 90% of whole-machine steady-state I2R; measured stand-vs-step comparison: total heat -71% when actually standing. The suspended test acquitted the actuator (0.21 rad sag -> 0.0008 rad in air) and mechanics vetoed the easy fix ("降 kp 救不了热" - equilibrium torque equals the external load regardless of kp). The fix (moving_gate: hard-gate the three shaping terms on |cmd| > eps) was designed - then DEFERRED by the user because the real robot at the time stood fine: "真机不表现该问题, 为真机不存在的病改奖励表不划算", with written trigger conditions (C-ladder stand row persistently failing, or real robot starting to step/drift at cmd=0) and the known hazard tag (this is exactly the reward-change class that triggers the B-arm signature). When the real robot later DID step and cook, the booked trigger fired and moving_gate moved from debt to to-do with its benefit re-priced: "停止空烧 hip_roll,稳态发热降 ~71%".

    Change

    moving_gate designed with the gate_by_cmd convention; deferral, triggers, and expected heat recovery all pre-registered instead of patching the reward for a then-sim-only symptom.

    Outcome

    The cmd=0 stepping went from mystery to closed arithmetic; the thermal measurement (hip_roll +20 degC) quantitatively confirmed the 90%-of-heat prediction; the reward change waited for real-world justification instead of spending a risky revision early.

    Mechanism

    Clock-driven shaping terms define a perpetual-motion bounty unless gated by command; the resulting idle gait is not a training bug but the table's optimum. Its cost surfaces on hardware as stall-torque heating set by statics (mass x lateral offset), which no gain change can remove - only removing the motion (gating) or widening the stance can.

    Applies when

    • the policy steps in place or creeps at zero command
    • specific joints run hot in idle behaviors
    • deciding when a known reward flaw justifies a risky mid-lineage fix
    “塑形合计 −3.32/步 … 净: 站定亏 42 倍 … 两颗 hip_roll 4.29 / 4.09 N·m(各占限扭 25%),占全机稳态 I²R 的 90% … 吊挂实测 hip_roll 跟踪误差 0.21 rad → 0.0008 rad … 降 kp 救不了热 … 真机不表现该问题, 为真机不存在的病改奖励表不划算”
    train/README.md § C1 FAIL 节 (cmd=0 的 sim/真机分歧记账) / C2 真机 A/B 三b 发热定性
  • When the eval proxy has a known systematic bias, gate on within-proxy differences, not absolutesrelative-metrics-survive-proxy-bias
    Mechanism understoodwalkgate-batterygate-batterysim2simattribution

    Where the evaluator is known-biased, design gates as within-evaluator contrasts (left vs right, A vs B, pre vs post) that cancel common-mode error; reserve absolute thresholds for quantities whose proxy calibration has been checked.

    Symptom

    MuJoCo systematically overestimated yaw turn gain (1.58/2.45 where Isaac measured 1.04/1.02), and friction alignment recovered only part of the gap (mu 1.0 -> 0.6 pulled it to 1.43/2.27) - so any absolute turn-gain acceptance threshold in MuJoCo would be judging the proxy's bias, not the policy.

    Context

    The acceptance criterion was rewritten to use only the left/right difference of the turn gain (standard: <20%): both directions pass through the same biased proxy, so the bias largely cancels in the difference while the policy's chirality - the thing being gated - survives. The absolute-gain row was dropped: "转向只判左右差,不判绝对值". A parallel task was still opened to align MuJoCo contact parameters to Isaac's (mu, restitution), since drift proved highly friction-sensitive (straight-line drift -65 deg at mu 1.0 vs +1 deg at mu 0.4) - bias reduction and bias-robust metrics proceeded together.

    Change

    Gate metric changed from absolute turn gain to left-right gain difference; proxy-alignment work scheduled separately rather than blocking acceptance.

    Outcome

    Turn acceptance became meaningful across proxy versions (v5 70% -> v6 19% difference measured the real improvement) while the absolute bias question was pursued without holding the ladder hostage.

    Mechanism

    A systematic multiplicative or additive proxy bias applies to both arms of a mirrored measurement; differencing (or ratioing) mirrored conditions cancels the common-mode bias to first order, leaving the asymmetry signal. Metrics built this way remain valid while the proxy is imperfect - which it always is somewhere.

    Applies when

    • a sim proxy disagrees with the trainer or hardware in absolute terms
    • writing acceptance thresholds for direction-paired skills
    • proxy calibration work would otherwise block a ladder
    “转向只判左右差,不判绝对值 —— MuJoCo 的偏航增益系统性高估(实测 1.58/2.45 vs Isaac 1.04/1.02),摩擦只能解释一部分(μ 1.0→0.6 仅拉到 1.43/2.27)。差值是相对量。”
    train/WALK_V5_SPEC.md § 6. 验收
  • Real-robot trials of a new skill were staged by risk - a hanging dry run with the robot posed by hand, then one short try per category on a mat with the hardest last, then the composed behaviour (switch + walking) last - with the user present and a log every timestaged-hang-mat-floor-for-get-up
    Replicatedrecoveryreal-deployreal-acceptanceprocess

    Stage a new skill's hardware trials by risk - hanging dry run posed by hand, one short try on a mat per start category with the hardest last, the composed behaviour last - with an operator ready to cut enable and a log for every try; relax a safety ban only for short, attended runs and say so in writing.

    Symptom

    A get-up policy acts violently near the ground by design, and the first unstaged real run of the line was stopped as dangerous.

    Context

    The hanging checklist written with the first stamped recovery product (v2_5, 2026-08-11), to be ticked item by item with the user present: both machines on the same commit and firmware torque limits checked; the robot hung from a single point about 0.1 m off the ground; a dry run with the robot posed by hand into supine and prone to watch that the target stream is gentle (the beta contract keeps targets within +/-0.25 rad of the measured pose, so enabling causes no homing fling); the first floor try is supine only, on a mat, once, with torque and joint logs; categories are added one at a time, prone last; any kicking or oscillation cuts enable immediately. For the switch (08-14) the runbook orders: hang with the standing policy as the locomotion side, then on a mat push the robot over and let it recover, and only last swap in the walking policy. An exemption was also written: edge-standing policies stay banned from long or unattended runs, but a short single A/B with the user present, hung or on a mat, is allowed. The one-leg line reused the same order (hang, then floor with a spotter, 60 s segments with a temperature check).

    Change

    Real trials as a checklist of stages, each gated on the previous one, with the composed behaviour last.

    Outcome

    The line's first real get-up (v2_6, 08-11) came through this protocol and was reported "fairly stable"; no further hardware outcomes of the switch are recorded in the spec.

    Mechanism

    Each stage exposes one new risk (commanded targets without contact, a single category with contact, harder categories, then the interaction of two policies), so a failure is attributable and cheap.

    Applies when

    • first hardware trial of a recovery, jumping or other high-impact skill
    • switching between two policies on hardware for the first time
    • a policy with a known posture defect needs a comparison run
    “吊挂空跑: 手动摆到 supine/prone 姿态, 看目标流是否温和 (β 帽 7.5 N·m, 目标永远贴着当前 q ±0.25 rad —— 使能瞬间无归位甩动, 这是 β 契约附带保证) … 落地首试: supine 一类, 垫子, 单次; τ/q --log 全程记录 … 逐类别扩展 (prone 最后), 每类先单次”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §40 吊挂执行单(真机首试;需用户在场,逐项打勾)
  • The advisor's "runtime five-stage state machine + per-stage reference poses + RL residual" appeared in none of the three papers it cited - reading the originals changed the plan and downgraded two widely repeated industry claimsadvisor-paraphrase-vs-paper
    Replicatedrecoveryprocessprocessattribution

    Read the primary source behind any piece of advice before adopting its architecture; record where the paraphrase and the original differ, downgrade claims the originals do not support to speculation, and adopt what the verified sources actually share.

    Symptom

    After the violent first real run, an advisor proposed re-architecting recovery as a runtime staged state machine with reference poses and an RL residual, citing HoST, HumanUP and StableMimic.

    Context

    The three papers were read in full on 2026-08-09 and tabulated (deployment form, what the "stages" really are, hard constraints, references). HoST: one end-to-end policy, height-gated rewards in training, action anchored as q + beta*a with a beta curriculum. HumanUP: two training stages with the same observation/action; the vendor's three-stage state machine is the baseline it beats (41.7% vs 78.3%); Stage II tracks an 8x slowed Stage I trajectory (4x too violent, 10x does not converge). StableMimic: a learned soft gate. On 08-10 more sources were checked the same way: the Agility page does not say Digit's self-righting was learned in simulation (only step recovery is stated as RL), so the claim was downgraded to speculation; HoST's support for the 12-DoF armless Mini Pi exists in its code repository, not in the paper text.

    Change

    Adopted only what the originals share: a hard action bound or anchored action space, strong smoothing including a second-difference term, a slowed hidden reference, heavy DR and real fallen states. The staged runtime state machine was not adopted.

    Outcome

    The next re-rooting candidates came straight from the verified material, and the beta-anchored action space (HoST, with the Mini Pi configuration as the nearest real-robot precedent) became V2, the lineage that later stood up on hardware.

    Mechanism

    A paraphrase compresses a paper into the advisor's own architecture; only the original shows what was actually deployed, what was a baseline, and which numbers came with which ablation.

    Applies when

    • an advisor, agent or summary proposes an architecture with citations
    • an industry claim ("X learned it in sim") is about to justify a design
    • several papers are cited for one combined recipe
    “⇒ **顾问的核心形态"runtime 五阶段状态机 + 每阶段参考姿态 + RL residual"在三篇引文 里均不存在**,其中 HumanUP 还点名 state machine 是局限。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 三篇引文精读判决(2026-08-09 全文核对;顾问转述与原文有出入)
  • The fallen-state reset was designed, not sampled from SO(3) - fixed category shares with jitter, a low drop that settles physically, equal left/right shares for mirror augmentation, and a numeric check before trainingfallen-pose-reset-distribution
    Observed oncerecoverytraining-runcurriculumdomain-randomizationprocess

    Build a fallen-start distribution from named, physically plausible categories with jitter and a settle phase, keep mirrored categories at equal probability, and check the realized shares and penetration numerically before spending a training run on it.

    Symptom

    A get-up policy can only learn from the fallen states its resets produce; uniformly random orientations produce ground-penetrating and limit-jammed states the robot can never be in.

    Context

    R0 reset_root_fallen: supine 30%, prone 30%, side_l 15%, side_r 15%, mid (random axis 50-125 deg) 10%, +/-15 deg jitter, full yaw, dropped from 0.28-0.40 m and left to settle under physics, joints uniform inside the soft limits with a 5% margin plus small random velocities. Random SO(3) was rejected (the advisor agreed). side_l and side_r must have equal probability because mirror augmentation turns a left fall into a right fall. The advisor had proposed supine and prone only for R0; the spec included side and mid because the feasibility accounts showed physical solutions for all of them, and wrote "narrow back to supine+prone" down as the first fallback. With no display on the training box the reset was checked numerically instead of by eye.

    Change

    Category mix as above; realized shares, settle height and penetration measured over 512 envs before the first run. A fallen-state bank (real falls, settled and stored) was pre-registered for R2.

    Outcome

    Realized shares 29.3/31.6/16.4/17.8% against the config, settle +0.262 m, final penetration 0/512 (a 0.10 m peak at the write instant, ankle links only, pushed out within 80 ms because the 0.28 m drop floor is shorter than a fully extended leg). R0's failure was a reward basin, not a reset artifact. The fallen-state bank stayed unbuilt through V3.1 (checklist item open); R0.3 later re-sliced the prone share into roll_l/roll_r bands, which is what forced the acceptance distribution to be frozen separately.

    Mechanism

    A category-structured, physically settled start distribution keeps training on states the robot can actually occupy, and equal mirrored shares keep mirror augmentation a pure doubling of data rather than a bias.

    Applies when

    • designing reset distributions for get-up, recovery or multi-contact skills
    • mirror/symmetry augmentation is on and the task has chiral start states
    • no viewport is available to inspect resets on the training machine
    “角度 jitter ±15°、yaw 全域、0.28~0.40 m 低空放下由物理沉降,关节软限位内 均匀(留 5% 余量)+ 小随机速度。**不用 random SO(3)**(会采出穿地/极限卡死 等现实不可能状态,顾问同判) … side_l/side_r **概率必须相等**(镜像增强的样本同分布前提)”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §3 R0 任务定义 / §9 核查单
  • Every rung gets written stop criteria - hit any one, stop; tightened when priors say results should come fastpreregistered-stop-criteria-per-rung
    Replicatedomnitraining-runprocessgate-battery

    Freeze per-rung stop criteria (old-skill floors, new-skill deadline, oscillation signature, known pathology signatures) before training, stop on first hit - and shorten the deadline in proportion to how fast your mechanism says results should appear.

    Symptom

    Continued training past the point of degradation had previously destroyed capabilities (C1 trained away the root's backward skill); without hard stop rules, sunk-cost reasoning keeps runs alive too long.

    Context

    Standard rung protocol - eval_c_matrix every 100 iters at 5 seeds with a fixed criteria list, frozen before training ("开训前写死,事后不许改"): (1) stand or vx+0.15 at <=3/5 for two consecutive points; (2) vx+0.15 tracking <70% for two consecutive points (vs recorded root baseline); (3) oscillation signature - survival repeatedly crossing zero between adjacent checkpoints, stop on first occurrence, do not wait for confirmation; (4) new-skill-no-progress deadline (e.g. wz tracking <30% or still same-signed after iter 400); (5) a known lineage pathology signature (s1c arm-B: scatter -> half-recover -> collapse).

    Change

    Stop budgets scale with prior knowledge: when C4-redo4's feed-forward was already proven open-loop, the no-progress deadline was tightened from +500 to +100 iters ("前馈是开环就给 71~96% 的 … 若 +100 还没有 … 不值得再烧").

    Outcome

    Multiple rungs were stopped exactly on criterion (C4-redo criterion 4 at iter 1200; redo2/redo3 at absolute 900), converting each into a clean hypothesis test instead of a drifting run; no rung burned its full cap on a dead hypothesis.

    Mechanism

    Degradation of retained skills is a lineage-level injury that later training often cannot undo, so detection latency is capital loss; and a stop rule written before the run cannot be bent by hope. Tightening deadlines when priors predict immediate results converts "no result yet" into evidence against the mechanism.

    Applies when

    • starting any resumed/curriculum training rung
    • deciding whether to keep training a run that shows early regression
    • a mechanism-backed change should produce results immediately
    “每 100 iter eval_c_matrix.py --seeds 5,命中任一即停 … 存活在相邻 checkpoint 间反复跨越 0(出现即停)… 收紧版:绝对 iter 900(= +200)时 vy±0.10 跟踪仍 <30% 即停”
    train/C_LADDER_RUN.md § 3h. 预注册停梯判据(开训前写死,事后不许改)
  • The run policy never left the ground and fell in the second simulator from the frontal plane - its DR (gains and latency only) covered the actuator axis, not the frontal-plane contact and inertia disturbances the doubled stride amplified; "is DR on" is the wrong questionthin-dr-judged-by-channel-coverage
    Replicatedrundr-tuningdomain-randomizationsim2simattribution

    Judge a DR recipe by whether its randomized terms cover the channel where the skill can lose stability, not by whether DR is enabled; when a new skill lengthens single support or enlarges motion in one plane, add disturbances in the plane it destabilizes before training.

    Symptom

    run R1 (6,000 iterations, 78 min): no flight phase ever appeared, and every one of 13 checkpoints failed the eight-gate MuJoCo smoke. In Isaac: zero terminations in 6,000 iterations, 4.2 deg tilt. In MuJoCo at delay 2: 1/6 survived, falls within 1.9-6.2 s at 50.8-58.7 deg, the most saturated joints all roll joints.

    Context

    The run contract doubled sagittal travel (knee action scale 0.9, knee swing peak 1.14 rad) with a 0.60 s period and 0.40 duty - long single support - while roll/yaw scales were deliberately left at 0.5. DR copied the s1e recipe: kp/kd (0.9, 1.1) and latency on; mass, COM, joint friction and push all off; ground friction pinned at (1.0, 1.0). Flight was read two independent ways: Isaac's per-foot contact reward stayed 0.845-0.857, never above 0.87 - the arithmetic ceiling of a gait with zero flight - and 30 of 36 MuJoCo seeds had flight fraction exactly 0 (the nonzero six were all tumbling falls). Foot lift itself worked (46-59 mm against a 50 mm design point): the walk-era "not enough travel" failure did not recur.

    Change

    Verdict FAIL, with the pre-registered first knob (exploration noise 1.0 -> 1.2) explicitly rejected as aimed at a different axis. The lesson was generalized and applied at the next line's design review: the one-leg spec made push, body mass, base COM and friction DR mandatory for its permanent single support and banned the thin recipe.

    Outcome

    The run line did not continue past R1 in the sources. The one-leg V0 with the wider DR passed its friction-variant gate (mu 0.4 and 1.2) inside a 40/40 acceptance.

    Mechanism

    Randomizing gains and latency covers the actuator's axis; a skill whose failure lives in frontal-plane contact and inertia needs randomization on that channel (push, mass, COM, friction), or the trainer's exact plant becomes the only one the policy can stand on - the omni_s1 transfer trap a second time, this time with DR switched on.

    Applies when

    • a policy is flawless in the trainer and falls immediately in a second simulator
    • reusing a DR recipe from a skill with a different support pattern
    • failures concentrate on one axis (roll, yaw) the DR does not touch
    “**机理**: 矢状面行程翻倍 (膝摆动峰 1.14 rad) + T 0.60 + duty 0.40 的长单支撑, 把额状面扰动放大了一个量级; 而 roll/yaw 通道按 §3 **刻意没有放大** (仍 0.5), DR 又是 s1e 复刻的薄配方 (mass/COM/关节摩擦/push **四关全关**, 地面摩擦钉死 (1.0, 1.0))。 … 说明**薄 DR 的判据不能只看"有没有开 DR"**, 要看**开的那几项 是否覆盖失稳所在的通道** —— kp/kd 与延迟是执行器轴向的, 对额状面接触/惯性 扰动零覆盖。 … 0.87 正是「零腾空的走路步态」的天花板算术”
    git:Lucen V2@origin/run-line:train/README.md § run R1 FAIL (2026-08-09, run 21-30-30_run_r1): 腾空零, 但病根在额状面不在探索
  • Fall recovery was defined as the whole chain - any fallen pose, a stable stand, a clean hand-back to walking - and built as a second policy behind a deploy-side switch, not folded into the walking PPOrecovery-two-policies-and-a-state-machine
    Observed oncerecoveryprocessprocesscontract-freezereal-acceptance

    Define a recovery skill by the whole chain it must complete, including the hand-back to the next controller; if it is built as a separate policy, make the switching logic and its handoff contract a deliverable of their own, and keep the recovery observation contract a subset of the locomotion one so a unified policy stays possible later.

    Symptom

    A walking robot that falls needs a human to stand it back up. The design question on 2026-08-09 was whether to teach getting up inside the existing omni walking policy or beside it.

    Context

    The user set the goal as "any fallen pose -> stand up alone -> stand stably", and the spec named the real difficulty as the full chain fall -> recovery -> stable stand -> correctly initialised walking history and clock -> walking, making the deploy state machine a first-class deliverable. A unified single policy had a real-robot precedent (arXiv:2605.18611, a state-dependent gate near 37 deg tilt) but was deferred until a recovery policy and an omni policy were each reliable. The line ran on its own branch and worktree with every walk/stand/omni/run config path untouched. The development path copied the G1 learned get-up logic (arXiv:2502.12152): first find any feasible get-up (ugly accepted), then add smoothing, torque and real-robot constraints. The recovery contract kept the base 45-dim observation (command slice held at 0, no gait phase, no frame history - their reasons do not apply to a skill without a clock or a velocity task), so it stays a prefix of the 215-dim omni contract and a later merge is not foreclosed.

    Change

    Two policies and a deploy-side switch instead of one retrained walking policy; recovery got its own minimal contract (45 dims, full-range action, later the beta-anchored profile) and its own acceptance battery.

    Outcome

    The split held for the whole line: on 08-14 deploy_policy gained a second (PolicyIO, ONNX) pair behind --recovery-policy, each loaded under its own manifest contract, and the runbook runs stand_v1b or omni_c4_ff800 as the locomotion side with recovery_v3_1p1c. The literature scan of 08-10 found that every verified get-up implementation deploys one end-to-end policy (or softly gated experts) and stages only on the training side - so the runtime state machine here is the walk/recovery switch, not a staged get-up.

    Mechanism

    A separate policy keeps each reward table single-purpose and lets a proven walking lineage stay byte-frozen; the cost moves to the handoff, where every piece of state one policy leaves behind (history, clock, last action, command) must be reset for the other.

    Applies when

    • adding fall recovery or get-up to a robot that already walks
    • choosing between one unified policy and a switched pair of policies
    • designing the observation/action contract of a secondary skill
    “先做 recovery policy + omni policy 两个策略,部署侧状态机切换;不把 recovery 硬塞进现有 omni PPO。 … 任务定义:**任意跌倒姿态 → 自己站起来 → 稳定站立**。真正的难点不只是"起身", … omni walk**(§6 部署状态机是本 spec 的一等公民,不是附录)”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §0 目标口径与架构判决(用户 2026-08-09 定)
  • A joint-velocity penalty meant to slow the get-up cut joint speed 16% and left the get-up time unchanged - the knob never moved the variable, so the idea it was meant to test stayed untesteddof-vel-penalty-is-not-a-pacing-knob
    Mechanism understoodrecoveryreward-shapingreward-shapingattribution

    Before reading a result as a test of an idea, check that the knob actually moved the independent variable; velocity regularizers smooth a schedule they do not set, and a schedule driven by per-step task income moves only when that income's time structure does.

    Symptom

    The get-up took 0.6-0.9 s in Isaac with large torque demand; the user proposed getting up more slowly so less torque would be needed.

    Context

    The idea had support in the accounts: the acceptance bound is an upper bound of 5 s (5-8x margin), the quasi-static squat path peaks at 25% of the limits, and rolling over needs no momentum. R3.2 raised dof_vel from -1e-3 to -5e-3 as the single variable.

    Change

    dof_vel -1e-3 -> -5e-3 (child-run from R3.1).

    Outcome

    Get-up medians moved +0.02-0.04 s (noise); raw joint velocity -16%; torque demand median got worse (hip_pitch 46-48% -> 63-67%) as the new term competed with torque_headroom on the same joints; MuJoCo 98 -> 96%. Verdict FAIL on the knob, not on the idea, and the rung was not adopted. When pace was later attacked through the income's time structure (V2.5/V2.5b), the MuJoCo get-up moved into the 3.5-4.5 s design band.

    Mechanism

    The pace was set by base_height_progress paying for every step spent high (stand earlier, earn more); a velocity regularizer only smooths motion along the same schedule and does not change when the robot stands up.

    Applies when

    • trying to make a skill slower or gentler with smoothness penalties
    • an experiment's primary metric did not move and a verdict is being written
    • two penalties act on the same joints
    “**关键判读:`dof_vel` 罚只把关节速度压了 16%,而起身用时一点没变。** 也就是说**这一级根本没有把"慢下来"这个自变量推动起来** —— 所以它**不构成对 用户假说的检验** … 起身节奏由 `base_height_progress` 的逐步计酬决定(早站起来就多 拿),速度正则只在同一条时间轨迹上把动作抹匀,不改变何时站起来。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §25 R3.2(dof_vel −1e-3→−5e-3,慢一点起身)
  • Privileged signals (true velocity, foot force, foot height) go to the critic onlyobservation-honesty-critic-only
    Mechanism understoodwalkobservation-designobservation-honesty

    Treat the actor observation vector as a hardware contract: every element must exist on the real robot with realistic noise; privileged simulator truths belong in the critic only.

    Symptom

    Policies trained on ground-truth base linear velocity work in sim and fail on hardware, where only a drifting IMU and encoders exist - the policy has learned to depend on a signal that does not survive deployment.

    Context

    Many open-source locomotion stacks feed simulator ground-truth linear velocity to the actor. The reference team refused: the real robot has no ground-truth velocity. Asymmetric actor-critic keeps the training benefit of privileged information without deploying the dependency.

    Change

    Route ground-truth velocity, foot contact forces, and foot heights to the critic only; the actor observes exclusively signals that exist on hardware (IMU-derived quantities, encoders, commands, previous actions).

    Outcome

    Recorded as adopted doctrine in Lucen's experience log; the trained actor's input contract matches what the real robot can actually produce.

    Mechanism

    The critic is discarded at deployment, so it may consume any privileged state to reduce value-estimation variance; the actor's observation set is a deployment contract - anything in it that hardware cannot supply (or supplies with different noise/drift) becomes a train/deploy distribution shift the policy was never trained to handle.

    Applies when

    • designing actor/critic observation spaces
    • reviewing a config where the actor sees base_lin_vel or contact forces
    • sim policy is strong but real robot drifts, oscillates, or falls without obvious actuator cause
    “很多开源代码库把真值线速度喂给策略,Asimov 团队没有,因为真机上没有真值速度,只有会漂的 IMU 和编码器;用完美速度训练出来的策略会依赖它,然后在硬件上失效。真值速度、足底力、足高统统只给 critic”
    Experience.md § 观测空间的诚实性 (line 7)
  • The 1.5 Hz step-frequency gate was retracted - the author had misread his own actuator data, and the limit fought pendulum dynamicsgate-threshold-retracted-frequency
    Mechanism understoodwalkgate-batterygate-batteryactuator-modelingprocess

    Every gate threshold must cite its measurement and survive a first-principles sanity check; when a gate keeps failing otherwise healthy behavior, re-derive the threshold from the raw data before enforcing it again - and retract wrong gates in writing.

    Symptom

    An acceptance criterion "step frequency <= 1.5 Hz" kept failing healthy policies (walk_v4 at 2.33 Hz), and an earlier attempt to force slower stepping (v3) had killed stepping altogether.

    Context

    The threshold had been derived from the author's own actuator frequency-response measurements - but re-reading the raw table showed the misread: amplitude ratio at 2.0 Hz is 0.88 (knee) / 0.83 (ankle), acceptable; the genuinely bad point was walk_v2's 3.8 Hz at 0.58. Mechanism check agreed: the leg as a compound pendulum (L ~ 0.30 m) has a natural frequency ~1.1 Hz, swing half-period 0.45 s - the observed 2.1-2.4 Hz sits near where the leg wants to swing, and forcing 1.5 Hz "是跟摆动动力学对着干" (fights the swing dynamics). The frequency definition itself was pinned by two independent methods (contact-event counting 2.39 Hz vs FFT 2.33 Hz, agreeing): reported numbers are cycle frequency = steps per leg per second.

    Change

    Gate retracted in writing: "步频 ≤1.5 Hz 应删除或放宽到 ≤2.5 Hz"; frequency definition standardized before entering any config.

    Outcome

    walk_v4/v6's 2.1-2.4 Hz reclassified from disease to normal; the v3 failure got its probable explanation (suppressing stepping to meet a wrong gate).

    Mechanism

    A gate is only as good as the measurement and the reading behind it; thresholds inherited from a misread plot become invisible design constraints that later training obeys at real cost. Cross-checking a threshold against first-principles dynamics (pendulum frequency) is a cheap way to catch such misreads.

    Applies when

    • an acceptance threshold repeatedly fails policies that look healthy
    • thresholds were set from a single person's reading of raw data
    • a forced compliance with a gate degrades the behavior it guards
    “我当初依据自己测的执行器频响定的,但看错了区间。… 2.0~2.4 Hz 的幅值比 0.83~0.88 是可接受的;真正不行的是 walk_v2 的 3.8 Hz。… 腿按复摆算(L≈0.30 m)自然频率约 1.1 Hz … 把它压到 1.5 Hz 是跟摆动动力学对着干(walk_v3 把迈步压没了,可能正是这个原因)。”
    train/WALK_DIAGNOSIS.md § ② 撤回"步频 ≤1.5 Hz"这条验收标准 —— 是我定错了
  • After seven patch-generations, freeze the lineage as a regression baseline, fix the structural debts, and retrain from zerofreeze-lineage-fix-structure-restart
    Observed oncewalkprocessprocesscontract-freezecurriculum

    When successive rungs keep trading one symptom for another, ask whether the remaining problems are structural (contracts, latency, sampling, reward-table architecture); if so, freeze the lineage as regression baselines, pay the structural debts, and restart minimal - carrying forward laws and instruments, not weights and weights' patches.

    Symptom

    The v5-v11 walk lineage had accumulated interacting patches (reward terms, gates, clamps, per-joint scales) faster than it converged on the user's goal; v12's spec itself was superseded before training by an external review's verdict that the remaining problems were structural, not parametric.

    Context

    The 2026-08-05 status banner records the pivot: the walk profile was rolled back wholesale to v10b parameters, the v5-v11 lineage frozen "只作回归对照" (kept only as regression baselines), and four structural debts were named as prerequisites for a from-zero straight-walk baseline: the action-latency FIFO (fixed with its own test), the ONNX manifest contract, discrete command sampling, and a minimal reward table. The v12 spec - fully designed, partially implemented - was suspended: "本规格挂起,不再按此开训".

    Change

    Strategy switched from "one more patch generation" to freeze-fix-restart: lineage checkpoints retained as comparison anchors, infrastructure hardened first, then a clean retrain with a minimal reward table (this restart produced the s* generation that later became the real-robot SOTA line).

    Outcome

    A designed-and-ready training generation was deliberately not run - the review's structural findings outranked sunk design cost; the restart line inherited seven generations of laws (calibrations, gate batteries, falsified fixes) without inheriting their entangled reward table.

    Mechanism

    Patch lineages accumulate coupled terms whose interactions eventually cost more to reason about than a restart costs to train; the knowledge worth keeping is the laws and instruments (measured plant values, calibrated gates, falsified directions), not the entangled weights. A restart on hardened structure converts the lineage's lessons into a clean initial design instead of another delta.

    Applies when

    • repeated rungs shuffle symptoms without net progress
    • an external review flags infrastructure/contract debts
    • deciding between another patch generation and a clean retrain
    “同日外部评审定调换路线:冻结 v5~v11 血统(只作回归对照),修结构性问题(latency FIFO 已修 tests/test_action_latency.py、ONNX manifest 契约、离散命令采样、最小奖励表)后从零训直行基线。本规格挂起,不再按此开训。”
    train/WALK_V12_SPEC.md § ⚠️ 状态 (2026-08-05)
  • A frame-history observation under zero DR memorizes the trainer's plant fingerprint - the estimator must see variation to learn estimationhistory-obs-needs-plant-variation
    Mechanism understoodomniobservation-designobservation-honestydomain-randomizationsim2sim

    If the observation carries history (stacked frames, RNN), keep at least minimal plant variation (gain/latency jitter) on from the first iteration - "nominal first, robust later" is structurally invalid for estimator-bearing contracts.

    Symptom

    omni_s1 (fresh 215-dim contract with a 5-frame history window, trained with DR fully off): training all green, yet the MuJoCo gate scored 0/20 on all eight doors - falls within 2 s, seven checkpoints, not one transferred.

    Context

    The history window exists precisely to let the actor implicitly estimate line velocity and actuator dynamics (the actor is denied base_lin_vel by observation honesty). Under a constant plant that implicit estimator has nothing to estimate - it learns the trainer's exact response fingerprint instead, and any other simulator's micro-differences are out-of-distribution: "5 帧窗按设计就是隐式估计器, plant 恒定时它学到 Isaac 精确响应的指纹". The planned "nominal-first-robust-later" staging was declared STRUCTURALLY incompatible with history observations: "估计器要见过变化才学估计, 否则学背诵" (an estimator must see variation to learn estimation, otherwise it learns recitation). Honest confound note kept: this is mixed with "zero DR does not transfer, period" - but both attributions prescribe the same fix, so no control was run.

    Change

    S1.1: minimum actuator jitter turned on from day one - kp/kd +/-10%, latency 0-1 frame (friction/COM/mass still nominal, no push - those stay for the S2 ladder).

    Outcome

    Transfer restored: survival 0/20 -> 20/20, speed 19/20, foot distance 20/20 (remaining failures moved to gait quality, a different disease); the staging doctrine was amended - history-carrying contracts never train under a frozen plant.

    Mechanism

    A recurrent/history channel fits whatever temporal structure minimizes loss; with a deterministic plant the cheapest structure is the plant's own impulse-response signature, yielding features that are simulator-specific rather than physics-general. Plant variation forces the channel to carry state-estimation features that transfer.

    Conflicts

    Attribution is explicitly confounded with the simpler "zero DR never transfers" reading ("与「零 DR 本身就不迁移」混杂 … 两种归因处方相同, 不做对照") - the source chose not to spend a control run separating them.

    Applies when

    • adding frame stacking or recurrence to an actor observation
    • a nominal-plant policy fails a cross-simulator gate within seconds
    • planning DR staging for a new contract
    “frame_hist × 零 DR = plant 指纹过拟合——5 帧窗按设计就是隐式估计器, plant 恒定时它学到 Isaac 精确响应的指纹, MuJoCo 的微小差异即 OOD, 2 s 内摔, 七个 checkpoint 无一迁移。「先标称后鲁棒」的分段与历史观测结构性冲突:估计器要见过变化才学估计,否则学背诵。”
    train/OMNI_V0_SPEC.md § 3. S1.1 修订记录 ①
  • A reward on a quantity the actor cannot observe teaches "produce less of it", never "correct it" - closed-loop correction needs an outer loopreward-observability-limit
    Mechanism understoodomniobservation-designobservation-honestyreward-shapingcontract-freeze

    Before adding a reward, check the actor can observe (or infer) the quantity: unobservable-error rewards buy only average suppression - route correction tasks to an outer loop whose commands stay in distribution, and do not break a frozen contract to add an observation a deploy-side loop can supply.

    Symptom

    Heading kept drifting despite world-frame yaw rewards, and a reviewer proposed heading-error rewards - raising the question of what yaw shaping can even teach this actor.

    Context

    The adopted architectural verdict: the actor's 45-dim base observation cannot see accumulated heading at all - projected_gravity is invariant to rotation about the gravity axis, and omega_z is a rate, not an angle. World-frame yaw-rate rewards are therefore privileged shaping that can only teach "少产生旋转" (generate less rotation), never "偏了以后拉回原线" (pull back to the line after drifting) - the policy cannot represent the error it would need to correct. The S1 gate (<=5 deg / 10 s) demands exactly the former, so the stack is right for its gate; active heading correction is assigned to the deployment outer loop (--heading P-loop converting heading error into in-distribution wz commands) plus small-wz training - and the 215-dim contract is explicitly NOT extended with a heading observation ("契约不加 heading 观测,冻结不动"). The reviewer's companion bias hypothesis was adjudicated with data: drift is bimodal - a basin mechanism decides whether you leave (seeds vary +/-16-46 deg vs -385 to -391 deg), and once out, rotation direction is constant (weight chirality; candidate root: the phase clock always swings left first).

    Change

    Yaw shaping kept as rate-tracking (three-layer stack); heading correction owned by the deploy outer loop; contract frozen; the "which behaviors need an outer loop" question settled by observability analysis rather than reward tuning.

    Outcome

    Stopped a contract change and a futile reward direction; drift work split correctly into rate-suppression (trainable) and error correction (outer loop), consistent with the earlier measured 10x drift reduction from the deploy-side loop.

    Mechanism

    A policy can only condition on its observation sigma-algebra; rewards on functions outside it shift the marginal action distribution (open-loop average effects) but cannot create feedback on the unobserved variable. Whether to add an observation, an outer loop, or accept average-shaping is decided by the task's gate: suppression gates need shaping, correction gates need the variable in some loop's view.

    Applies when

    • adding rewards on accumulated/世界-frame quantities (heading, position)
    • deciding between a new observation, an outer loop, and shaping
    • a drift symptom persists across reward-weight changes
    “actor 的 45 维基座观测不到累计航向(projected_gravity 对绕重力轴旋转不变,ωz 是速率不是角度)——世界系 yaw 奖励是特权塑形,只能教「少产生旋转」,不能教「偏了以后拉回原线」。… 主动纠偏闭环 = S3 把小 wz 进分布 + deploy --heading 外环 … 215 契约不加 heading 观测,冻结不动。”
    train/OMNI_V0_SPEC.md § 3. 评审④判决(2026-08-06,S1.3 开训前)
  • Prove a new penalty actually fires - two ways a clearance term silently did nothinginert-reward-term-audit
    Mechanism understoodwalkreward-shapingreward-shapingplant-calibration

    Before training with a new reward term, log its realized per-step value under the current policy and confirm it is nonzero where intended - check coordinate zero-points against FK and check who occupies the term's gate; and never weaken the term that creates the states your new term needs.

    Symptom

    A newly designed swing-height clearance penalty could have trained as a no-op twice over, and the companion advice to lower feet_air_time actively backfired when tried.

    Context

    Instance 1 (zero-point offset): the proposed code used body_pos_w of the foot link, but that is the ankle_roll_link frame origin, which sits 0.0585 m above the ground even with the foot flat on it - so (0.03 - 0.0585) is always negative and the penalty is永远 0; the 0.0585 offset must be subtracted (verified identical in MuJoCo FK and Isaac). Instance 2 (gate occupancy): the clearance penalty fires only in swing phase; a dragging policy keeps both feet in contact, so the penalty is constantly 0 for exactly the policy it was meant to fix - and worse, any slight lift immediately incurs it, a reverse threshold. Lowering feet_air_time to 0.5 on that advice measurably collapsed air time to 0.0002 (below v3). Corrected understanding: "clearance 是把已有的摆动相抬高, 造出摆动相仍要靠 air_time" - air_time creates the swing phase, clearance raises it.

    Change

    Fixed the height zero-point; kept feet_air_time as the swing-phase creator with clearance layered on top; both errors documented as corrections to the team's own earlier advice.

    Outcome

    With both fixed, swing height rose from 22-23 mm (v2) to 29 mm (v5) to 34 mm (v6); the inert-term failure class entered the standing checklist.

    Mechanism

    A penalty's gradient exists only where its gate is occupied and its argument crosses its threshold; frame offsets shift the threshold out of reach, and phase gates can have zero occupancy under exactly the policy being treated. Terms interact as an ecology - one term must create the states in which another can act.

    Applies when

    • adding any gated or thresholded penalty (clearance, impact, slip)
    • a new term produces no behavioral change at any weight
    • body-frame positions are used in reward code
    “body_pos_w 是 ankle_roll_link 坐标系原点,平放触地时仍高出地面 0.0585 m。… (0.03 − 0.0585) 恒为负 → 惩罚永远是 0 … clearance 惩罚只在摆动相生效,拖地时两脚始终触地 → 惩罚恒 0;而一旦轻微抬脚就立刻扣分,对正在拖地的策略是反向门槛。… 正确认识:clearance 是"把已有的摆动相抬高",造出摆动相仍要靠 air_time。”
    train/WALK_DIAGNOSIS.md § walk_v4 独立验收 — 本文档给的两处代码/建议是错的
  • A pull-assist curriculum keyed to a global success share was satisfied by the categories that already worked and withdrew before prone learned anything - conditioning the criterion on prone took it from 2.5% to 98.7%curriculum-criterion-conditioned-on-lagging-category
    Mechanism understoodrecoverytraining-runcurriculumgate-battery

    Measure a curriculum's advancement criterion on the population the scaffold is meant to help; a global success share is met by whatever already works, and the help is withdrawn before the lagging case learns.

    Symptom

    Under the beta-anchored action, supine and side stood reliably while prone still sat (1.3%). A pull-assist curriculum added to help it was withdrawn completely within ~790 iterations and prone moved only to 2.5% (noise).

    Context

    The pull assist follows HoST: an upward force on the base, active only when the torso is within 30 deg of vertical, scaled by body weight (HoST's 200 N on G1 = 0.583 BW -> 56 N here, steps of 5.6 N, ten levels to zero); the product must pass with no assist. It had already taught sit-to-stand in V1. In V2.1 its advancement criterion was the standing-time share over all envs (threshold raised to 0.55 because the share was already ~0.53).

    Change

    V2.2: PullAssistForce with gate_category = "prone" - only envs whose first step classifies them as prone count toward the criterion - and the threshold back at 0.35. A feasibility signal was pre-registered: if prone's share stayed near zero under the full 56 N, return to the roll-over path instead of adding force.

    Outcome

    The prone-conditioned curriculum withdrew level by level only as prone itself passed: prone 98.7%, and the four-category gate passed for the first time on the line (98.6% overall, re-falls 0%, torque gate PASS).

    Mechanism

    A pooled success share is filled by the categories that already succeed (supine/side ~53%), so the scaffold is removed on their account before the lagging category has used it.

    Applies when

    • an assist, guide force or easier setting is withdrawn by a success threshold
    • one task category lags while the pooled metric looks healthy
    • a curriculum ran to completion without changing the lagging category
    “**教训:全局站立占比阈会被存量类别(supine/side ~53%)凑够,拉力在 prone 学会前就撤光了 —— metric 设计失误,不是拉力机制失效**(它在 V1 教会过 坐→站)。 … **V2.2(已启动)**:`PullAssistForce` 加 `gate_category="prone"` —— 达标判据 只统计 prone 类 env”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §34 V2.1 判决(2026-08-10)
  • Slowing the gait clock at deployment is out-of-distribution and backfires - lower the commanded speed instead, or train the knobcycle-time-override-is-ood
    Mechanism understoodwalkreal-deployreal-acceptanceattributioncontract-freeze

    Any deployment override must correspond to a dimension the policy was trained to handle; to make a parameter field-adjustable, randomize it in training and observe it - otherwise use the levers inside the trained envelope (commands) and leave the knob alone.

    Symptom

    Real-robot feedback "walks very fast and unstable" suggested slowing the gait; a deploy-side --cycle-time override existed, making "just slow the clock" a one-flag temptation.

    Context

    A sim sweep of the override on walk_v6 @cmd 0.3 showed monotone degradation away from the trained 0.40 s cycle: at 0.50 s tilt jumped 7.9 -> 13.2 deg and landing force 1.52x -> 2.24x; at 0.80 s (half speed) clearance collapsed to 3 mm - dragging again - with 20 deg tilt. Meanwhile the legitimate lever, lowering the commanded speed with the clock untouched, improved everything monotonically: cmd 0.1 gave 104% tracking, 6.7 deg tilt, minimum slip - the most stable operating point. The file distinguishes the two "slows" explicitly: lower command = smaller steps at the same 2.5 Hz rhythm; a slower rhythm itself requires retraining - randomize cycle_time (e.g. 0.40-0.65 s) during training and expose it as an observation, and only then does --cycle-time become a field-adjustable knob.

    Change

    Deployment guidance: never ship a cycle-time override the policy was not trained under; respond to "too fast/unstable" with lower commands; schedule clock variability as a training-time (contract-level) change if a field knob is wanted.

    Outcome

    The sweep quantified the trap before hardware paid for it (dragging and 2.2x landing force at slowed clocks); cmd 0.1 documented as the stable demo point.

    Mechanism

    The policy is a function fitted around the training distribution; a deploy-side override moves an input (phase rate) to values never seen, so behavior degrades unpredictably - the knob LOOKS like a capability because it exists in the code, but capability lives in the training distribution, not the interface.

    Applies when

    • a deploy tool exposes overrides (clock, scale, gains) beyond the training distribution
    • hardware feels "too fast/aggressive" and a quick knob exists
    • deciding between a deploy-side tweak and a retrain
    “0.80s | 1.25Hz | 0.165 | 3mm(拖地) | 20.0° … 慢一半直接崩 … 策略按 0.40 训练, 别的周期属分布外。… 降指令速度才是有效杠杆 … cmd 0.1 是最稳的工作点。… 要节奏本身变慢必须重训 —— 训练期把 cycle_time 随机化(如 0.40~0.65s)并作为观测的一维, 部署时 --cycle-time 就成了现场可调的旋钮。”
    train/WALK_DIAGNOSIS.md § 2026-08-01 追加: 调慢步态时钟(--cycle-time)在仿真里是反效果
  • Calibrate a wall penalty by measuring healthy and sick policies - healthy pays ~0, the disease pays a wallcalibrate-threshold-between-healthy-and-sick
    Mechanism understoodwalkreward-shapingreward-shapinggate-battery

    Calibrate every threshold penalty by evaluating its exact formula on replays of at least one healthy and one sick policy: place the threshold between their distributions, size the weight so the sick policy pays a decisive fraction of tracking while the healthy one pays ~0, and pre-compute neighboring thresholds for cheap adjustment.

    Symptom

    Real walk_v7 occasionally clipped its own legs (stance narrowed to 133 mm mean vs nominal 214.5); a foot-distance penalty was needed, but an uncalibrated threshold/weight risked either doing nothing or becoming a reverse barrier.

    Context

    The term (relu(d_min - lateral foot distance), measured in the base yaw frame because world-frame y is meaningless after turning) was calibrated before training by replaying three known policies through the exact reward formula at the acceptance operating point: healthy v5 (183 mm) pays 0.6% of tracking - effectively free; narrowed v7 (133 mm) pays 22% - effective widening pressure; collapsed v8 (93 mm) pays 55% - a wall. d_min 0.16 was placed deliberately between healthy and sick, with alternative thresholds (0.14/0.18) pre-computed in the tool output for later adjustment. The shape self-check was named as a standing question: "先问'零代价的选项是什么'" - the zero-cost region must be exactly the desired behavior.

    Change

    feet_lateral_distance added at d_min 0.16 / weight -10, with telemetry expectation pre-registered (should decay toward 0 as stance learns >160 mm; if bow-legged over-widening >214 appears, only then discuss an upper bound).

    Outcome

    v9_probe on hardware: no leg contact ("没碰腿(N2 兑现)"), stance min 145/126 mm green; the term's zero-cost design left healthy gait untaxed.

    Mechanism

    A relu threshold penalty defines a free region and a priced region; its correctness is entirely in where the boundary sits relative to the healthy and pathological distributions. Replaying known-good and known-bad policies through the exact formula measures both distributions in the term's own currency, making the threshold and weight a placement decision instead of a guess.

    Applies when

    • adding any relu/threshold-style wall penalty
    • a safety margin (foot distance, joint limit, clearance) needs enforcement without taxing normal behavior
    • choosing between candidate thresholds for a new term
    “形状自检 (v6 横向组/v8-B 的教训 —— 先问"零代价的选项是什么"): 标称站距付 0, 健康步态付 ~0, 收窄才付费 … v5(健康) 183 mm … 0.6% ≈ 免费 | v7(收窄) 133 mm … 22% —— 有效推宽 | v8(塌陷) 93 mm … 55% —— 墙 … d_min=0.16 恰在 v5(183)与 v7(133)之间”
    train/WALK_V9_SPEC.md § 2. N2 —— 脚距惩罚(已完成权重预标定)
  • A real-robot verdict is (policy x deployment stack) - when the stack changes materially, old verdicts expirestale-verdicts-under-old-stack
    Mechanism understoodwalkreal-acceptancereal-acceptanceattributionprocess

    Date every hardware verdict with the deployment-stack version it was measured under; after any material stack change, re-test before trusting old condemnations or old praises - with the interpretation of each possible result written down first.

    Symptom

    walk_v5 stood condemned as "kicks wildly" and walk_v6 as "cannot walk unassisted" - but those verdicts were issued under an earlier deployment stack (--heading did not exist yet, several fixes had just landed); only v7 had ever run under the current unified stack.

    Context

    New sim evidence sharpened the doubt: a same-harness four-version sweep showed v6 was the HEALTHIEST archive at the real operating point (cmd 0.15: dominant frequency locked at 2.50, steps 35:35 perfectly symmetric, foot distance 203/190 mm best of four, mean tilt 4.2 deg, saturation 15%). A full re-test under the unified stack was scheduled with per-version questions and a pre-filled interpretation table ("判读表(预填假设,回来对号)"): e.g. v6 walks + frequency ~2.5 -> old verdict was the stack's fault, v6 becomes the comparison champion; v6 walks but at ~1.25 -> period-doubling on hardware = confirmed plant gap, actuator fitting promoted to mainline; v5 no longer kicks -> the kicking was an old-stack artifact.

    Change

    All four versions re-queued on hardware under one stack (same torque limits, slew profile, heading loop, logging), with the version-specific legacy profile pinned; verdicts held provisional until re-issued.

    Outcome

    The re-test design separated policy properties from stack artifacts before any policy was permanently written off - and turned each outcome into a specific conclusion via the pre-filled table.

    Mechanism

    A deployed behavior is produced by the policy plus everything between it and the motors (heading loop, slew limits, torque caps, clock); verdicts implicitly condition on that whole stack. Fixing the stack invalidates the conditioning, so old failures may be stack artifacts and old successes may not survive either.

    Applies when

    • deployment tooling (limits, filters, loops) changed since a policy was last judged
    • deciding which historical policy is the rightful baseline
    • a sim sweep contradicts an old hardware verdict
    “只有 v7 在完整的今日部署栈下上过真机 … v5"左右乱踢"、v6"未能自主"的判决全部来自更早的栈(--heading 尚不存在, 部分修复刚落地)——判决已过期。且 2026-08-02 四代同机仿真横测翻出了新证据:v6 在真机工况(cmd 0.15)下是四代里最健康的仿真档案”
    train/REAL_SWEEP_V5_V8.md § 0. 为什么重测
  • One fixed acceptance matrix for every rung - new skill must PASS while every old skill stays within a regression budgetfixed-acceptance-matrix-per-rung
    Replicatedomnigate-batterygate-batteryprocesscurriculum

    Freeze one acceptance matrix for the whole ladder; every promotion requires the new skill's PASS plus bounded regression on every prior skill, measured against the parent's baseline under the current (bug- fixed) metric code - and include command transitions, not just steady states.

    Symptom

    Sequential skill training silently trades old skills for new ones (C1 trained away the root's backward ability); without a constant measurement frame, each rung's numbers are incomparable and regressions hide.

    Context

    The C ladder ran the same 13-cell command matrix at 20 seeds per cell at every rung (stand; vx +0.15/+0.30; vx -0.10/-0.20; vy +/-0.10; wz +/-0.20; two vx&wz combos; two vx&vy combos), with promotion requiring "新技能 PASS 且旧技能不明显退化" - old-skill regression budget <=2/20 against the parent's recorded 20-seed baseline. For the transition-rich final rung, a command-switch block was added (forward->stop, stop->backward, forward->turn, left->right, turn->forward; survival + re-track within 2 s) because "每个 steady command 都会做 ≠ 命令切换不会摔" - steady-state success does not imply switch safety, and the joystick does switches. The C4 product's gate ran 260 cells (13 x 20) all 20/20.

    Change

    Battery frozen once, reused verbatim per rung; baselines re-measured per parent (and re-measured again after the metric-frame fix, since old baselines were taken with the buggy coordinate reading - "旧基线是坏坐标系的, 不可引用").

    Outcome

    Regressions were caught at the rung that caused them (C1's backward loss, C2's vx+0.30 decay), and cross-rung comparisons stayed valid for the ladder's whole life.

    Mechanism

    A constant matrix makes every rung's output a point in the same metric space, so "did we lose anything" is a table diff, not a judgment call; the per-skill regression budget converts previously earned PASSes into standing constraints on all future training.

    Applies when

    • designing gates for sequential skill addition
    • promoting a checkpoint to be the next rung's root
    • after any evaluation-code fix (old baselines must be re-measured)
    “新技能 PASS 且旧技能不明显退化才晋级。… C5 追加:命令切换验收(steady ≠ transition) forward→stop、stop→backward、forward→turn、left→right、turn→forward,各 20 seed,判存活 + 切换后 2 s 内是否重新跟上。”
    train/C_LADDER_RUN.md § 5. 固定验收矩阵(每级跑同一张,每项 20 seed)

Next page