Skip to content

Training Coach

The coach reads a training run and returns a diagnosis, in-whitelist proposals and an experiment plan — and every claim cites one of the cards below. This is that corpus: 22 methodology rules distilled from a real sim2real programme, and 171 episode cards behind them, each with the sentence in the war history it came from.

Doctrine

A report may cite any of these as doctrine-N.

  1. doctrine-1Contract freeze and fingerprint discipline

    The policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.

    Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).

    Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.

    contract-freeze-and-checkerlegacy-profile-pinningderived-asset-staleness-checkgain-profile-belongs-in-the-stamprecovery-two-policies-and-a-state-machinewalk-recovery-fsm-handoff

  2. doctrine-2Attribution by resolved training params - never eval-override knobs

    Capability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.

    Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).

    Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.

    kd-bandwidth-mu-law-attributiondeploy-knob-attribution-before-retrainingcycle-time-override-is-oodresolved-config-is-source-of-truth

  3. doctrine-3PASS gates become constraints; FAIL gates become objectives

    Once a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.

    Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.

    Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".

    fixed-acceptance-matrix-per-rungpreregistered-stop-criteria-per-rung

  4. doctrine-4One variable per ladder rung - counted against what the checkpoint saw

    A rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.

    Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.

    Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.

    resume-state-dr-auditorthogonal-batch-with-ablation-orderfill-the-missing-factorial-cell

  5. doctrine-5Pre-register risks, readings, and stop criteria before the ladder

    Before a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.

    Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).

    Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.

    preregister-risks-and-fork-readingspreregistered-stop-criteria-per-rungpreregistered-real-expectationsfeasibility-accounts-lock-design-point

  6. doctrine-6Plant parameters are measured, never invented

    Every plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.

    Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).

    Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.

    friction-measured-not-guessedarmature-n2-rotor-inertiatorque-limit-shape-by-measured-peakslatency-lerp-reverse-extrapolationfeasibility-accounts-lock-design-pointsim-api-friction-columnsget-up-feasibility-accounts-before-trainingsingle-support-gain-authority-probe

  7. doctrine-7Sim2sim gate before sim2real - under deployment conditions

    Every checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).

    Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).

    Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.

    sim2sim-gate-before-sim2realeval-plant-honesty-contact-paramspipeline-latency-is-plant-not-drbody-frame-velocity-api-audittorque-penalty-bought-by-leg-bracingtorque-disagreement-between-simulators-unresolved

  8. doctrine-8Observation honesty - the actor's inputs are a hardware contract

    The actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.

    Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).

    Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.

    observation-honesty-critic-onlyhistory-obs-needs-plant-variationreward-observability-limitdeploy-heading-loop-and-align-training

  9. doctrine-9Reward economics are audited in realized currency

    Reward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.

    Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).

    Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.

    realized-contribution-auditreward-cost-of-ignoring-auditgate-new-reward-terms-by-commandignore-floor-diagnosiscalibrate-threshold-between-healthy-and-sickcalibration-threshold-with-withdrawal-clauseinert-reward-term-auditseated-basin-dead-exp-kerneltail-torque-needs-hinge-on-computed-demand

  10. doctrine-10The zero-cost option must be the desired behavior

    For every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.

    Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).

    Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.

    penalize-the-slip-not-the-jointsaturation-cheating-zero-rate-costmoving-gate-42x-stand-taxtermination-closes-degenerate-basinpenalty-gate-is-an-escape-hatchsoft-limit-penalty-charges-nominal-poseunpriced-foot-attitude-is-a-free-variableenumerate-cheapest-cheats-before-trainingbinary-band-reward-fake-touchdown

  11. doctrine-11Measurement discipline: independent referees, signs, distributions

    A disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.

    Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).

    Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.

    independent-referee-for-metric-disputesbody-frame-velocity-api-auditsame-sign-response-is-yaw-biasmedian-hides-bimodal-distributionratio-metrics-need-absolute-checkheading-integral-not-body-ratesame-distribution-reward-comparisonmultiseed-sign-test-for-driftsingle-impulse-recovery-is-chaotic

  12. doctrine-12The deployment pipeline is plant

    Irreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.

    Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).

    Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.

    pipeline-latency-is-plant-not-drpower-scale-hurts-nonforward-axesdeploy-scaling-not-training-equivalentteleop-command-band-per-axislatency-dr-covers-measured-pipelinedeploy-rate-limiter-windupslew-anchor-is-an-integratorbeta-anchored-action-targetpower-derating-cuts-full-range-contract

  13. doctrine-13DR budget is finite; its distribution is the measured support

    Robustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.

    Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).

    Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.

    push-dr-conditional-budget-conservationdr-tail-plant-continuationconstant-value-dr-overfits-margintask-shaping-before-plant-hardeningcom-randomization-forces-leg-spreadcom-dr-rollback-on-symptomthin-dr-judged-by-channel-coveragefriction-priority-re-measured-after-plant-change

  14. doctrine-14Gates measure what hardware feels: posture, margins, stripped assists

    Acceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.

    Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).

    Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.

    task-metrics-vs-posture-metricsstand-gate-posture-not-survivalconstant-value-dr-overfits-marginsuspension-probe-removes-free-stabilizerhip-roll-sum-predicts-lateral-driftchirality-scored-separatelylow-speed-commands-reward-draggingend-state-confusion-matrixfrozen-acceptance-distribution-and-pinned-seedvideo-as-acceptance-recordepisode-length-bounds-what-a-gate-seesremoved-wall-returns-on-hardware

  15. doctrine-15Fork and root selection: recoverability, maturity, frozen rewards

    Choose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.

    Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).

    Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.

    fork-root-recoverable-shortfallroot-maturity-vs-product-qualityfine-tune-reward-change-falsified

  16. doctrine-16Curricula: verified engagement, lineage counters, disease-phase gating

    Automatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.

    Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).

    Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.

    auto-curriculum-engagement-checkcurriculum-counter-lineage-stepsgate-penalties-to-the-disease-phaseaggregate-metrics-mask-subgroup-failurebucket-share-is-not-a-gradient-levercurriculum-criterion-conditioned-on-lagging-categoryper-step-income-drives-speed-time-gatetime-gate-vs-wide-stance-retire-the-fix

  17. doctrine-17Probe before training: feasibility first, hypotheses in tables

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.

    Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).

    Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).

    open-loop-probe-before-reward-tuninghypothesis-table-code-auditperiod-doubling-evidence-racesuspended-test-isolates-actuator-authorityconfiguration-probe-wall-not-slopeprone-dead-end-is-foot-placementdof-vel-penalty-is-not-a-pacing-knobamplitude-cut-falsified-yaw-fix

  18. doctrine-18External advice is recomputed locally; values transfer as ratios

    Every external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.

    Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).

    Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.

    external-advice-audit-against-own-arithmetictransfer-ratios-not-absoluteslatency-dr-covers-measured-pipelinereference-structure-fk-amplitude-divisioncycle-average-tracking-for-gait-quantitiesadvisor-paraphrase-vs-paper

  19. doctrine-19Hardware sessions are scripted experiments, not tuning sessions

    Real-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.

    Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).

    Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.

    risk-ordered-real-deploymentbattery-bracketed-real-abknow-zero-command-behaviorpush-test-chirality-protocolno-field-tuning-protocolreversible-single-variable-field-experimentssim-veto-needs-real-confirmationfirst-real-get-up-violent-stage-one-policystaged-hang-mat-floor-for-get-uppower-cycle-preflighttwo-machine-config-disciplinefall-guard-becomes-a-statehardware-log-is-the-attribution-input

  20. doctrine-20Close questions in writing; restart when the debt is structural

    Audited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.

    Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).

    Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.

    frozen-verdicts-semantic-boundariesstale-verdicts-under-old-stackzero-offset-calibration-shifts-envelopeplant-swap-invariants-vs-shiftsfreeze-lineage-fix-structure-restartminimal-reward-table-with-provenancewrite-hardware-verdicts-back

  21. doctrine-21Name the quantity in the space it lives in

    A goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.

    Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).

    Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.

    joint-space-proxy-for-task-space-quantitycontact-detector-single-signal-lieszero-partial-credit-is-not-an-iteration-problemheading-integral-not-body-rate

  22. doctrine-22Continuation needs a live gradient; a release is chosen by a scan

    Continue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.

    Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).

    Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.

    converged-continuation-is-poisoncheckpoint-choice-is-a-full-gate-scanstance-decided-by-get-up-pathstop-stacking-roll-back-and-auditcurriculum-history-is-part-of-the-productcontinuation-budget-not-from-zeroroot-maturity-vs-product-quality

Experience cards

131 cards matching “get-up-feasibility-accounts-before-training”.

  • Before training a one-leg stand the spec named the cheapest cheats - hopping on the support foot, a raised foot resting unloaded, a leg tripod - and gave each a countermeasure and a gate; one still appeared and was caught by exactly those gatesenumerate-cheapest-cheats-before-training
    Observed onceonelegreward-shapingreward-shapinggate-battery

    Before training, list the cheapest behaviours that would satisfy each reward term without doing the task, give each a countermeasure in the reward and a gate in acceptance, and prove the intended behaviour is reachable with a probe - then treat any gate the policy games as evidence about the reward, not the gate.

    Symptom

    The literature's single-leg benchmark reports eight state-of-the-art general policies holding a clean one-leg stand 0 times out of 90 - they survive by sneaking steps and hops - so the task's first adversary was the policy's own cheating.

    Context

    The spec's shape self-check ("what is the zero-cost option?") listed, for the one-foot bucket: the cheapest cheat, a foot resting on the ground without load, countered by a 5 N contact threshold plus positive swing income; the second cheapest, small hops on the support foot to reset balance, countered by a continuous support-air penalty plus a gate of zero support-foot flight segments. The probe that preceded training had already seen a third: early low-lift postures "survived" by pressing the swing foot at 78-95 N, a leg tripod, removed by folding the shank back. The two-foot bucket was checked too: its zero-cost behaviour is ordinary standing, with no odd base state.

    Change

    Countermeasures and gates written before training: swing-contact and support-air penalties, gate 2 (zero swing-foot contact frames above 5 N), gate 3 (zero support-foot flight segments).

    Outcome

    The first run still found the unloaded-foot cheat (a binary reward band gave it no gradient to lift) - and it was caught, by the contact gates and the cross-simulator comparison, not discovered on hardware. The retrained V0 passed all gates 40/40, including zero support-foot flight after the flight detector was corrected.

    Mechanism

    A policy optimizes the reward, not the intent; the cheapest behaviours that satisfy the reward are predictable from the reward's structure, and a gate written for each before training turns a silent cheat into a failed row.

    Applies when

    • designing rewards for balance, contact or "hold still" tasks
    • benchmark policies are known to cheat the task
    • writing acceptance gates for a new skill
    “文献里 8 个 SOTA 通用策略在单脚站基准上 0/90 干净保持, 全靠偷步偷跳活命,这是本任务的第一反作弊对象 … 单脚桶下最便宜的作弊是"脚虚放地上不受力"——接触判定 >5 N 力阈(沿用),配 swing_height_band 正收入拉开。 … 第二便宜是"支撑脚小跳重置"——support_air_penalty 连续罚 + 验收门支撑脚腾空段=0 双保险。”
    git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §0 目标口径 / §5 形状自检(零成本选项是什么)
  • Every gate on a penalty is an exit - to stop paying a stance tax the policy parked 2 deg outside a 30 deg uprightness gate (lunging), and, from scratch, just under a height gate (crouching); a positive band and an always-on guard fixed bothpenalty-gate-is-an-escape-hatch
    Replicatedrecoveryreward-shapingreward-shapinggate-battery

    Never gate a penalty on a state the policy can leave by getting worse; use always-on guards for what must never happen and positive, gated bands for what you want, and count every gate on a penalty as one more escape route to check in the logs.

    Symptom

    V2.9 added stance_width_task = relu(0.34 m - foot spacing) x standing gates, weight -10. Three checkpoints scored 0% on acceptance: standing height reached, feet on the ground, angular rate low, but the torso leaned 32.5/32.3/31.9 deg in a fore-aft lunge. The training dashboard read "tax paid off, base_height at full value".

    Context

    The penalty was gated by uprightness (tilt < 30 deg) and standing height; its tax had no time gate (500 steps x -1.65) while the standing income sat behind a 3 s zero gate (~325 steps), so leaning just past 30 deg lost a little gated income and saved the whole tax. base_height has no upright gate, so the lunge still collected it. The first metric also measured full horizontal spacing, so a staggered lunge counted as "wide".

    Change

    Two laws written down: a penalty may only carry gates the policy cannot escape by getting worse (make it an always-on guard) or it becomes a positive band ("not earned" is not "escaped"); and width is measured laterally in the base yaw frame. From scratch (V3.1 P1) with the lateral metric but the same gates, the policy parked just under the height gate instead (base_height 1.176/1.5, h ~ 0.30 m against a 0.3264 gate), the beta curriculum never advanced in 1,700 iterations, and the run was stopped early. P1b flipped the penalty into a positive band +2.0 x clamp(lateral / 0.34) x standing gates; P1c added yaw_guard = -5 x relu(|hip_yaw| - 30 deg), always on, no gate, no exemption.

    Outcome

    P1b: lateral stance 0.364 m, all four categories 100%, MuJoCo mu 1.0/0.4 both 100%. P1c: all six acceptance criteria passed for the first time on the line (hip-yaw saturation 1.2%), the guard's tax converging to -0.016 (almost never touched).

    Mechanism

    A gated tax that is not paid is saved, so the policy moves to the cheapest state just outside the gate; a gated income that is not earned is simply lost, so a positive band has no exit. HoST's style penalties are ungated or binary - the same law seen from the other side.

    Applies when

    • adding a penalty multiplied by an uprightness, height, phase or contact gate
    • a policy settles just beyond a gate threshold
    • training metrics look paid-up while acceptance collapses
    “**罚项的门 = 策略的逃生门**。带直立门的负项可以靠"变得更差"(倾出门外) 全时免税;HoST 的 style 罚全部无门控/二值恰是同一律的反面实证。修律: 负项只许挂"变差逃不掉"的门(上限护栏 always-on),或改正向 band (收不到 ≠ 逃掉)。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §46 定案(血统内第四败 + 两条新律)
  • Write each config's expected hardware signature before the session - and if reality disagrees, change the books, not the conclusionpreregistered-real-expectations
    Replicatedomnireal-acceptancereal-acceptanceprocesssim2sim

    Before hardware runs, write per-config expected signatures and the disagreement rule (hardware outranks sim; discrepancies get recorded, not reconciled); validate the harness by checking it reproduces at least one known real behavior.

    Symptom

    Hardware impressions are easily narrated after the fact; without written expectations, any real-robot outcome can be made to "match" the sim story.

    Context

    The S2 acceptance sheet carried a section titled "sim 侧预注册预期 (事后核对, 不许事后改)" - per-configuration behavioral signatures written before the session: s1e@0.8 the disturbance king (push 159/160, zero chirality, all-mu 20/20) at the cost of speed gates 0/20 and zero-command wander ~0.98 m with -29.5 deg/20 s rotation; fric-3000 "walks accurately but is easier to push over"; fric-2400 neither. Credibility check included: the sim harness had reproduced the already-recorded real behavior (pace in place + right drift + net rotation -30 deg/20 s), which "提高本单全部预期的可信度". The anomaly clause fixed the epistemics in advance: if results systematically disagree with sim, "不改结论改账" - don't massage the conclusion, write the discrepancy into the books, and per the earlier zero-warning lesson, hardware wins.

    Change

    Every hardware session ships with a pre-registered expectation table (signature per config), a baseline-match credibility check, and a written precedence rule for disagreement.

    Outcome

    The A/B session became falsifiable: agreement confirms the proxy, disagreement is booked as a proxy-bias finding rather than argued away.

    Mechanism

    Pre-registration converts qualitative hardware sessions into tests of the sim-to-real mapping itself; a reproduced known behavior calibrates trust in the remaining predictions; and fixing "who wins on disagreement" beforehand prevents authority from drifting to whichever source flatters the plan.

    Applies when

    • planning any hardware acceptance or A/B session
    • the sim harness's credibility in this regime is unestablished
    • post-session write-ups tempt narrative fitting
    “⚠️ sim 复现了真机已记录的「原地踏步 + 右漂 + 净旋 −30°/20s」—— harness 与真机行为对得上, 提高本单全部预期的可信度。… 结果与 sim 系统性不符 → 不改结论改账: 写进 README 该节, 按 「Isaac 指标三次零预警」的教训, 以真机为准。”
    train/REAL_RUN_S2.md § 2. sim 侧预注册预期 (事后核对, 不许事后改) / 4. 异常处置
  • The standing-pose reward had been pulling toward the narrow stance the whole line was fighting - a zero-training kinematic audit of the target vector found it, after first auditing the wrong nominalpose-target-geometric-audit
    Mechanism understoodrecoveryattributionreward-shapingattributionplant-calibration

    Before training on a pose target, audit it with forward kinematics - is it geometrically consistent (feet flat, intended stance, intended width) and is it the frame you think it is (action nominal vs standing default)? A posture term's target may itself be the attractor you are fighting.

    Symptom

    Several rungs aimed at widening the stance failed; the stance stayed narrow as if something kept pulling it back.

    Context

    The audit (MuJoCo forward kinematics, no training): every 5 deg of hip roll widens the stance ~5.5 cm (0.271 m at 5 deg, 0.383 m at 15 deg); at 47 deg of hip yaw a wide stance cannot be flat-footed (residual foot tilt ~0.7 x hip roll), which explained the stalled rungs. The first report also said the stand_pose nominal (hip roll 25, knee 60) has a 63 deg residual foot tilt - but that was the action frame's nominal (the limit-midpoint squat), which stand_pose never used, despite a docstring warning not to mix them. stand_pose's real target was DEFAULT_JOINT_POS: the contract's all-zero pose, legs parallel, ~0.22 m apart.

    Change

    The disease statement was corrected in writing: the narrow stance was not an accidental by-product of proxy traps but the target stand_pose had been actively rewarding. A stored "narrow the stance" knife was marked toxic. V2.8 moved the target to a flat 15-deg stance (sigma 3 -> 1.5, flat_feet margin 5 -> 20 deg).

    Outcome

    V2.8 still failed in-lineage (stance unchanged, feet nearly overlapping, mu 0.4 transfer 2%) - see stance-decided-by-get-up-path - and the from-scratch V3.1 removed both roll joints from stand_pose and put width into a task-space term, which is what finally produced a 0.355 m flat stance. The 63 deg finding was kept as a warning: an action nominal used as a standing target would be a ready-made pit.

    Mechanism

    A posture term with a sharp kernel around the wrong target is an active attractor; every other term fighting it pays twice.

    Applies when

    • a posture keeps returning despite penalties against it
    • a reward uses a default or nominal pose as its target
    • the contract has more than one "nominal" (action frame vs standing pose)
    “上文"stand_pose 的 nominal (hip25/knee60) 残倾 63°"**审计错了对象**:那是 **动作参考系 nominal**(限位中点蹲),stand_pose 从未指向它(函数 docstring 原文即警告"两者别混",还是混了 —— 记档)。 … **修正后的病根陈述:窄站距不是代理陷阱的意外副产物,而是 stand_pose 一直在主动奖励的目标本身**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 勘误(实现时抓到):审计混了两个 nominal —— 真病根比误诊的更直白
  • The action-delay was implemented as lerp - beyond one step it extrapolated BACKWARD, so a whole lineage trained on a fictitious actuatorlatency-lerp-reverse-extrapolation
    Mechanism understoodinfraactuator-modelingactuator-modelingplant-calibrationsim2sim

    Unit-test plant-model code (delays, filters, randomizers) against hand-computed truth across its FULL configured range, not just the nominal case - a delay must be a queue, and any interpolation used outside [0,1] is a silent plant corruption that training will faithfully absorb.

    Symptom

    Every walk/stand model up to v11 had been trained on a silently wrong plant: the action-latency implementation lerp(cur, prev, lag) is only an interpolation for lag <= 1 - at lag 3 it computes 3*prev - 2*cur, a REVERSE extrapolation. With the configured (0, 0.06) s at 50 Hz (lag in [0,3]), about 2/3 of environments were adapting to actuator dynamics that do not exist.

    Context

    Listed as evidence item #1 for the full restart: "全部旧模型训在错误 plant 上" and the head suspect for the real robot's wild kicking. The fix replaced it with a true FIFO delay line (commit 7f13793) plus its own regression test (tests/test_action_latency.py) - but every exported ONNX predated the fix, which is part of why the lineage was frozen rather than patched.

    Change

    Delay implementation rewritten as an honest FIFO with unit tests; the restart baseline trained on the corrected plant from day one.

    Outcome

    A generation-scale training investment was revealed to have a corrupt plant underneath; the class of bug (plausible-looking math that silently changes meaning outside its valid range) got a permanent test.

    Mechanism

    lerp(a, b, w) leaves the segment for w > 1; used as a delay it fabricates high-gain inverted dynamics precisely in the largest-delay draws, so the policy learns compensation for an actuator that cannot exist - and DR then trains robustness to the artifact rather than to reality. No training metric can catch this: the sim is self-consistent, just wrong.

    Applies when

    • implementing or auditing action delay / filtering in a trainer
    • a lineage behaves as if compensating dynamics nobody modeled
    • deciding whether old checkpoints are salvageable after a plant bug
    “动作延迟旧实现 lerp(cur, prev, lag) 在 lag>1 时是反向外插(w=3 → 3·prev−2·cur),配置 [0,0.06]s@50Hz 即 lag∈[0,3],约 2/3 env 在适应不存在的执行器动态。7f13793 已换真 FIFO,但所有 ONNX 均训于修复之前 —— 真机"乱踢"的头号嫌疑。”
    train/OMNI_V0_SPEC.md § 0. 为什么从零 (1)
  • Training the final recipe from scratch in one run - every mechanism the lineage had accumulated - produced 0% and a seated robot; the order in which the lineage acquired those mechanisms was part of why it workedcurriculum-history-is-part-of-the-product
    Observed oncerecoverytraining-runcurriculumprocess

    A recipe that ends a lineage is not a recipe for a from-scratch run: consolidate it as ordered curriculum phases matching how the lineage acquired its mechanisms, check that every curriculum criterion is reachable from the starting policy, and read the run's raw term values, not the total reward, before calling it green.

    Symptom

    V3.0 trained the lineage's whole final recipe from scratch in one 9,000 iteration run - full beta curriculum, prone-conditioned pull assist, friction DR, the 3 s zero gate on standing income, flat_feet - testing the proposition "the product is defined by its configuration, not by its training history". Training looked all green (reward 26.33, episode length 500, 100% time-outs).

    Context

    Read in raw units against v2_6c at the same weights, the green was a seated equilibrium: base_height 0.377 vs 0.640, stand_pose 0.205 vs 0.516, flat_feet 0.0000 (zero because it sits outside its height gate, not because the feet were flat). Acceptance: 0.0% in Isaac at the deployed authority, 0.0% on every MuJoCo friction level, and still 0.0% at the training-time authority (100% seated at 0.222 m, upright and still).

    Change

    The full stdout (316k lines) was read: the beta curriculum's criterion (standing share over 0.35) was met zero times, so beta never left the wide setting and the policy had no experience at the deployed authority; the pull curriculum was stuck on the same criterion. The lineage had escaped the seated basin with immediate income (V2.0-V2.2) and only then added the zero gate to cure rushing (V2.5); from scratch, the zero gate removed the early "stand fast, earn more" gradient needed to escape. Verdict "the curriculum history is part of the product", limited to n = 1. v2_6c stayed the product.

    Outcome

    V3.1 kept the order as explicit phases: P1 from scratch with immediate income (zero gate off) until the curricula advance, P2 adding the zero gate. P1b/P1c escaped the seated basin and passed; P2 was later judged net negative and dropped (time-gate-vs-wide-stance-retire-the-fix).

    Mechanism

    Mechanisms that refine a competent policy (time gates, tight authority) can delete the gradient a naive policy needs, and a curriculum whose advancement criterion the naive policy never meets freezes at its first level.

    Conflicts

    The spec limits the falsification to "this recipe + this curriculum criterion" (n = 1, no seed sweep, no criterion tuning). V3.1's phased run succeeding is consistent with the ordering reading but changed other terms too.

    Applies when

    • consolidating a long lineage of continuation fixes into one clean recipe
    • a from-scratch run with all mechanisms enabled plateaus early
    • curriculum state is not logged or never advances
    “命题:产物由配置定义,而非训练史定义。 … 训练 log(完整 stdout 316k 行)`[beta_anchor]` 仅初始 1 行,**达标 0 次** … V2.0~V2.2 靠**即时计酬**爬出坐姿盆地(§33),站立巩固后 V2.5 才装归零门 治"过快"(§39)。**课程史是产品的一部分。**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §42 V3.0 判决(2026-08-11):从零单 run 全机制 FAIL 于坐姿盆地
  • A 10 s acceptance episode left 6-7 s of standing to observe - a narrow stance held for that window and split on hardware; a gate cannot see instability slower than its own horizonepisode-length-bounds-what-a-gate-sees
    Observed oncerecoverysim-evalgate-batteryreal-acceptance

    Size the standing phase of an acceptance episode, and its disturbances and floor friction, to what deployment will impose; a pass on a short static window certifies only that window.

    Symptom

    v2_6 passed every simulation gate (success 99.6%, re-falls 0-1%) and then, on the real robot, stood up and slid into the splits several times; prone starts stood and then fell backwards.

    Context

    Acceptance ran 10 s episodes; a ~1-2 s get-up left roughly 6-7 s of static standing on a nominal floor with no disturbance. The narrow stance's lateral margin and the straight-knee stance's lack of any flex buffer are both failure modes that need time, disturbance or lower friction to show.

    Change

    The gap was booked as a known blind spot of the gate ("long-duration standing stability") alongside the task-space stance criterion; later rungs added MuJoCo friction sweeps at mu 0.4 to every checkpoint scan.

    Outcome

    The spec through §50 records the blind spot but no longer standing window or disturbance row in the recovery acceptance itself.

    Mechanism

    An acceptance episode observes only the dynamics that unfold within its horizon under its conditions; slow drifts and disturbance-triggered failures are outside it by construction.

    Applies when

    • a policy passes sim gates and fails on hardware after a delay
    • acceptance episodes are short relative to deployment use
    • stability is judged without pushes or friction variation
    “**sim 门为什么没逮住**:10 s episode 起身后只站 ~6-7 s,静态窗口内窄站距 撑得住;真机站立时长/扰动谱在门口径之外 —— 长时站立稳定性记为口径缺口。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §43 判读:sim 门为什么没逮住
  • Slowing the gait clock at deployment is out-of-distribution and backfires - lower the commanded speed instead, or train the knobcycle-time-override-is-ood
    Mechanism understoodwalkreal-deployreal-acceptanceattributioncontract-freeze

    Any deployment override must correspond to a dimension the policy was trained to handle; to make a parameter field-adjustable, randomize it in training and observe it - otherwise use the levers inside the trained envelope (commands) and leave the knob alone.

    Symptom

    Real-robot feedback "walks very fast and unstable" suggested slowing the gait; a deploy-side --cycle-time override existed, making "just slow the clock" a one-flag temptation.

    Context

    A sim sweep of the override on walk_v6 @cmd 0.3 showed monotone degradation away from the trained 0.40 s cycle: at 0.50 s tilt jumped 7.9 -> 13.2 deg and landing force 1.52x -> 2.24x; at 0.80 s (half speed) clearance collapsed to 3 mm - dragging again - with 20 deg tilt. Meanwhile the legitimate lever, lowering the commanded speed with the clock untouched, improved everything monotonically: cmd 0.1 gave 104% tracking, 6.7 deg tilt, minimum slip - the most stable operating point. The file distinguishes the two "slows" explicitly: lower command = smaller steps at the same 2.5 Hz rhythm; a slower rhythm itself requires retraining - randomize cycle_time (e.g. 0.40-0.65 s) during training and expose it as an observation, and only then does --cycle-time become a field-adjustable knob.

    Change

    Deployment guidance: never ship a cycle-time override the policy was not trained under; respond to "too fast/unstable" with lower commands; schedule clock variability as a training-time (contract-level) change if a field knob is wanted.

    Outcome

    The sweep quantified the trap before hardware paid for it (dragging and 2.2x landing force at slowed clocks); cmd 0.1 documented as the stable demo point.

    Mechanism

    The policy is a function fitted around the training distribution; a deploy-side override moves an input (phase rate) to values never seen, so behavior degrades unpredictably - the knob LOOKS like a capability because it exists in the code, but capability lives in the training distribution, not the interface.

    Applies when

    • a deploy tool exposes overrides (clock, scale, gains) beyond the training distribution
    • hardware feels "too fast/aggressive" and a quick knob exists
    • deciding between a deploy-side tweak and a retrain
    “0.80s | 1.25Hz | 0.165 | 3mm(拖地) | 20.0° … 慢一半直接崩 … 策略按 0.40 训练, 别的周期属分布外。… 降指令速度才是有效杠杆 … cmd 0.1 是最稳的工作点。… 要节奏本身变慢必须重训 —— 训练期把 cycle_time 随机化(如 0.40~0.65s)并作为观测的一维, 部署时 --cycle-time 就成了现场可调的旋钮。”
    train/WALK_DIAGNOSIS.md § 2026-08-01 追加: 调慢步态时钟(--cycle-time)在仿真里是反效果
  • Cross-simulator gate (Isaac Lab to MuJoCo) comes before any hardware attemptsim2sim-gate-before-sim2real
    Replicatedwalksim2sim-gatesim2simprocessgate-battery

    Gate every policy through a second simulator with an independently built plant before hardware; treat sim2sim failure as a contract or overfitting bug, and sim2real failure after a sim2sim pass as a plant/actuator gap.

    Symptom

    A policy that only ever ran in its training simulator carries untested dependencies on that simulator's solver, contact model, and defaults; the first place those dependencies surface should not be the real robot.

    Context

    Standing order of operations for the whole Lucen program, recorded as the opening line of the experience log: train in Isaac Lab, gate in MuJoCo, only then go to hardware. The MuJoCo side is the same plant used for evaluation batteries, so a sim2sim pass also validates the exported policy + contract (obs ordering, scales, defaults) outside the training stack.

    Change

    Pipeline rule adopted - every checkpoint must pass the MuJoCo evaluation battery (sim2sim) before it is considered for real deployment (sim2real).

    Outcome

    Used generation after generation as the cheap filter; hardware sessions only ever received policies that had already survived a second simulator.

    Mechanism

    Two simulators disagree exactly where a policy is overfit to simulator-specific artifacts (contact softness, integrator, default parameters, obs conventions); a cross-sim transfer catches contract bugs and solver overfitting at zero hardware risk, so real-robot failures that remain are attributable to genuine plant/actuator gaps.

    Applies when

    • planning the path from training to first hardware trial
    • exported policy behaves differently outside the training framework
    • triaging whether a real-robot failure is contract vs plant
    “先sim2sim - 从isaaclab 到mujoco / 再sim2real”
    Experience.md § opening lines (1-2)
  • A 2-degree joint-zero calibration fix moved the whole runnable envelope - re-test old "cannot run" verdicts after recalibrationzero-offset-calibration-shifts-envelope
    Observed onceomnireal-deployreal-acceptancehardwareplant-calibrationattribution

    Date every hardware verdict with the calibration state; after any zero/mount recalibration, re-test previously condemned policy-power combinations and previously "unexplainable" posture offsets before attributing either to training or model.

    Symptom

    s1d was on record as "only runs at power 0.7" (kicked wildly at 0.8); after a calibration pass, the same policy ran 12 s at 0.8 with no kicking at all.

    Context

    The calibration had fixed a 2.08 deg zero offset on r_hip_roll - exactly the constant error source on the dominant joint of the kicking oscillation loop ("恰是乱踢振荡环主导关节的常值误差源"). The three-generation post-calibration hardware sweep also closed a second case: the robot's mysterious "backward lean" disappeared after calibration, and the sim-real posture difference collapsed from opposite-sign 5+ deg to same-sign ~2 deg ("后仰案实质了结") - the lean had been a sensing/zero artifact, not a mass-model error. Booked consequence: if the s1d recovery re-verifies, "真机可跑档整体 上移" - every policy's runnable power envelope shifts up, and downstream lineages' hardware expectations get revised.

    Change

    Joint-zero and mount calibration promoted from setup chore to a variable that dates hardware verdicts: verdicts about which power/scale levels a policy can run are conditioned on the calibration state they were measured under.

    Outcome

    One policy rehabilitated at a higher power level; one standing sim-real posture discrepancy closed without touching model or training; a pending re-verification booked rather than asserted.

    Mechanism

    A constant joint-zero error acts as a persistent disturbance injected at the feedback loop's most-loaded joint; near an oscillation threshold, removing a 2-degree bias is the difference between a stable and an unstable loop. Since the error is additive and machine-side, it shifts every policy's stability envelope simultaneously - which is why verdicts must carry their calibration date.

    Conflicts

    The s1d rehabilitation awaited one confirming re-run at the time of writing ("待复核一跑坐实") - the offset-as-cause reading is the head suspect, not a closed verdict.

    Applies when

    • a policy oscillates at a power level others tolerate
    • sim and real disagree on a constant posture offset
    • deciding whether to re-test old hardware verdicts after maintenance/calibration
    “发现①:s1d@0.8 能跑了(旧账「只有 0.7 能跑」)——12s 无乱踢。头号嫌疑 = 标定修正:r_hip_roll offset 修 2.08°,恰是乱踢振荡环主导关节的常值误差源。… 发现②:「后仰」标定后消失 … sim-real 姿态差从反号 5°+ 收敛到同号 2°,后仰案实质了结。”
    train/README.md § 真机 @0.8 三代横评(2026-08-07 标定后)
  • Push DR helped one lineage and hurt another at the same dose - robustness budget is conserved and gets borrowed, not createdpush-dr-conditional-budget-conservation
    Replicatedomnidr-tuningdomain-randomizationcurriculumattribution

    Before opening a disturbance-DR rung, measure whether the untrained policy already meets the spec; if training it anyway, expect the benefit to be conditional on the lineage's existing DR load, grade the intensity, and audit retained margins - budget spent elsewhere will be borrowed back.

    Symptom

    The push rung's outcome flipped with the lineage: direct +/-0.6 m/s push failed outright on first attempt (base walking collapsed - kd1.2 scan 0/3 from iter 3300, sim2sim self-falls with pushes OFF - no PASS point existed); staged +/-0.3 then gave the narrow-kd single-working-point lineage real gains (push survival 1/5 -> 4/5) while the SAME dose made the dual-working-point balanced-band lineage WORSE (20-seed survival 18 -> 12/20 plus across-the-board push regression).

    Context

    The four-ladder verdict ("四梯定案", s2e/s2f at both intensities) named the pattern: "push DR 收益条件性" - the benefit is conditional on how much robustness budget the lineage has already spent. The law candidate: "DR 总预算守恒, 平衡带鲁棒性从抗扰余量借" - total DR budget is conserved; a lineage already covering a wide plant band pays for push tolerance out of its disturbance margin. Both S2 ladders therefore closed at the friction rung, with the decisive numerator: untrained push tolerance already met the 4-6 N*s requirement, so the rung was not needed at all ("⑥ push 不训(收益条件性,免训 ±0.6 已达 标)"). The same accounting later justified the C-before-S2 ordering ("push/μ 两轮已实证 DR 预算有限且会被重分配") and trimmed the second S2 pass to three rungs.

    Change

    Push removed from the standing ladder; graded intensity retained as the method IF a lineage ever needs push training; "does the untrained policy already meet the disturbance spec" instituted as the first check before opening any disturbance rung.

    Outcome

    Two rungs (push, ground mu) deleted from the second S2 pass on measured grounds; the ladder's real yield was re-stated honestly as precision, not robustness (speed gate 0 -> 20/20, zero-command drift 0.98 -> 0.06 m, but push 159 -> 125/160).

    Mechanism

    A fixed-capacity policy allocates representation and margin across the training distribution; adding a disturbance axis to a lineage that already spans a wide plant family forces reallocation - the new tolerance is bought with existing margins. Lineages with narrow plant coverage have free budget, so the identical DR dose lands as gain. Benefit is a property of (dose x lineage state), never of the dose alone.

    Applies when

    • proposing push/perturbation training on a hardened lineage
    • the same DR rung helped one lineage and hurt another
    • accounting where a ladder's robustness gains actually came from
    “push DR 收益条件性 —— s2e⑥a (单工作点血统 kd 窄带) ±0.3 得抗推 1/5→4/5; s2f⑥ (双工作点平衡带血统) 同档反而 20-seed 存活 18→12/20 且抗推全面倒退。规律候选: DR 总预算守恒, 平衡带鲁棒性从抗扰余量借。两阶梯均以 ⑤ 摩擦级收官 … 抗推 4~6 N·s 免训已达标。”
    train/OMNI_V0_SPEC.md § 4. ⑥ push 四梯定案 (2026-08-07)
  • Before resuming a checkpoint, diff the current cfg against what the checkpoint was trained withresume-state-dr-audit
    Replicatedomnitraining-runfork-selectiondomain-randomizationattributionprocess

    "One variable per rung" counts variables against what the checkpoint actually experienced: audit the checkpoint's logged training config and align every unintended difference before resuming.

    Symptom

    Two consecutive rungs (C1 back-mode, C2' forward-turn) failed from the same root with the same full-regression signature despite adding different new modes - so the mode was not the cause.

    Context

    Both runs resumed s1e-500 with the then-current cfg, which carried PD band (0.8,1.2) plus three DR events (base_com, joint_friction, push_robot) accumulated by later lineages. Verified on the training machine from the source of truth (the run's logged params/env.yaml): s1e-500's actual training state was PD +/-10% (0.9,1.1) and all three DR events None. Resuming it under the new cfg meant eating 4 new plant variables plus a new mode at once - the intended "1 variable" was actually 5. A worse variant (c1_redo from s2e_pd-1400) added push +/-0.3 to a root that had never seen it: near-total collapse within +100 iters.

    Change

    C2 aligned the cfg to the checkpoint's training state before resuming (PD back to (0.9,1.1), three DR events off) - making the new mode the only true variable. Permanent rule recorded: compare the checkpoint's training-time DR with the current cfg before any resume.

    Outcome

    C2 trained successfully from the same root that had "failed" twice (wz 20/20 with genuine sign-antisymmetric response by iter 700-800); the A/B falsification ("两个不同模式同签名崩") plus the env.yaml verification closed the attribution.

    Mechanism

    A resumed policy is instantly evaluated (and its value function trained) under whatever plant distribution the cfg specifies; every DR term the checkpoint never adapted to is a distribution shift applied on day one, compounding with the intended change. Single-variable discipline is therefore a property of (cfg diff) x (checkpoint history), not of the cfg diff alone.

    Applies when

    • resuming or forking any checkpoint under an evolved config
    • a resumed run degrades broadly within the first few hundred iterations
    • two different changes from the same root fail with the same signature
    “A/B 定谳:两个不同模式同签名崩 → 病因不是模式,是「从 s1e-500 续训」。… s1e-500 训练态 = kp/kd ±10% (0.9,1.1),base_com / joint_friction / push_robot 全 None;而 cfg 里带着 (0.8,1.2) + … 三个 DR —— 从它续训等于一次吃 4 个新 plant 变量 + 新模式 … 永久教训:续训前必须比对 checkpoint 的训练态 DR 与现行 cfg。单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」。”
    train/C_LADDER_RUN.md § 3b. 这不是重复实验 —— 前两次的病根已定位并修掉
  • Friction DR was demoted after a measurement (94% success at mu 0.4 with no friction randomization) and promoted again when the action contract changed and mu 0.4 fell to 76% - DR priorities belong to a plant and contract, not to a taskfriction-priority-re-measured-after-plant-change
    Observed oncerecoverydr-tuningdomain-randomizationsim2sim

    Re-measure transfer along the friction axis for every new action contract or plant, not once per task; a DR priority settled under one action parameterization does not carry to the next.

    Symptom

    Getting up is all scraping and pushing against the ground, and training pinned friction at 1.0, so friction looked like the first thing to randomize.

    Context

    The MuJoCo gate on R0.5 (5 categories x 10 seeds x 4 friction levels) measured 100/100/98/94% at mu 1.0/0.8/0.6/0.4: degradation showed first as time (prone 3.2 -> 5.3 s), not failure, so friction DR was demoted and the DR budget earmarked for mass/COM. After the switch to the beta-anchored action space, V2.2 read 90/94/90/76%: mu 0.4 was now the weak row.

    Change

    V2.3 (single variable): friction DR static (1.0, 1.0) -> (0.2, 2.0), dynamic (0.15, 1.6), the HiFAR range keeping the base dynamic/static ratio; restitution untouched. Continued from v2_2.

    Outcome

    Isaac nominal 99.8% (DR did not hurt the nominal plant); MuJoCo 98/98/96/92% - mu 0.4 76 -> 92%, mu 1.0 back to R3.1's 98% with bounded torque.

    Mechanism

    How much a policy leans on friction depends on how it moves; the spec records that the sensitivity rose after the action contract changed but does not establish why.

    Applies when

    • changing the action space, gains or authority of an existing skill
    • deciding which DR axis to spend the next rung on
    • an earlier sweep justified leaving an axis unrandomized
    “**μ 砍到 0.4(训练值的 40%)仍有 94%**,退化先体现在**用时**(prone 3.2→5.3 s) 而不是成败。μ≥0.8 完全无损。→ **§17 曾把"摩擦随机化提到 R4 第一项"当作优先 事项,这条实测把它降级了**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §21 MuJoCo 复核门 ② 摩擦依赖
  • The latency DR range must cover the measured deployment pipeline - 0-20 ms could not even reach the real 1-2 control stepslatency-dr-covers-measured-pipeline
    Mechanism understoodwalkactuator-modelingactuator-modelingdomain-randomizationhardware

    Measure end-to-end action latency in control steps on your own stack (including cross-process queue boundaries), set the DR range to cover it with margin, and never import a delay count without its control frequency.

    Symptom

    Action latency was randomized over 0-20 ms (0-1 control step at 50 Hz), but the measured deployment path is 1-2 steps: the deploy process writes the target, an independently running BusWorker picks it up on its NEXT cycle, plus CAN round-trip - the training range could not cover the robot's actual latency at all.

    Context

    Fix: widen action_latency_s to 0-0.06 (0-3 steps). The external reference's "uniform 6 steps" was explicitly NOT copied - that number depends on his unknown control frequency; locally, a sweep at 0/1/2/3 steps showed walk_v5 survives all with insensitive metrics, so 6 steps "在我们这里没有依据" (has no local basis). The range was set from the measured pipeline with margin, not from a foreign constant.

    Change

    action_latency_s (0, 0.02) -> (0, 0.06), justified by pipeline analysis (writer/worker cycle boundary + bus time) and bounded by the local latency sweep.

    Outcome

    The DR band now brackets the true deployment latency; the policy trains against the delay it will actually face instead of a fictional sub-step world.

    Mechanism

    Latency DR only immunizes against delays inside its support; a range below the physical pipeline guarantees an untrained distribution shift at deployment. The correct range comes from tracing the pipeline's worst case (queueing boundaries + transport), and foreign step-counts are meaningless without the control rate they were measured at.

    Applies when

    • setting or auditing action-delay randomization
    • deployment uses a separate bus/worker process from the policy loop
    • importing delay-modeling numbers from other projects
    “现行 0~20 ms = 0~1 个 50Hz 控制步, 而实测部署链路是 1~2 步(deploy 写 STATE.target 后, 独立跑的 BusWorker 下一轮才取走下发, 再加 CAN 往返)——现在的区间覆盖不到真机的实际延迟。… 不照抄参考来源的"统一 6 步": 那取决于他的控制频率(未知), 而我们扫过 0/1/2/3 步 … 6 步在我们这里没有依据。”
    train/WALK_V7_SPEC.md § ⑤ action_latency_s 0~0.02 → 0~0.06
  • A DR tail the robot never has is pure cost - stage deterministic plant levels instead of one wide uniformdr-tail-plant-continuation
    Mechanism understoodomnidr-tuningdomain-randomizationcurriculumactuator-modeling

    Set every DR range from the measured deployment distribution and cut tails that hardware cannot produce; when an axis changes the controller's character (delay, major gain regimes), prefer staged deterministic levels with gates over one wide uniform.

    Symptom

    Two consecutive lineages (s1e, s1f) trained under uniform latency DR (0, 0.06 s = 0-3 frames) both converged to drag-glide gaits - buying survival under heavy delay by giving up swing (3.6 mm) - even though the real pipeline never exceeds ~2 frames.

    Context

    The account: roughly 1/3 of training quality was spent on the >2 frame tail that hardware never presents ("uniform 尾部 ~1/3 训练质量 花在真机不出现的 >2 帧上"). The deeper reading came from the user: uniform 0-3 frames is not merely tail-heavy - it "把性质不同的控制系统 混进同一 PPO batch" (mixes qualitatively different control systems into one PPO batch); a 0-frame and a 3-frame plant demand different controllers, and one policy trained on the mixture serves neither. The S2 v2 ladder therefore redefined latency "从「随机化参数」重新定 义为 actuator/control plant 的一部分": deterministic FIFO levels, staged 1 frame then 2 frames (lo=hi so fractional interpolation degenerates to exact N frames, synonymous with the harness --delay N), each level gated by the fixed acceptance battery - a plant continuation, not a randomization.

    Change

    Latency DR replaced by staged deterministic levels covering the measured 1-2 tick reality with no tail; each stage a separate continuation rung with the standard gate and rollback.

    Outcome

    The s2_lag1 rung showed the clean-signal benefit immediately (survival 20/20, heading 6x recovery) with the swing cost booked honestly (21 -> 12 mm, half-pass, ladder paused for adjudication); the drag-glide attractor from uniform tails did not recur.

    Mechanism

    DR asks one policy to cover a plant family; when part of the family is fictitious, the policy pays real capability for fictitious robustness, and when family members demand structurally different controllers, gradient averaging produces a compromise controller optimal for none. A measured, discrete plant set matches the actual deployment support and keeps each rung's training signal coherent.

    Applies when

    • policies converge to degenerate gaits that buy worst-case survival
    • a DR range extends well past the measured hardware range
    • choosing between wide randomization and a staged ladder on an axis
    “两轮实证(s1e/s1f)宽尾延迟 DR 逼出拖地滑行 … uniform 0~3 帧不止尾重,而是把性质不同的 控制系统混进同一 PPO batch;1→2 帧确定性分级 = plant continuation,训练信号干净得多—— latency 从「随机化参数」重新定义为 actuator/control plant 的一部分。”
    train/OMNI_V0_SPEC.md § 4. v2 阶梯 (2026-08-07 用户定)
  • A pull-assist curriculum keyed to a global success share was satisfied by the categories that already worked and withdrew before prone learned anything - conditioning the criterion on prone took it from 2.5% to 98.7%curriculum-criterion-conditioned-on-lagging-category
    Mechanism understoodrecoverytraining-runcurriculumgate-battery

    Measure a curriculum's advancement criterion on the population the scaffold is meant to help; a global success share is met by whatever already works, and the help is withdrawn before the lagging case learns.

    Symptom

    Under the beta-anchored action, supine and side stood reliably while prone still sat (1.3%). A pull-assist curriculum added to help it was withdrawn completely within ~790 iterations and prone moved only to 2.5% (noise).

    Context

    The pull assist follows HoST: an upward force on the base, active only when the torso is within 30 deg of vertical, scaled by body weight (HoST's 200 N on G1 = 0.583 BW -> 56 N here, steps of 5.6 N, ten levels to zero); the product must pass with no assist. It had already taught sit-to-stand in V1. In V2.1 its advancement criterion was the standing-time share over all envs (threshold raised to 0.55 because the share was already ~0.53).

    Change

    V2.2: PullAssistForce with gate_category = "prone" - only envs whose first step classifies them as prone count toward the criterion - and the threshold back at 0.35. A feasibility signal was pre-registered: if prone's share stayed near zero under the full 56 N, return to the roll-over path instead of adding force.

    Outcome

    The prone-conditioned curriculum withdrew level by level only as prone itself passed: prone 98.7%, and the four-category gate passed for the first time on the line (98.6% overall, re-falls 0%, torque gate PASS).

    Mechanism

    A pooled success share is filled by the categories that already succeed (supine/side ~53%), so the scaffold is removed on their account before the lagging category has used it.

    Applies when

    • an assist, guide force or easier setting is withdrawn by a success threshold
    • one task category lags while the pooled metric looks healthy
    • a curriculum ran to completion without changing the lagging category
    “**教训:全局站立占比阈会被存量类别(supine/side ~53%)凑够,拉力在 prone 学会前就撤光了 —— metric 设计失误,不是拉力机制失效**(它在 V1 教会过 坐→站)。 … **V2.2(已启动)**:`PullAssistForce` 加 `gate_category="prone"` —— 达标判据 只统计 prone 类 env”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §34 V2.1 判决(2026-08-10)
  • Train with self-collisions ON (filtering nested-link ghost pairs) - the reward wall prevents, the physics makes cheating impossibleself-collision-physics-plus-reward-wall
    Mechanism understoodwalkplant-calibrationplant-calibrationsim2simreward-shaping

    Never train a contact-risk behavior with self-collisions disabled; enable them with an audited filter list for nested/overlapping pairs (zero contacts across a pose sweep), record the fps cost, and keep a calibrated distance penalty as the preventive layer on top.

    Symptom

    walk_v8 logged 107 frames of leg-on-leg contact while still earning 0.751 tracking score - because training-side self-collisions were OFF, leg clipping was literally imperceptible to the policy ("碰腿在训练里 根本感知不到").

    Context

    Enabling self-collisions naively is its own trap: an Isaac audit had shown PhysX auto-filters adjacent bodies (base-hip clean for free) but nested links generate ghost forces - calf and ankle_roll overlap 65 mm at the zero pose, producing 12x body-weight phantom forces. The v10 recipe: enable self-collisions, explicitly filter only the two nested pairs (l/r calf-ankle_roll), then run a zero-contact audit at three poses (nominal stand, walk crouch, swing-extreme) requiring contact count = 0, adding any residual pair to the filter and re-auditing; a 500-iter sanity run for NaN and an fps-cost record (measured -8.8%). Redundancy with the reward-side foot-distance wall was argued, not assumed: "N2 离得远(奖励侧预防),SC 碰了疼(物理侧兜底)" - the reward keeps distance at range, the physics makes contact hurt - so the v8-style "clip legs and still score" outcome becomes physically impossible.

    Change

    enabled_self_collisions=True + 2-pair filter + three-pose zero-contact audit (re-verified at 0.00 N after the later mass update) + fps budget recorded.

    Outcome

    Leg contact entered the training signal; the audit protocol caught the nested-pair ghost-force hazard before it corrupted training; combined with the calibrated distance wall, later versions held contact = 0 on hardware and in sim.

    Mechanism

    A hazard absent from the training physics cannot be learned about, no matter the reward; but collision meshes that interpenetrate at rest inject large fictitious forces if enabled blindly. Filtered enabling plus a pose-swept zero-contact audit gives true contact physics with no phantom energy - and layering prevention (reward) with consequence (physics) covers both learning and enforcement.

    Applies when

    • real robot self-contacts while training scored it healthy
    • enabling self-collisions on a model with nested collision meshes
    • deciding between reward-side and physics-side fixes for clipping
    “PhysX 自动过滤相邻体(base↔hip_pitch 免费干净),幽灵力只在 calf↔ankle_roll(零位嵌套 65mm,12 倍体重)。… 与 N2 互补不冗余:N2 离得远(奖励侧预防),SC 碰了疼(物理侧兜底)—— v8 那种 107 帧互碰拿 0.751 跟踪分的事从此物理上不可能。”
    train/WALK_V10_SPEC.md § 4. SC —— 训练侧自碰撞(范围已探明,比想象便宜)
  • Run acceptance under measured contact parameters - honest condim/torsional-friction flipped a false PASS into a real-matching FAILeval-plant-honesty-contact-params
    Mechanism understoodwalksim2sim-gatesim2simplant-calibrationgate-battery

    Pin the evaluation plant's contact model to measured values (contact dimension, torsional/rolling friction, mu) before trusting any gate that involves slip, impact, or drift - a gate can only fail a policy for physics its simulator contains.

    Symptom

    walk_v5 passed the old acceptance battery yet failed on the real robot (footfall force, drift, kicking) - the evaluation plant was flattering the policy.

    Context

    The battery was re-run under "honest contact parameters" - condim 4 (adding torsional contact) with measured torsional friction 0.035 - and v5 then FAILED exactly the rows corresponding to its real problems: heading 185 deg (limit 30), support-foot yaw slip 284 deg (limit 80), landing force 1.72x (limit 1.5x), suspended tilt 45.9 deg (limit 10). The slip physics depends on torsional friction, which the default contact model (condim 3) does not even simulate - a slip problem is invisible to an evaluator that cannot represent yaw friction at the foot. Term-sizing measurements for the new rewards were likewise taken under the same honest parameters (cmd 0.45, skipping the 5 s start transient).

    Change

    Acceptance harness pinned to condim 4 / torsion 0.035 (measured); verdicts issued under defaults declared non-citable for these rows.

    Outcome

    Sim acceptance verdicts began agreeing with hardware ("现在失败, 与真机一致"); the v6 fixes could be developed and validated against an evaluator that could actually see the disease.

    Mechanism

    An evaluator is a plant model too: contact dimensionality and friction values decide which failure modes exist in the simulation at all. Evaluating under default contact parameters tests the policy in a world where its real failure is physically impossible, producing structurally false PASSes.

    Applies when

    • sim acceptance passes policies that fail on hardware
    • slip/drift/impact gates run under default simulator contact settings
    • setting up a cross-simulator evaluation harness
    “accept_v2.py 已加三条判据, walk_v5 在诚实的接触参数下(--condim 4 --torsion 0.035)现在失败, 与真机一致:直行 15s 航向累计 <30° | 185° ✗ … 落脚力峰值 <1.5× 体重 | 1.72× ✗”
    train/WALK_V6_MINIMAL.md § 5. 验收
  • Removing a hand trim re-exposed the plant offset it had been silently compensating - and a slope scan told bias from sensitivityhand-trims-hide-plant-offsets
    Mechanism understoodwalkplant-calibrationplant-calibrationattributionreal-acceptance

    Treat hand-tuned trims as undocumented plant measurements: before deleting one, find what it compensates and re-house that knowledge in the model or the reward budget; diagnose posture errors with a sensitivity sweep to distinguish constant bias from gain problems.

    Symptom

    After switching from the old hand-trimmed default to the clean geometric zero, the retrained stand policy's only regression was torso lean: 1.8 deg -> 4.1 deg backward.

    Context

    The old default's ankle-pitch trim (-0.0489/+0.0628) had been pre-compensating a fore-aft COM mismatch; removing the trim removed the hidden compensation, and the posture reward alone was too weak to win it back. A COM sensitivity scan settled what kind of problem this was: sweeping base COM offset -50 to +50 mm gave nearly identical slopes for old and new policies (~0.026 deg/mm) - "不是质心敏感度问题, 是恒定偏置" (not a sensitivity problem, a constant bias). Fix landed in stand_v1b: posture corrected to +0.24 deg while keeping symmetry (<=0.1 deg) and low effort (0.259), disturbance rejection better than both predecessors. Model credibility was checked the honest way: v0's sim prediction at the real COM position (-22 mm) was -2.31 deg lean vs real measured 2.2-3.1 deg - "预测精准命中" - which is what licensed trusting v1b's -0.52 deg prediction. (Side flag from the same file: a sign convention had been documented wrongly in early comments - gravity_base[0] > 0 is forward lean.)

    Change

    Trims retired in favor of explicit modeling: symmetric geometric default plus a posture-reward budget sized to carry the real COM offset; the offset itself known (real COM ~22 mm behind model).

    Outcome

    stand_v1b passed acceptance as the standing lineage's final version; the walk-line requirement "加大躯干姿态惩罚权重" was upgraded from suggestion to mandatory, since walking amplifies what standing tolerates (real walk_v1 hit 26 deg lean vs sim 7.4).

    Mechanism

    Hand trims are plant knowledge stored in the wrong place - invisible, asymmetric, and stale after recalibration; removing them re-exposes the raw plant error. A sensitivity sweep separates the two possible diagnoses (slope change = control problem; parallel offset = constant plant bias), each with a different fix.

    Applies when

    • cleaning up hand-tuned offsets/trims in defaults or calibration
    • a posture bias appears after a default or calibration change
    • deciding whether a lean is a COM-sensitivity or constant-offset issue
    “两者斜率几乎相同(≈0.026°/mm),v1 只是整体多后仰约 2.4° —— 不是质心敏感度问题,是恒定偏置。成因:旧 default 的踝俯仰 trim(−0.0489/+0.0628)本就预补偿了前后质心偏差,换成零位 default 后这份补偿没了 … v0 在真机质心处(−22 mm)的 sim 预测为 −2.31° 后仰,真机实测 2.2~3.1° 后仰 —— 预测精准命中。”
    train/RETRAIN_v2.md § 4b. stand_v1 独立验证结果 / 4c. stand_v1b 验收结果
  • Every acceptance run records video of the very rollout that produced the numbers - the seated basin, edge-standing feet, tangled legs and the narrow stance were all seen on video before, or instead of, a metric catching themvideo-as-acceptance-record
    Replicatedrecoverysim-evalgate-batteryreal-acceptancemeasurement

    Make video a default output of every acceptance run, rendered from the same rollout the metrics come from (fixed views including the feet), and watch it - posture failures are visible before any gate row exists for them; never let video replace or override the numeric gate.

    Symptom

    Posture failures in the recovery line kept arriving as things the numbers had no row for: a kneeling W-sit, feet standing on their outer edges, crossed legs, a stance too narrow to hold on hardware.

    Context

    From 2026-08-09 (user decision) accept_recovery renders offscreen by default, following one env for the whole episode and archiving the clip. The MuJoCo gate's --video renders three views (side, front, feet) of the same rollout the metrics come from, with all plant modelling (delay, push); the older replay-based renderer produced an independent trajectory without delay and was not used for acceptance. Video never gates: if rendering breaks, --no-video keeps the numeric gate running.

    Change

    Video as a default acceptance artifact, named per policy, category and view, reviewed by the user.

    Outcome

    The R0.1 prone clip showed the same kneel-sit as R0; the R3.1 failure clip showed the crossed legs; the v2_5 feet view showed edge standing and led to V2.6; the v2_6 videos led the user to order a real-robot A/B between v2_5b and v2_6. Once, the recorder did not start (P1c final acceptance) and the visual material had to be produced separately.

    Mechanism

    Metrics exist only for failure modes someone anticipated; video shows the unanticipated ones, and rendering the metric rollout itself guarantees the picture and the numbers describe the same episode.

    Applies when

    • setting up an acceptance pipeline for posture-sensitive skills
    • numbers pass but a human reviewer is uneasy
    • sim videos are rendered by a separate replay tool
    “**验收存视频(用户定 2026-08-09)**:`accept_recovery.py` 默认开 Isaac 离屏 渲染,跟拍一个 env 的整局并归档 … 跪坐这类盆地在数字表出现前肉眼先看见,视频是验收的 定性存档,数字门不受它影响”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §5 R0 验收门(预注册)验收存视频(用户定 2026-08-09)
  • mj_objectVelocity returns inertial-principal-axis frame - one API assumption poisoned eval and observations for a whole linebody-frame-velocity-api-audit
    Mechanism understoodomnisim2sim-gatemeasurementsim2simobservation-honestyattribution

    Verify every frame-sensitive API against a hand-computed truth (rotate raw qvel yourself, or command a known world velocity and check where it lands) before trusting any evaluation or observation built on it - especially when a model's inertial frame is rotated from its body frame.

    Symptom

    Sidewalk vy read ~0 under every condition; separately, whole-policy performance was mysteriously mediocre in sim2sim while training-side numbers looked fine. Four training rungs were declared FAIL partly on these readings.

    Context

    base_link's URDF inertial frame is rotated 90 deg about x relative to the body frame (iquat = [0.7071, 0.7071, 0, 0]). mj_objectVelocity(flg_local=1) rotates into ximat - the inertial principal-axis frame - not the body frame, and returns center-of-mass point velocity, not body-origin velocity. Consequences measured: the "vy" column was actually vertical velocity vz (walking at cmd 0.25: old reading +0.0093 vs true -0.0424); the angular velocity fed to the policy in sim2sim was [wx, wz, -wy] - a different quantity than Isaac and the real IMU provide. RMS check over 8 s of walking: y/z axes swapped between v6[:3] and the qvel truth.

    Change

    Fixed sim2sim and both probes to compute ang_b = qvel[3:6] and lin_b = xmat.T @ qvel[0:3] (identical quantity to Isaac's root_ang_vel_b / root_lin_vel_b), with a standalone reproduction script (frame_bug_repro_0809.py).

    Outcome

    Re-scoring the "failed" C4 lineage under correct coordinates reversed the verdicts: c4r4 checkpoints showed vy 80-126% tracking (old reading: +/-2%) and vx+0.30 at 91-95% where the old metric said 0/5 - the bad frame both mis-measured vy and, via corrupted policy observations, systematically depressed all measured performance. Final product passed 260/260 cells.

    Mechanism

    A simulator API's frame convention is part of the observation contract; when the model's inertial frame is rotated relative to the body frame, frame-agnostic use of a "local" velocity silently permutes axes. Feeding a policy an axis-permuted angular velocity is an observation corruption that degrades behavior everywhere, not just on the axis being studied.

    Applies when

    • building or auditing a cross-simulator evaluation harness
    • one measured axis reads near-zero under all conditions
    • sim2sim scores are inexplicably worse than training-side metrics
    • URDF/MJCF inertial frames are rotated relative to body frames
    “base_link 的 iquat = [0.7071, 0.7071, 0, 0] … mj_objectVelocity 用的是这个 … 喂给策略的 base_ang_vel 是 [wx, wz, −wy] —— MuJoCo 侧观测与 Isaac / 真机 IMU 不是同一个量;vy_mean 报的是竖直速度 vz —— 前进 cmd 0.25 时旧读数 +0.0093,真值 −0.0424。”
    train/C_LADDER_RUN.md § 3l. ⚠️ mj_objectVelocity 读的是惯性主轴系 / 3m. 一 bug 坐实
  • A curriculum ramp keyed to the process step counter re-fires on every resume - and shipped policies never saw the penaltycurriculum-counter-lineage-steps
    Mechanism understoodomnicurriculumcurriculumattribution

    Key every curriculum/ramp schedule to lineage-cumulative progress, not per-process counters; and audit where your shipped checkpoints sit relative to every ramp - a penalty that no product ever experienced is not part of your training.

    Symptom

    vx+0.30 died at a fixed relative time in every resumed run: resume at 500 -> slide at 1100, zero at 1300; resume at 700 -> slide at 1300, zero at 1500 - absolute depths offset by exactly the resume offset, relative timetable identical.

    Context

    ramp_reward_weight (the saturation penalty ramp) read env.common_step_counter, which restarts at 0 for every run including --resume. So start_step=600 meant "600 iters after THIS resume", not "lineage iteration 600". The A/B arm pair was the clean proof: their env.yaml differed only in log_dir, only resume point distinguished them, and the omni CurriculumManager had exactly one active term - nothing else could produce that timetable. Second consequence: every shipped checkpoint (s1e-500 at +500, C2-700 at +200, A800 at +100) was selected before its run's +600, so the saturation penalty weight was 0.000 for every product ever shipped - explaining saturation 33% and raw |action| 1.9 against clip 1.0 (hip_roll in bang-bang), i.e. half the heat budget.

    Change

    Two independent recommendations recorded: (1) make the ramp count lineage steps (add the checkpoint's iteration offset on resume) or pin terminal weights in downstream rungs instead of ramping; (2) give the saturation penalty its own rung - never mixed into a skill-learning rung (that would be two variables again).

    Outcome

    Explained the recurring +600 death of the highest-amplitude command and the persistent actuator saturation of all shipped products with one root cause; honest caveat booked (at +800 the weight is only -0.086, small, but vx+0.30 is the command demanding the largest action amplitude, so it is squeezed first).

    Mechanism

    Resumable training splits "the lineage" from "the process"; any schedule keyed to process-local counters silently re-applies its transient to every descendant run, and any product-selection habit that picks checkpoints early systematically samples the pre-ramp regime - the curriculum exists in the config but never in any shipped policy.

    Applies when

    • resumed/forked training with any scheduled reward or DR ramp
    • a metric dies at a fixed offset after each resume
    • shipped policies show behavior a late-schedule penalty should prevent
    “ramp_reward_weight 读的是 env.common_step_counter,它每个 run 从 0 开始,--resume 也不例外。… 原始 run / 臂B | 500 | 1100 = +600 | 1300 = +800;臂A | 700 | 1300 = +600 | 1500 = +800 … 所有出品其实从没见过饱和罚。… 这解释了 sat_max_pct 33%、raw |a| 最大 1.9(clip 是 1.0)—— hip_roll 一直在 bang-bang,而罚它的那一项权重恒 0。热账的一半在这里。”
    train/C_LADDER_RUN.md § 3g. 系统性问题:saturation_ramp 每次 resume 归零
  • A suspended (no-load) test acquits or convicts the actuator before you blame authoritysuspended-test-isolates-actuator-authority
    Mechanism understoodomnireal-acceptancehardwarereal-acceptanceattribution

    Before attributing a failure to actuator authority, measure no-load tracking error and steady-state torque fraction; blame authority only if the task fails while the error grows with demanded force - and then fix gains or targets, not training.

    Symptom

    hip_roll sagged 0.21 rad on the ground and saturation questions loomed over the sidewalk plan - was the roll axis physically too weak (authority), or was something else limiting it?

    Context

    Before C4, the roll-authority question was settled by measurement triage: suspended test (--suspend, feet off ground) showed hip_roll tracking error 0.0008 rad - actuator acquitted; the entire 0.21 rad ground sag is load-induced. Steady-state torque was 25% of limit - 75% margin remains. Since sidewalk needs lateral force, not exact angles, authority was ruled "not a hard limit", with a pre-registered criterion for when it WOULD become one: sidewalk fails to track AND roll error keeps growing - then the fix is raising hip_roll kp or lowering the vy target, not more training.

    Change

    Hypothesis "roll authority insufficient" demoted from blocker to a monitored branch with an explicit trigger condition; C4 proceeded.

    Outcome

    Later open-loop probes confirmed the actuator could produce the behavior (sidewalk feed-forward ran at full amplitude, 5/5 survival), and the eventual C4 failure causes were measurement and reward, never authority.

    Mechanism

    Suspended vs loaded comparison separates the actuator's closed-loop competence from the load path: tiny no-load tracking error means the motor/controller is fine and any loaded deviation is statics (gravity / stiffness budget, kp trading error for force). Torque-fraction measurement then bounds how much force headroom actually remains.

    Applies when

    • suspecting an axis is "too weak" for a new skill
    • large position sag on a loaded joint
    • deciding between hardware fix, gain change, and more training
    “吊挂(--suspend)实测 hip_roll 跟踪误差 0.0008 rad → 执行器无罪,地面下垂 0.21 rad 全是负载所致;稳态占限扭 25% → 仍有 75% 扭矩余量。… 判据:若 C4 出现「侧走跟不动且 roll 误差继续变大」,那才是权限账 … 解法是提 hip_roll 的 kp 或降 vy 目标,不是硬训。”
    train/C_LADDER_RUN.md § 3d. roll 权限:已部分澄清,不是硬上限
  • The deploy-side walk/recovery switch - into recovery at tilt > 65 deg held 0.3 s, back at tilt < 15 deg with angular rate < 1 rad/s and straight knees held 1 s, a 15 s timeout, last action cleared both ways - and the handoff steps the design required are only partly implementedwalk-recovery-fsm-handoff
    Observed oncerecoveryreal-deployreal-acceptancecontract-freezeprocess

    Specify a deploy-time controller switch as hysteretic, time-filtered predicates the robot can measure (proxy what it cannot, e.g. straight knees for height), a timeout that ends in a safe stop, and a complete handoff (history, clock, last action, command ramp) - then test that the code performs every handoff step, because the design document is not the implementation.

    Symptom

    With a recovery policy and a locomotion policy as separate networks, the robot needs a switch: when is it "fallen", when is it "up", and what state must be reset so the next policy does not act on the previous one's history.

    Context

    The 08-09 design: enter recovery when fallen (tilt > 55 deg or height < 0.60 x 0.384 m) for 150 ms, leave for a stand-hold when upright (tilt < 12 deg, height > 0.85 x 0.384 m, feet steady, |omega| < 0.8) for 400 ms - wide entry, strict exit, hysteresis - then a mandatory handoff trio before walking resumes (reset the walking policy's observation history, restart its phase clock at 0, clear its previous action and latency buffer) and a command ramp instead of a jump. The 08-14 implementation in deploy_policy (--recovery-policy): each policy under its own manifest contract (walk: nominal + scale; recovery: beta-anchored), the same rl_default gains, the tilt cutoff disabled; RECOVERY at tilt > 65 deg for 0.3 s; LOCO again at tilt < 15 deg and |omega| < 1 and knees straight (< 0.35 rad) for 1 s - the robot's computer has no height estimate, and straight knees stand in for height so a V3.0-style upright kneel cannot pass as standing; RECOVERY longer than 15 s ends in a safe stop; last_action cleared on both switches; command forced to 0 during RECOVERY; power scaling applies to LOCO only.

    Change

    An open account was written down with it: the recovery end state (0.355 m stance, hip yaw -/+27 deg) is outside the walking policy's training start distribution, so the first test must use stand / zero command as LOCO, and the long-term fix is to widen the walking policy's initial states rather than bend recovery's stance to suit walking.

    Outcome

    The spec records only a successful compile (py_compile); the hang test was left to be done on site. The operator runbook carries the three-step procedure (hang with stand as LOCO, mat and push, then omni walk as LOCO) but no outcome.

    Mechanism

    Two policies trained separately each assume their own history, clock and last action; a switch that carries any of them across feeds the next policy a state it never saw - the design called this "the walking policy seeing a ghost history".

    Conflicts

    The 08-09 design requires resetting history, clock and previous action plus a command ramp; the 08-14 implementation records last_action clearing and a zero command during recovery; the one-leg spec of 2026-09-14 lists the handoff hygiene as specified but not implemented - reset_history() is called by nothing (a 215-dim policy would carry four frames of pre-fall history), the phase clock is not zeroed, and there is no command ramp back to LOCO. No hardware run of the FSM is recorded in either source.

    Applies when

    • switching between separately trained policies on hardware
    • a policy with history or phase observations is re-enabled mid-run
    • the robot lacks a sensor the switching criterion was designed around
    “**判据**:进 RECOVERY = 倾角 >65°(`--fall-tilt-deg`)持续 0.3 s;回 LOCO = §5 真机可测子集:倾角 <15° ∧ |ω|<1 ∧ **膝直 <0.35 rad(NX 无高度观测, 高度门用膝直代理 —— 防 V3.0 型"跪坐但直立"误判)** 持续 1 s (`--recover-hold`);RECOVERY 单次 >15 s(`--recovery-max-time`)安全停。 … **切换卫生**:两向切换 last_action 清零;RECOVERY 态 cmd 强制 0”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §50 FSM 双策略调度(2026-08-14,用户令):deploy_policy --recovery-policy
  • The shipped checkpoint was chosen by scanning checkpoints on the full gate - neighbours 100 iterations apart failed 1 and 38 cells, late checkpoints degraded - never by taking the last one, and training stopped on signals, not on a schedulecheckpoint-choice-is-a-full-gate-scan
    Replicatedonelegsim-evalfork-selectiongate-battery

    Choose a release checkpoint by running the full acceptance battery over a band of checkpoints (including the transfer axis), stop training on measured signals rather than a fixed iteration count, and expect adjacent checkpoints to differ sharply.

    Symptom

    Gate results moved sharply and non-monotonically between checkpoints of the same run, and the last checkpoint was often not the best.

    Context

    One-leg V0r1: the 2,000 neighbourhood was best; from 2,500 on the nominal gates degraded (late overtraining); 2,000 itself had one real micro-hop (17.7 mm over 5 frames); 2,300 was all green and shipped. V0r2: failed cells per checkpoint 2,000:19, 2,100:38, 2,200:1, 2,300:3, 2,400:27, 2,500:12, 3,000:18 - 2,200 shipped. The recovery line learned the same from the other side: stopping v2_6 early at a scheduled point left a policy whose re-fall rate had spiked to 9-22% before consolidation healed it ("stop on signals, not on the schedule"), and a continuation's transfer decayed checkpoint by checkpoint while Isaac stayed perfect.

    Change

    The acceptance rule "scan checkpoints, do not look only at the last one" is written into the one-leg gates (called the S1 discipline); release candidates are chosen from the scan.

    Outcome

    Both one-leg releases were mid-run checkpoints (2,300 and 2,200) chosen by the full 40-cell battery.

    Mechanism

    PPO keeps changing the policy after the gates saturate; with no gradient toward the gate's conditions, later checkpoints wander, so gate quality is a noisy function of iteration.

    Applies when

    • picking which checkpoint of a run to export and stamp
    • a run is stopped at a fixed iteration budget
    • final-checkpoint results are worse than mid-run smoke tests
    “Isaac 侧 S1 纪律: 验收扫 checkpoint,不是只看最后一个。 … 扫描判决: 2000 邻域最优——2500+ 标称面退化(⑤③② 散挂, 晚期过训), 2000 有一例真微跳(L s100 μ1.2, 17.7mm/5帧), 2300 全绿。”
    git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §6 验收门 / §8 核查单 5 与 7
  • The trainer read a stale USD after the URDF mass update - regenerate derived assets and gate on an automated equality instrumentderived-asset-staleness-check
    Mechanism understoodinfraplant-calibrationplant-calibrationprocesssim2sim

    For every derived plant artifact (USD from URDF, generated value files), pair the generation step with an automated source-vs-derived equality instrument, prove the instrument can fail, gate training on its PASS, and re-run physical audits after every regeneration.

    Symptom

    Measured link masses had been committed to URDF/MJCF (total 9.58 -> 9.792 kg, weighed values), but Isaac reads the derived USD asset - which still carried the old masses: a silent 2.2% mass fork between the training plant and the evaluation plant.

    Context

    The v12 checklist made USD regeneration a hard precondition ("硬性 前置") and, crucially, backed it with an instrument: check_usd_mass.py compares USD vs URDF per-link mass AND inertia trace, validated by showing it FAILs the stale asset naming 9 offending links, then PASSes after re-conversion (13/13 links consistent, total 9.7920). The self-collision filter audit was re-run after the regeneration (three poses, 0.00 N) because a regenerated asset invalidates physical audits done on the old one. This milestone was also where the three plants first aligned: "三边 plant 首次对齐(armature+摩擦+ 实称质量)就在这一代".

    Change

    convert_urdf re-run on the training machine, regenerated asset committed, check_usd_mass.py PASS required as an acceptance gate for the generation; dependent audits repeated post-regeneration.

    Outcome

    The 2.2% plant fork was closed before it could distort a generation's acceptance numbers; the staleness class of bug now has a permanent detector instead of a memory.

    Mechanism

    Source-of-truth edits do not propagate to derived binary assets by themselves; any consumer reading the derivative silently trains or evaluates on the old plant. An automated equality check between source and derivative - proven able to fail - turns an invisible staleness into a red gate, and regeneration invalidates every audit performed on the old artifact.

    Applies when

    • editing masses/inertia/geometry in URDF or MJCF sources
    • a trainer or evaluator consumes converted/derived assets
    • plant numbers differ between simulators for no visible reason
    “25ba997 把连杆质量更新为实称值(总重 9.58→9.792 kg,URDF/MJCF 已改),但 Isaac 读的是 train/assets/laika_v2.usd —— 仍是旧质量。… 否则 Isaac(9.58)与 MuJoCo(9.79)质量分叉 2.2%,v12 验收数字失真。验收门:python tools/check_usd_mass.py 必须 PASS … 对旧资产实测 FAIL/9 连杆点名,仪器已验证”
    train/WALK_V12_SPEC.md § 7. 核查单 (⚠️ 先重转 USD)
  • When a contract default changes, old policies must run under a pinned legacy profile - a silent clock swap is out-of-distribution on hardwarelegacy-profile-pinning
    Mechanism understoodwalkreal-deploycontract-freezereal-acceptanceprocess

    Treat every trained policy as bound to the contract values of its training era: version the deployment profiles, pin old policies to their era's profile in every command, and never let a changed default silently apply to an old artifact.

    Symptom

    The walk profile's gait clock moved from 0.40 s to 0.50 s for new training, but versions v5-v9 were all trained at 0.40 s - running them under the updated default would silently feed a 25% slower phase clock to policies that never saw one.

    Context

    The re-test runbook hard-codes --policy-profile legacy_walk_040 into every command for the old versions, with the warning not to omit the flag: the mismatch is invisible (no error, no crash) but puts the policy out of distribution on hardware, where the same file had already documented that off-clock operation collapses gait quality.

    Change

    Deployment profiles versioned per training era; historical policies permanently associated with their era's profile; runbooks write the profile flag explicitly rather than relying on defaults.

    Outcome

    Old policies stayed runnable and comparable after the contract moved on; the silent-mismatch failure mode was closed by convention.

    Mechanism

    Changing a shared default rebinds every old artifact to a contract it was not trained under; unlike a schema break, a value change produces no error - only degraded, unexplainable behavior. Version-pinned profiles make the binding explicit and permanent.

    Applies when

    • changing any default in the deployment contract (clock, scales, gains) while old policies remain in use
    • writing runbooks that mix policy generations
    • a re-tested old policy behaves worse than its era's records
    “2026-08-02 起 walk profile 的时钟改为 0.50(WALK_V10_SPEC §3)。v5~v9 全是 0.40 训的,本文件所有命令已改带 --policy-profile legacy_walk_040 ——不要省掉这个 flag,否则是拿慢 25% 的相位时钟静默喂旧策略(分布外,真机危险)。”
    train/REAL_SWEEP_V5_V8.md § 1. 预检 ⚠️ 时钟改为 0.50
  • Curriculum-gate a penalty to the phase where its disease occurs - early on it only taxes explorationgate-penalties-to-the-disease-phase
    Replicatedwalkcurriculumcurriculumreward-shapingaction-rate

    For penalties aimed at late-stage pathologies (freezing, saturation, degenerate attractors), ramp the weight in only after exploration noise has decayed; anchor the terminal weight to measured healthy-vs-sick raw values, and shift all related tripwires to after the ramp completes.

    Symptom

    The action_saturation penalty, applied from iteration 0 in v8a, taxed exploration itself: with init_noise_std 1.2 the sampled actions paid ~-2.45/step before any policy had formed - while the disease it targets (clamp freezing) is a LATE pathology (v9 froze at iteration ~2624).

    Context

    v10 re-introduced the same penalty behind a curriculum gate: weight 0 until iter 1000, ramping linearly to -1.0 by iter 2000 - present only when the disease can occur, absent while exploration noise dominates. The trust argument was evidence, not hope: in v8a the term, while active, had pulled joint_pos_ref from 0.041 up to 0.155 and climbing - proof it can extract a policy from the frozen pit. Weight magnitudes were anchored to measured raw values (healthy v5 0.310 / v6 0.106 vs frozen v7 1.145 / v9 1.22 per step: at -1.0 healthy pays 6-18% of tracking, frozen pays 60%+, standing ~0). v10c then isolated the gated term as THE anti-freeze mechanism by single variable, upgraded to untouchable status in v11: "S 的门控机制(v10c 单变量铁案:任何情况下 不许撤,只许调终值)" - and v11 dared to relax other penalties only because S stood guard.

    Change

    action_saturation gated 0 -> -1.0 over iters 1000-2000 (later terminal value tuned -1.0 -> -0.5 with the gate mechanism itself frozen); tripwires adjusted to respect the gate's timing (freeze check moved to iter 2500-3000 to give the ramped term its effect window).

    Outcome

    Freezing stopped recurring while early training kept full exploration; the mechanism graduated from experiment to invariant within two versions.

    Mechanism

    A penalty's incidence depends on who occupies its support: early in training that is exploration noise (whose suppression starves learning), late it is the converged pathology. Time-gating aligns the penalty's presence with its target's presence, buying the constraint without the exploration tax - and tripwire timing must then be computed from the gate schedule, not from ungated precedents.

    Applies when

    • a structural penalty punishes exploration in early training
    • a late-onset pathology (freeze/saturation) needs a standing guard
    • deciding when a curriculum ramp should engage
    “v8a 实锤它的病根是"罚在采样动作上"——init_noise_std 1.2 的早期等于罚探索(~−2.45/步);而冻结是晚期病(v9 速率 2624 才死平)… 门控让它只在病发期在场。… v8a 里它在场时 joint_pos_ref 从 0.041 爬到 0.155 且仍在升——有从低谷爬出的实证。”
    train/WALK_V10_SPEC.md § 2. S 保险 —— action_saturation 课程门控
  • Sort external training advice into adopt / already-have / modify / would-trap by recomputing it on your own configexternal-advice-audit-against-own-arithmetic
    Replicatedomniprocessprocessattributionreward-shaping

    Never apply external tuning advice directly: recompute each claim on your own reward table and probe data, classify it adopt / have / modify / trap, and record why - and verify external citations actually exist.

    Symptom

    External AI/literature advice for the omni ladder arrived plausible-sounding but was written without knowledge of this robot's actual reward table, contract, and history; following it blindly would have broken single-variable discipline and, in one case, made sidewalk unlearnable.

    Context

    Before the C ladder, every external suggestion was audited: 3 adopted (ellipsoid command sampling; staged wz bands; command-switch acceptance), 3 already present (unified reward; frame history - the frozen 168-dim 5-frame window; per-100-iter acceptance), 2 modified (stand share kept at 20% to avoid a second variable; back share NOT raised because the probe showed backward works untrained 20/20@67%, so oversampling would only crowd out forward), and 1 flagged as a trap: "start vy very small (0.06-0.15)" - on THIS reward table vy was only an L2 tax, so ignoring a vy=0.06 command costs 0.4% of the vx tracking scale, 28-180x cheaper than ignoring forward, with quadratic shrinkage making small commands weaker still. A separate retrieval-reliability note: two search agents returned fabricated verbatim quotes from arXiv PDFs (2 papers, verified fake and discarded); only HTML/abstract/source-verifiable material was used.

    Change

    Advice classified only after recomputing each claim with local numbers; the "start small" trap was replaced by adding a gated lateral tracking term (the ladder's only true reward surgery) instead of shrinking the command.

    Outcome

    The adopted items (ellipsoid modes, staged wz, transition acceptance) entered the ladder; the trap was avoided; one external factual error (calling s1g the mainline start - it was falsified 0/20) was caught. Later, one initially-dismissed item (sigma=0.15 too narrow) turned out right for a different reason than claimed - see cycle-average-tracking-for-gait-quantities.

    Mechanism

    External advice encodes the advisor's reward table and robot, not yours; the transfer-validity test is whether the claim survives recomputation under your own arithmetic (reward margins, probe baselines, contract freeze). Items that survive become experiments; items that don't become documented traps.

    Applies when

    • incorporating LLM or literature advice into a training plan
    • advice conflicts with locally measured baselines
    • an external claim depends on reward-table details the advisor cannot know
    “其建议 C4「先很小,vy = ±0.06~0.15」—— 在我们这张奖励表下会让侧走学不起来 … 忽略侧走比忽略前进便宜 28~180 倍,且指令越小激励越弱(平方缩放)—— "先很小"在稳定性上对、在梯度上正好把信号缩没了。”
    train/C_LADDER_RUN.md § 1. 外部 AI 训练建议的评估(采纳 / 已有 / 要改 / 会踩坑)
  • Four consecutive fixes were each continued from the previous fix's degraded state until the user stopped the ladder - "change parameters, don't stack errors" - rolled back to the last good checkpoint and audited the target geometry firststop-stacking-roll-back-and-audit
    Observed oncerecoveryprocessprocessfork-selectionattribution

    When successive rungs each start from the previous rung's output and the target symptom does not move, stop, roll back to the last good checkpoint and re-derive the next change from an audit; keep the measurements, discard the stacked remedies.

    Symptom

    After the real-robot splits, the V2.7 ladder tried to widen the stance: a term swap (A), a new stance knife (b), more iterations (甲), a doubled weight (乙). Stance barely moved while hip yaw ratcheted 46.7 -> 49.9 -> 52.5 deg toward its 60 deg limit.

    Context

    Each rung started from the previous rung's output. The user ruled on 2026-08-11 that things had gone wrong from V2.7-A: go back to v2_6 and rethink which parameters to change instead of stacking errors.

    Change

    乙 was killed at start and not counted; the product baseline rolled back to v2_6c model_29399; every measurement and law learned on the ladder was kept ("the data is real; what stacked was the treatment"). Before any new training, a zero-training kinematic audit of the stance targets was run.

    Outcome

    The audit found the stand_pose target itself rewarding the narrow stance (pose-target-geometric-audit) and showed geometrically why the yawed stance could not be widened with flat feet - so 乙 was proven unnecessary without running it. The next in-lineage attempts still failed, which is what established that the stance is set by the get-up path.

    Mechanism

    A rung continued from a degraded state inherits its compensations, so each new fix answers the previous fix's side effects; the yaw ratchet was the visible trace of that stacking.

    Applies when

    • three or more corrective rungs in a row without progress on the target metric
    • a side-effect metric ratchets in one direction across rungs
    • a new rung is being planned from the latest (not the best) checkpoint
    “用户裁:"从 V2.7-A 开始就出问题了,应该回到 2.6 再思考如何改变参数而不是 错误叠加。"认账:A 的补丁 → b 的新刀 → 甲的加时 → 乙的加权,每级都从上级 的**退化状态**续(yaw 46.7→52.5° 的棘轮就是叠加痕迹)。 … 数据是真的,叠加的是处置。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 方法论裁定(2026-08-11,用户):V2.7 全阶梯叫停,回滚 v2_6
  • Three same-shaped judging errors - task metrics (survival, tracking, displacement) cannot stand in for posture metricstask-metrics-vs-posture-metrics
    Replicatedomnireal-acceptancereal-acceptancegate-batteryattribution

    Keep validated posture-class rows (tilt, per-joint L/R asymmetry, temperature) in every acceptance battery alongside task rows; when operator feel contradicts the gates, suspect the metric class before the operator - and never build a new skill on what is actually an asymmetry defect.

    Symptom

    The C2 product judged "full pass" on task metrics (A800: turn-gap 7 pp, vx+0.30 19/20) felt WORSE in the operator's hands than the half-pass 700: A800 tilted up to 12.90 deg (700: 6.64), drifted left while standing, showed larger per-joint asymmetries, and ran its hip_rolls 5 degC hotter.

    Context

    The re-judgment catalogued three same-type metric errors in one campaign: (1) stand judged by SURVIVAL - missed 0.5-1.4 m wandering; (2) stand ranked by DISPLACEMENT - ordering was opposite to real feel (tilt ordering matched); (3) chirality judged by wz-tracking GAP - measured turning symmetry while the robot's actual disease was postural left/right asymmetry, "两个不同的东西,且结论相反". Common pattern named: "我一直用「任务指标」当判据,而真机手感对应的是「姿态 指标」… 任务类指标不能替代它". The fix was already in the data: the per-joint left/right asymmetry table (printed identically by sim2sim and deploy) agreed with hardware in direction on every row - "判据可用、有预测力,我只是没把它写进 PASS 条件". Shipping decision followed the posture read: product reverted to 700 ("又一次「买到 精度、卖掉别的」"), and the C4 root moved to 700 as well, with the sharpest line of the episode: A800's left-drift "like sidewalking" is probably its frontal-plane asymmetry defect, not a capability - "在缺陷上建能力是危险的".

    Change

    Two posture quantities with demonstrated real-robot predictive power promoted into every PASS battery: tilt-max median and per-joint left/right asymmetry (both sim-computable, deploy-homologous); motor-temperature readout added to session close-out.

    Outcome

    Deployment flipped to the posture-better checkpoint; the hip_roll temperature table (43-48 degC vs 25-28) confirmed the earlier 90%-of-heat account; the run-line acceptance battery inherited the posture rows from birth ("任务类替代不了姿态类").

    Mechanism

    Task metrics measure goal attainment under the evaluator's episode definition; posture metrics measure the body state trajectory that operators, motors, and downstream skills actually experience. The two can rank candidates oppositely because task success tolerates postural pathology - so a battery without posture rows is blind to exactly what hardware feel reports first.

    Applies when

    • hardware feel disagrees with a green acceptance table
    • choosing between checkpoints that split task vs posture metrics
    • selecting the root for a skill that resembles an existing defect
    “共同模式:我一直用「任务指标」(存活 / 跟踪率 / 位移)当判据,而真机手感对应的是「姿态指标」(倾角、逐关节左右不对称)。→ 验收判据里必须有姿态类指标,任务类指标不能替代它。… A800 的「左飘像 side walk」很可能 … 是它更大的额平面不对称的表现 —— 在缺陷上建能力是危险的。”
    train/README.md § C2 选点改判 (2026-08-09): 手性判据第三次选错指标
  • Verify changes in the run's resolved config (and checkpoint md5), never in the source you editedresolved-config-is-source-of-truth
    Replicatedomniprocessattributionprocesscontract-freeze

    Attribution and single-variable claims must be made on the resolved per-run config (and checkpoint hashes), not on source diffs; verify every intended variable landed before burning compute, and verify every rollback byte-level against the historical resolved config.

    Symptom

    An intended arm-B config change never reached the training run - the run was grid-identical (117/117 cells) to its C2 predecessor - and the burn was only understood afterwards.

    Context

    The repo's discipline hardened around the logged resolved config (logs/<run>/params/env.yaml) as the only source of truth: (1) the C2 root-cause analysis was performed against the checkpoint's logged env.yaml, not the code ("以真相源 23-19-25/params/env.yaml 核实"); (2) C4 added a pre-flight: grep the landed env.yaml for the new keys, and compare the first checkpoints of the two arms - identical md5 means the variable did not land, stop immediately; (3) the C4 full rollback was accepted only after starting a 1-iter run and byte-comparing its resolved env.yaml against the historical 700-era file (identical except 4 dormant schema fields, each verified to be at its no-op default).

    Change

    Standing pre-flight and post-change verification: dump/diff the resolved config that the run actually consumed; use checkpoint hash equality as a cheap "variable landed" detector between arms.

    Outcome

    Caught the not-landed variable class of failure; made the rollback provably equivalent to the historical training state rather than believed-equivalent.

    Mechanism

    Between edited source and the running experiment sit layered overrides, env-var switches, and registration logic; only the resolved, serialized config reflects their composition. Diffing at that level tests the actual experiment; diffing source tests intent.

    Applies when

    • launching an A/B pair or any single-variable rung
    • rolling back to a historical training state
    • a run behaves as if a change was never applied
    “开训前先验落盘 cfg(上一轮臂B 的改动没进 run,与 C2 逐格 117/117 相同):grep -E "base_com|joint_friction|push_robot|track_lin_vel_y_exp" logs/<run>/params/env.yaml 另:两臂第一个 checkpoint 的 md5 若相同 = 变量没进去,立刻停。”
    train/C_LADDER_RUN.md § 3d. ⚠️ 开训前先验落盘 cfg / 3l. 回退清单(验证)
  • Choose the fork root by which candidate's shortfalls are recoverable, not by headline scorefork-root-recoverable-shortfall
    Mechanism understoodomnifork-selectionfork-selectionprocess

    When picking a checkpoint to fork from, rank candidates by whether their weaknesses are trainable-back, not by current headline metrics; prefer the candidate whose deficits the upcoming training directly pays for.

    Symptom

    Multiple candidate checkpoints for the omni-command ladder root, each best at something different: fric-3000 had the best tracking precision (vx 88-91%) and hardened plant robustness; s1e-500 had lower precision (vx 82%) but was the only candidate that could still walk backward.

    Context

    Root selection ran as a data probe, not a preference vote: 6 candidates x 8 out-of-distribution omni commands x 20 seeds = 960 cells (probe_omni_0808.json). s1e-500 @pw1.0 survived 20/20 in all eight conditions including backward at 67% tracking; the deep-trained fric lineage scored backward 0-3/20 despite better forward precision.

    Change

    Decision criterion made explicit: list what each candidate exclusively wins at, then ask which of those wins the loser could train back. fric-3000's exclusive wins (precision, plant robustness) are both retrainable - precision is directly optimized by the reward, plant hardening is a planned later pass. s1e-500's exclusive wins (backward plasticity 20/20 vs 2/20, disturbance margin 159/160 vs 125/160, push chirality symmetry 40/40 vs 17/40) had all been shown unrecoverable - push-level rungs failed twice, chirality never recovered even with mirror augmentation on. Root = s1e-500.

    Outcome

    s1e-500 carried the whole C ladder; its backward skill was preserved through C2/C4 gates (regress budget <=2/20 enforced), and the final C4 product passed a 260-cell battery at 20/20 everywhere.

    Mechanism

    Training can re-earn anything the objective directly pays for, but capabilities that earlier training destroyed and never restored (plasticity, symmetry, robustness margins) are empirically one-way doors. The information-bearing comparison is therefore recoverability of each candidate's deficit, which the team stated as "独占项的可恢复性正好相反 —— 这就是判据" (the exclusive items' recoverability is exactly opposite - that is the criterion).

    Applies when

    • selecting a resume/fork root among several checkpoints
    • one candidate is more precise but another retains a skill the rest lost
    • planning a task-extension ladder from an existing lineage
    “fric-3000 赢在精度(vx 88~91%…)与 plant 鲁棒性 → 两样都训得回来…;s1e-500 赢在可塑性(C1 20/20 vs 2/20)、抗扰余量(159/160 vs 125/160)、手性对称(推 ±6 N·s 40/40 vs 17/40)→ 三样都训不回来 … 独占项的可恢复性正好相反 —— 这就是判据。”
    train/C_LADDER_RUN.md § 0. 为什么根是 s1e-500(数据,不是偏好)
  • Adapting a lineage to one plant increment needs hundreds of iterations, not thousands - long runs only buy specializationcontinuation-budget-not-from-zero
    Mechanism understoodomnitraining-runcurriculumprocess

    Budget continuation rungs by increment class (hundreds of iterations for plant pins and smooth shifts, ~1000-1500 only for behavior-demanding changes like push), enforce a hard cap with frequent evaluation, and treat remaining budget as a reason to stop, not to continue.

    Symptom

    The default "6000 iterations per rung" (a from-zero-scale budget) was about to be applied to continuation rungs whose only change is one plant/DR increment - overspending compute and, worse, giving each rung thousands of iterations to specialize away retained skills.

    Context

    The 2026-08-07 budget table replaced the default with "最低适应窗口 + 每 100 iter 验收 + hard cap" scaled to the increment's difficulty: fixed-latency levels 300-500 (cap 500-800; the base has already seen in-band values, this only pins the plant); PD full-band 700 (cap 1000; kp+/-20%/kd+/-30% clearly widens the actuator family); COM +/-20 mm 500 (cap 800; a smooth dynamics shift); friction DR 700 (cap 1000; contact and actuator friction change the gait/contact solution together); push 1000 (cap 1500; a non-static plant change requiring recovery behavior - hardest). Rationale: "续训适应一个 plant 增量不需要从零量级的预算,跑长了只是给特化时间". The C ladder reused the scheme (per-rung caps 500-2000 by increment type), and the deep-training hazard got its own name when long runs sold quality ("深适应卖质量" - deep adaptation sells quality).

    Change

    Per-rung iteration budgets set by increment class with hard caps and 100-iter watch loops; checkpoint selection inside the window by the smoke curve, never "run to cap because budget remains".

    Outcome

    S2/C rungs completed in 300-1500 iterations each; the recurring late-run degradations (collapse valleys at 1500+, vx+0.30 decay) fell outside most rungs' caps instead of inside their runs.

    Mechanism

    A continuation rung's learning problem is local robustification around an existing optimum - low sample complexity; iterations past adaptation are spent sharpening onto the current distribution, which is exactly how retained skills and margins erode. Budgets sized to the increment bound both compute and the specialization damage window.

    Applies when

    • planning iteration budgets for a robustification or command ladder
    • a continuation run keeps improving its training metric late
    • retained skills decay in the back half of long continuation runs
    “「最低适应窗口 + 每 100 iter 验收(watch_ckpt --every 100)+ hard cap」—— 续训适应一个 plant 增量不需要从零量级的预算,跑长了只是给特化时间 … ⑥ push | 1000 | 1500 | 非静态 plant 变化,要学 recovery 行为,最难”
    train/OMNI_V0_SPEC.md § 4. 每级 iter 预算(2026-08-07 用户定)
  • Every new penalty ships with a pre-registered withdrawal clause - if healthy gait must pay above the cap, the term stands downcalibration-threshold-with-withdrawal-clause
    Replicatedwalkreward-shapingreward-shapingprocess

    Introduce every new penalty with: the zero-cost-option audit, a replay-calibrated weight formula (healthy pays a fixed small fraction of tracking), and a pre-registered withdrawal condition - and let the clause fire without argument when the calibration says the term cannot be afforded.

    Symptom

    Three same-shaped crashes had established a failure archetype: v4's clearance, v8a's landing window (weight off by 58x uncalibrated), and v6a's bare hip_yaw suppression all combined a zero-cost "don't move" option with a fee on any motion - a reverse barrier that pushes policies toward standing still.

    Context

    The v11 landing-window penalty was therefore introduced under a calibration-threshold protocol: (1) shape chosen with the window tightened (h_gate 0.03 -> 0.02, because 0.03 equaled the clearance target and priced the entire descent); (2) weight from a FORMULA, not judgment: measure the term's raw value on healthy replays (v5/v10b), set w = -(0.10-0.15 x tracking reward) / raw_healthy; (3) withdrawal clause pre-registered: if healthy gait must pay >15% of tracking no matter the tuning, the term is withdrawn to the next version rather than forced in - "不硬上". The companion hip_yaw quieting term ran the same protocol (calibrate on replays, healthy pays <=5%) and was later retired entirely when a structural fix (zero action scale) made its shaping tax unnecessary.

    Change

    Penalty introduction protocol: shape audit (what is the zero-cost option?), replay-based weight formula, healthy-pay cap with a written stand-down condition - all before training.

    Outcome

    The landing term was in fact withdrawn under its clause (v12 records "P5 落地窗口罚 已撤 … 维持撤下"), demonstrating the protocol firing as designed instead of the fourth same-type crash.

    Mechanism

    A penalty's damage mode is mispricing healthy behavior; since the healthy price is measurable in advance on replays, both the weight and the go/no-go decision can be computed rather than discovered by a ruined training run. The withdrawal clause converts "make it work" pressure into a clean deferral.

    Applies when

    • adding any motion-taxing penalty to a working gait
    • a proposed term's weight has no measurement behind it
    • a previous same-shaped term crashed training
    “权重公式而非拍脑袋:先在 v5/v10b 回放上量 h_gate=0.02 的原始值,w = −(0.10~0.15 × 跟踪奖励) / raw_健康;标定门槛:若健康步态无论如何要付 >15% 跟踪,本项撤下留 v12,不硬上 (v4 clearance/v8a-B/v6a 三次同型翻车的教训:代价为零的"不动"选项 + 一动就收费 = 反向壁垒)。”
    train/WALK_V11_SPEC.md § 6. P5 —— 落地窗口罚(三代欠账,标定门槛制)
  • Deployment power derating damages non-forward axes far more than forward - sweep it in sim before deployingpower-scale-hurts-nonforward-axes
    Replicatedomnireal-deployreal-acceptanceactuator-modelingattribution

    Treat deployment power/torque scaling as a plant parameter: evaluate the policy in sim at the exact deployment scale, expect non-dominant axes to degrade first under derating, and either deploy at the training power or train with power randomization.

    Symptom

    Policies deployed at power-scale 0.8 (a safety derating of commanded torque) looked fine walking forward but were weak at backward and turning, inviting the wrong diagnosis "the skill was not trained well".

    Context

    Measured repeatedly: on s1e, going 1.0 -> 0.8 cost forward 18% but backward 58%; on C4-ff800, turn tracking was +25%/+40% at pw0.8 vs +75%/+58% at pw1.0, backward 51-52% vs 97-103%, while forward stayed 96-98% at both. Sim evaluation numbers in the plan were all pw1.0, but the robot was being run at 0.8.

    Change

    Pre-deploy protocol added: sweep the exported policy across power in sim (for PW in 0.8 0.9 1.0: eval_c_matrix --power $PW --seeds 20) and deploy at the first level where both turn directions reach >=50%. For C4 the recommendation was raise the robot to pw1.0 - the sweep showed it nearly free (saturation 47%->33%, left foot-clipping danger zone 25%->6%, cost only tilt 6.7->8.3 deg).

    Outcome

    Turning "weakness" resolved without any retraining; the sim sweep correctly predicted the real-robot signature at both power levels.

    Mechanism

    Forward walking is the reward-dominant, torque-cheapest skill with the most margin; backward/turn/sidewalk live closer to the torque envelope, so a uniform torque derating consumes their margin first. Training ran at power 1.0 (the trainer does no power scaling), so deploying at 0.8 is a systematic underactuation the policy never experienced.

    Applies when

    • deploying with any torque/power derating or safety scale
    • secondary skills (backward, turn, lateral) underperform on hardware while forward walking looks fine
    • choosing the deployment power level for a new policy
    “power 衰减对非前进轴的伤害远大于前进轴(s1e:前进 1.0→0.8 掉 18%,后退掉 58%)。转向是非前进轴,0.8 下很可能明显跟不动。”
    train/C_LADDER_RUN.md § 3c. A-2 上机前先定部署力度档 / 3p. 二
  • Guessed joint friction was 2.5x low and damping 5x high - measure, then DR around nominalfriction-measured-not-guessed
    Mechanism understoodwalkplant-calibrationplant-calibrationdomain-randomization

    Measure frictionloss and damping separately (they need different rigs), put the measured value at DR center, and express DR as an additive band around that nominal - a DR range around a guessed value can exclude the real robot entirely.

    Symptom

    Old MJCF friction values were invented, not measured; when finally measured, every guessed value was wrong by a large factor in some direction.

    Context

    Joint friction split into Coulomb (frictionloss, tau_c) and viscous (damping b). Measured with the robot hung from a crane (吊机测) while armature was measured no-load; the two measurements are deliberately separated. Old MJCF: frictionloss 0.05, damping 0.1, DR joint_friction range [0, 0.1] "凭空拍的" (made up out of thin air).

    Change

    Replace guessed values with measured ones - frictionloss: RS06 0.15 / RS02 0.12 / RS00 0.13 N*m (old 0.05, i.e. 2.5x too low); damping: 0.02 N*m*s/rad on all three motor types (old 0.1, i.e. 5x too high). DR reshaped from an absolute made-up range [0, 0.1] to an additive band around measured nominal: joint_friction_add [-0.05, +0.10].

    Outcome

    "摩擦定稿(与 armature 一起, plant 参数第一次全部来自实测)" - friction frozen as part of the first fully-measured plant; DR now brackets a measured truth instead of spanning an invented interval.

    Mechanism

    Coulomb friction and viscous damping have opposite behavioral signatures (constant-torque threshold vs velocity-proportional drag); guessing both wrong in opposite directions gives a plant that is simultaneously too easy to start moving and too hard to move fast. DR centered on a wrong nominal makes the policy robust to a family of plants that does not contain the real one.

    Applies when

    • plant friction/damping values have no measurement provenance
    • DR ranges are absolute intervals rather than bands around a nominal
    • policy is over- or under-damped on hardware relative to sim
    “测关节摩擦。吊机测,而armature应该空机测试。… frictionloss τ_c (N·m) │ 0.15 │ 0.12 │ 0.13 │ 0.05(低 2.5×) … damping b (N·m·s/rad) │ 0.02 │ 0.02 │ 0.02 │ 0.1(高 5×) … DR │ joint_friction_add: [−0.05, +0.10] 叠标称 │ 旧 [0, 0.1] 凭空拍的”
    Experience.md § 摩擦定稿表 (lines 12-25)
  • Fine-tuning through a reward-table change was falsified (scatter, half-recover, collapse) - continuation training is legal only with the reward frozenfine-tune-reward-change-falsified
    Replicatedomnicurriculumcurriculumfork-selectionreward-shapingprocess

    Never fine-tune through a reward-table change - retrain from zero; reserve checkpoint continuation for frozen-reward plant/DR widening, reset noise_std when branching, and watch for the scatter/half-recover/collapse signature as the abort trigger.

    Symptom

    The s1c A/B experiment: arm B fine-tuned from an existing checkpoint under the revised reward (same contract, same network, changed reward + small DR) and failed with a characteristic signature - scatter, half-recover, fall back ("打散→半恢复→摔回"); arm A trained from zero under the same config won decisively (full shaping lifted swing to 21.6 mm within 500 iters; shipped at 5500).

    Context

    Verdict recorded: "从零 + 强塑形是本机唯一验证过的发育路径" (from-zero plus strong shaping is this machine's only validated development path). The signature became a standing stop criterion in every later rung that touched a reward ("s1c B 臂签名,出现即停"). Crucially the boundary of the law was drawn explicitly when S2 continuation training was proposed: "当年证伪的是「奖励表中途改版的 fine-tune」… S2 续训奖励表全程冻结、只逐级加宽 plant/DR,属另一类" - continuing a checkpoint with the reward FROZEN while widening plant/DR one rung at a time is a different class and was allowed (and then worked, powering the whole S2/C lineage) - with the honest fallback that if frozen-reward continuation ever collapses, that rung retrains from zero and the doctrine gets re-examined with data. Fine-tune arms also need mechanical care: reset the checkpoint's collapsed noise_std (terminal 0.033 "会杀死探索") and account for iteration counters re-zeroing (curriculum gates fire immediately).

    Change

    Reward changes and lineage continuation permanently separated: reward revisions -> from-zero retrain; plant/DR widening -> frozen-reward continuation with per-rung gates; the B-arm signature promoted to a universal tripwire.

    Outcome

    No later reward revision was attempted by fine-tune; frozen-reward continuation carried S2 (PD/COM/friction rungs) and the C command ladder successfully from the s1e root.

    Mechanism

    A trained policy sits in an optimum of its reward's geometry; changing the reward moves the optimum but leaves the policy's exploration noise near-zero and its value function calibrated to the old returns - it disassembles the old solution faster than it can assemble the new one. Widening DR under a frozen reward instead keeps the optimum's identity and asks only for local robustification.

    Applies when

    • proposing to fine-tune an existing policy under a revised reward
    • planning a robustification ladder from a validated checkpoint
    • a continued run scatters then partially recovers then collapses
    “B 臂 fine-tune 证伪(打散→半恢复→摔回——从零 + 强塑形是本机唯一验证过的发育路径)。… 当年证伪的是「奖励表中途改版的 fine-tune」(B 臂,塑形突变致终盘摔回);S2 续训奖励表全程冻结、只逐级加宽 plant/DR,属另一类;若 s2_lag1 续训本身塌方,回退方案 = 该级从零重训,续训教义再议(拿数据说话)。”
    train/OMNI_V0_SPEC.md § 3. S1.3 / 4. 与 s1c fine-tune 证伪的关系
  • An auto-curriculum ratchet capped out at iter 248 and never engaged - stage difficulty manually or verify engagementauto-curriculum-engagement-check
    Observed oncewalkcurriculumcurriculumdomain-randomizationprocess

    Prefer manually staged difficulty with gated transitions; if you use an automatic curriculum, instrument its internal state and alarm when it stops engaging - a saturated curriculum is constant DR wearing a curriculum's name.

    Symptom

    A curriculum mechanism intended to grow difficulty adaptively (s1f's ratchet) hit its cap at iteration 248 and never bit again - for 96% of the run its effect was equivalent to constant DR, i.e. the curriculum existed in name only.

    Context

    When external advice suggested graded wz bands (start ±0.15, then ±0.30), the team agreed with grading but explicitly rejected automatic curriculum, citing the s1f episode. The same logic had already been paid for with push grading: ±0.6 failed twice, ±0.3 was feasible - grading matters, but the grade transitions were made by hand at verified checkpoints.

    Change

    Ladder policy: difficulty staged manually, one band per rung, each transition gated by the acceptance battery; automatic ratchets not used unless their engagement is monitored and demonstrated.

    Outcome

    Every C-ladder band change (wz ±0.12-0.25 first, wider later) was an explicit, attributable rung; no silent constant-DR-in-disguise runs recurred.

    Mechanism

    Adaptive curricula couple their own state machine to noisy training metrics; a ratchet that saturates early stops adapting but keeps its name, so the operator believes difficulty is progressing when it is frozen. Manual staging costs more decisions but each decision is observable and reversible.

    Applies when

    • choosing between auto-curriculum and staged bands for a new skill
    • a curriculum's difficulty parameter plateaus early in training
    • post-hoc attribution of what difficulty a lineage actually saw
    “C2 的 wz 分级(±0.15 → ±0.30,不要一上来 ±0.6)—— 与我们付过学费的 push 分级同形(±0.6 两轮 FAIL,±0.3 才可行)。但必须手动分级,不做自动课程 —— s1f 的课程化棘轮 iter 248 封顶未咬合,96% 时长等价常量 DR”
    train/C_LADDER_RUN.md § 1. 采纳 3 条 (C2 的 wz 分级)
  • Model CAN polling skew - joint observations are 6-9 ms stale by read ordercan-timing-skew-modeling
    Observed oncewalkactuator-modelingactuator-modelingsim2simhardware

    If joints are read sequentially over a shared bus, reproduce the per-group observation staleness in sim (or randomize it over the measured range) - synchronous observations are a privileged fiction.

    Symptom

    Policies trained with synchronous joint observations degrade on hardware where motors are polled sequentially over CAN - hip data is already 6-9 ms old by the time ankle data arrives.

    Context

    Menlo's core sim2real finding on a leg platform of almost identical mass to Lucen's. CAN is a sequential bus: one poll cycle reads motors in a fixed order, so the observation vector mixes timestamps. Lucen runs 6:6 motors on a dual CAN split, so the problem transfers one-to-one.

    Change

    Explicitly model CAN timing skew in sim: give the joint groups different observation delays matching physical read order. Menlo went further - running real firmware in the loop with a motor simulator between MuJoCo and firmware injecting 0.4-2 ms uniformly distributed delay.

    Outcome

    Reported by the reference team as their core sim2real enabler on a same-scale platform; recorded in Lucen's experience log as directly applicable ("这个问题一模一样").

    Mechanism

    A policy exploits any cross-joint temporal coherence present in training observations; when hardware breaks that coherence per bus position, the learned feedback acts on inconsistent state estimates, injecting phase error exactly at the control bandwidth.

    Conflicts

    Second-hand episode: outcome numbers are the reference team's report, not a Lucen-run experiment; Lucen adopted the requirement but the corpus has no Lucen A/B of skew-on vs skew-off.

    Applies when

    • robot polls actuators sequentially over CAN/RS485 or similar shared bus
    • sim2real degradation appears as jitter or oscillation not seen in sim
    • designing the observation/delay model before a training run
    “电机走 CAN 是顺序轮询的,髋部电机的数据到踝部电机上报时已经陈旧了 6-9 ms,他们直接在仿真里显式建模了 CAN 时序偏斜,按读取顺序给三组关节不同的观测延迟。… 注入 0.4–2 ms 的均匀分布延迟。你们是 6:6 双 CAN 分总线,这个问题一模一样”
    Experience.md § 执行器 + 时序建模 (line 6)

Next page