Skip to content

Training Coach

The coach reads a training run and returns a diagnosis, in-whitelist proposals and an experiment plan — and every claim cites one of the cards below. This is that corpus: 22 methodology rules distilled from a real sim2real programme, and 171 episode cards behind them, each with the sentence in the war history it came from.

Doctrine

A report may cite any of these as doctrine-N.

  1. doctrine-1Contract freeze and fingerprint discipline

    The policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.

    Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).

    Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.

    contract-freeze-and-checkerlegacy-profile-pinningderived-asset-staleness-checkgain-profile-belongs-in-the-stamprecovery-two-policies-and-a-state-machinewalk-recovery-fsm-handoff

  2. doctrine-2Attribution by resolved training params - never eval-override knobs

    Capability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.

    Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).

    Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.

    kd-bandwidth-mu-law-attributiondeploy-knob-attribution-before-retrainingcycle-time-override-is-oodresolved-config-is-source-of-truth

  3. doctrine-3PASS gates become constraints; FAIL gates become objectives

    Once a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.

    Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.

    Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".

    fixed-acceptance-matrix-per-rungpreregistered-stop-criteria-per-rung

  4. doctrine-4One variable per ladder rung - counted against what the checkpoint saw

    A rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.

    Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.

    Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.

    resume-state-dr-auditorthogonal-batch-with-ablation-orderfill-the-missing-factorial-cell

  5. doctrine-5Pre-register risks, readings, and stop criteria before the ladder

    Before a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.

    Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).

    Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.

    preregister-risks-and-fork-readingspreregistered-stop-criteria-per-rungpreregistered-real-expectationsfeasibility-accounts-lock-design-point

  6. doctrine-6Plant parameters are measured, never invented

    Every plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.

    Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).

    Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.

    friction-measured-not-guessedarmature-n2-rotor-inertiatorque-limit-shape-by-measured-peakslatency-lerp-reverse-extrapolationfeasibility-accounts-lock-design-pointsim-api-friction-columnsget-up-feasibility-accounts-before-trainingsingle-support-gain-authority-probe

  7. doctrine-7Sim2sim gate before sim2real - under deployment conditions

    Every checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).

    Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).

    Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.

    sim2sim-gate-before-sim2realeval-plant-honesty-contact-paramspipeline-latency-is-plant-not-drbody-frame-velocity-api-audittorque-penalty-bought-by-leg-bracingtorque-disagreement-between-simulators-unresolved

  8. doctrine-8Observation honesty - the actor's inputs are a hardware contract

    The actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.

    Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).

    Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.

    observation-honesty-critic-onlyhistory-obs-needs-plant-variationreward-observability-limitdeploy-heading-loop-and-align-training

  9. doctrine-9Reward economics are audited in realized currency

    Reward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.

    Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).

    Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.

    realized-contribution-auditreward-cost-of-ignoring-auditgate-new-reward-terms-by-commandignore-floor-diagnosiscalibrate-threshold-between-healthy-and-sickcalibration-threshold-with-withdrawal-clauseinert-reward-term-auditseated-basin-dead-exp-kerneltail-torque-needs-hinge-on-computed-demand

  10. doctrine-10The zero-cost option must be the desired behavior

    For every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.

    Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).

    Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.

    penalize-the-slip-not-the-jointsaturation-cheating-zero-rate-costmoving-gate-42x-stand-taxtermination-closes-degenerate-basinpenalty-gate-is-an-escape-hatchsoft-limit-penalty-charges-nominal-poseunpriced-foot-attitude-is-a-free-variableenumerate-cheapest-cheats-before-trainingbinary-band-reward-fake-touchdown

  11. doctrine-11Measurement discipline: independent referees, signs, distributions

    A disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.

    Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).

    Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.

    independent-referee-for-metric-disputesbody-frame-velocity-api-auditsame-sign-response-is-yaw-biasmedian-hides-bimodal-distributionratio-metrics-need-absolute-checkheading-integral-not-body-ratesame-distribution-reward-comparisonmultiseed-sign-test-for-driftsingle-impulse-recovery-is-chaotic

  12. doctrine-12The deployment pipeline is plant

    Irreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.

    Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).

    Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.

    pipeline-latency-is-plant-not-drpower-scale-hurts-nonforward-axesdeploy-scaling-not-training-equivalentteleop-command-band-per-axislatency-dr-covers-measured-pipelinedeploy-rate-limiter-windupslew-anchor-is-an-integratorbeta-anchored-action-targetpower-derating-cuts-full-range-contract

  13. doctrine-13DR budget is finite; its distribution is the measured support

    Robustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.

    Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).

    Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.

    push-dr-conditional-budget-conservationdr-tail-plant-continuationconstant-value-dr-overfits-margintask-shaping-before-plant-hardeningcom-randomization-forces-leg-spreadcom-dr-rollback-on-symptomthin-dr-judged-by-channel-coveragefriction-priority-re-measured-after-plant-change

  14. doctrine-14Gates measure what hardware feels: posture, margins, stripped assists

    Acceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.

    Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).

    Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.

    task-metrics-vs-posture-metricsstand-gate-posture-not-survivalconstant-value-dr-overfits-marginsuspension-probe-removes-free-stabilizerhip-roll-sum-predicts-lateral-driftchirality-scored-separatelylow-speed-commands-reward-draggingend-state-confusion-matrixfrozen-acceptance-distribution-and-pinned-seedvideo-as-acceptance-recordepisode-length-bounds-what-a-gate-seesremoved-wall-returns-on-hardware

  15. doctrine-15Fork and root selection: recoverability, maturity, frozen rewards

    Choose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.

    Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).

    Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.

    fork-root-recoverable-shortfallroot-maturity-vs-product-qualityfine-tune-reward-change-falsified

  16. doctrine-16Curricula: verified engagement, lineage counters, disease-phase gating

    Automatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.

    Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).

    Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.

    auto-curriculum-engagement-checkcurriculum-counter-lineage-stepsgate-penalties-to-the-disease-phaseaggregate-metrics-mask-subgroup-failurebucket-share-is-not-a-gradient-levercurriculum-criterion-conditioned-on-lagging-categoryper-step-income-drives-speed-time-gatetime-gate-vs-wide-stance-retire-the-fix

  17. doctrine-17Probe before training: feasibility first, hypotheses in tables

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.

    Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).

    Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).

    open-loop-probe-before-reward-tuninghypothesis-table-code-auditperiod-doubling-evidence-racesuspended-test-isolates-actuator-authorityconfiguration-probe-wall-not-slopeprone-dead-end-is-foot-placementdof-vel-penalty-is-not-a-pacing-knobamplitude-cut-falsified-yaw-fix

  18. doctrine-18External advice is recomputed locally; values transfer as ratios

    Every external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.

    Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).

    Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.

    external-advice-audit-against-own-arithmetictransfer-ratios-not-absoluteslatency-dr-covers-measured-pipelinereference-structure-fk-amplitude-divisioncycle-average-tracking-for-gait-quantitiesadvisor-paraphrase-vs-paper

  19. doctrine-19Hardware sessions are scripted experiments, not tuning sessions

    Real-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.

    Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).

    Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.

    risk-ordered-real-deploymentbattery-bracketed-real-abknow-zero-command-behaviorpush-test-chirality-protocolno-field-tuning-protocolreversible-single-variable-field-experimentssim-veto-needs-real-confirmationfirst-real-get-up-violent-stage-one-policystaged-hang-mat-floor-for-get-uppower-cycle-preflighttwo-machine-config-disciplinefall-guard-becomes-a-statehardware-log-is-the-attribution-input

  20. doctrine-20Close questions in writing; restart when the debt is structural

    Audited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.

    Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).

    Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.

    frozen-verdicts-semantic-boundariesstale-verdicts-under-old-stackzero-offset-calibration-shifts-envelopeplant-swap-invariants-vs-shiftsfreeze-lineage-fix-structure-restartminimal-reward-table-with-provenancewrite-hardware-verdicts-back

  21. doctrine-21Name the quantity in the space it lives in

    A goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.

    Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).

    Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.

    joint-space-proxy-for-task-space-quantitycontact-detector-single-signal-lieszero-partial-credit-is-not-an-iteration-problemheading-integral-not-body-rate

  22. doctrine-22Continuation needs a live gradient; a release is chosen by a scan

    Continue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.

    Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).

    Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.

    converged-continuation-is-poisoncheckpoint-choice-is-a-full-gate-scanstance-decided-by-get-up-pathstop-stacking-roll-back-and-auditcurriculum-history-is-part-of-the-productcontinuation-budget-not-from-zeroroot-maturity-vs-product-quality

Experience cards

161 cards matching “removed-wall-returns-on-hardware”.

  • With no term on foot attitude the robot stood on the edges of its feet (ankle roll -26 deg) from R0.5 to V2.5b; pricing it fixed that and turned stance width into the next free variable - the user saw both on video before any metric flagged themunpriced-foot-attitude-is-a-free-variable
    Replicatedrecoveryreward-shapingreward-shapingreal-acceptance

    List every posture quantity the hardware cares about (foot attitude, stance width, yaw, knee flexion) and make sure each is priced by a term with real gradient at the observed error; pricing one exposes the next, so re-inspect the feet-level video after every change.

    Symptom

    Watching the v2_5 video the user said the ankle roll after standing looked very strange - the feet were not flat. The feet view showed the right foot standing on its outer edge; median ankle roll at t = 8 s was -/+26 deg (limit +/-35), mirrored. This was the physical form of the ankle_roll saturation criterion that had failed since R0.5.

    Context

    No term priced foot attitude: feet_contact_upright counts an edge contact as contact, and stand_pose's wide exp kernel (sigma^2 = 9) gives almost no gradient at 0.45 rad. HoST carries an "ankle parallel" term (+20); this table never had one. Edge standing was made a blocking precondition for hardware (continuous ankle load, unstable contact, wear).

    Change

    V2.6 (single variable): flat_feet = (|q_l_ankle_roll| + |q_r_ankle_roll|) x upright gate x height gate, a linear hinge with a 5 deg margin, weight -2, continued from V2.5b.

    Outcome

    Ankle roll -/+26 -> 5.3/5.2 deg, ankle_roll saturation 6.6% (< 10%), feet flat on the feet-view video; success 99.6%, re-falls 0-1%. Then the user watched v2_6: hip yaw constantly tense and the legs very close together. The numbers: hip roll -/+4.9/4.8 deg against a nominal 25 - feet flat and hips open 25 deg cannot coexist without ankle compensation, the new term taxed that compensation, and nothing priced stance width. "Foot attitude as a free variable" was fixed and "stance width became the new free variable" - which the real robot then exposed as splits.

    Mechanism

    An optimizer spends every posture degree of freedom no term prices; closing one reallocates the slack to the next unpriced one.

    Applies when

    • a standing or landing posture looks wrong on video while gates pass
    • a saturation criterion keeps failing on one joint
    • a new posture term was just added
    “用户看 v2_5 视频:"起身之后 ankle_roll 非常奇怪,脚根本不是平着站立"。 … 机理:奖励表**无任何脚掌姿态项** —— feet_contact_upright 边缘接触也算触地, stand_pose 的 exp 核(σ²=9)对 26°=0.45 rad 梯度≈0。脚掌姿态是自由变量。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §41 预注册 V2.6:flat_feet —— ⑥ 老账的物理形态被用户目视锁定
  • Model CAN polling skew - joint observations are 6-9 ms stale by read ordercan-timing-skew-modeling
    Observed oncewalkactuator-modelingactuator-modelingsim2simhardware

    If joints are read sequentially over a shared bus, reproduce the per-group observation staleness in sim (or randomize it over the measured range) - synchronous observations are a privileged fiction.

    Symptom

    Policies trained with synchronous joint observations degrade on hardware where motors are polled sequentially over CAN - hip data is already 6-9 ms old by the time ankle data arrives.

    Context

    Menlo's core sim2real finding on a leg platform of almost identical mass to Lucen's. CAN is a sequential bus: one poll cycle reads motors in a fixed order, so the observation vector mixes timestamps. Lucen runs 6:6 motors on a dual CAN split, so the problem transfers one-to-one.

    Change

    Explicitly model CAN timing skew in sim: give the joint groups different observation delays matching physical read order. Menlo went further - running real firmware in the loop with a motor simulator between MuJoCo and firmware injecting 0.4-2 ms uniformly distributed delay.

    Outcome

    Reported by the reference team as their core sim2real enabler on a same-scale platform; recorded in Lucen's experience log as directly applicable ("这个问题一模一样").

    Mechanism

    A policy exploits any cross-joint temporal coherence present in training observations; when hardware breaks that coherence per bus position, the learned feedback acts on inconsistent state estimates, injecting phase error exactly at the control bandwidth.

    Conflicts

    Second-hand episode: outcome numbers are the reference team's report, not a Lucen-run experiment; Lucen adopted the requirement but the corpus has no Lucen A/B of skew-on vs skew-off.

    Applies when

    • robot polls actuators sequentially over CAN/RS485 or similar shared bus
    • sim2real degradation appears as jitter or oscillation not seen in sim
    • designing the observation/delay model before a training run
    “电机走 CAN 是顺序轮询的,髋部电机的数据到踝部电机上报时已经陈旧了 6-9 ms,他们直接在仿真里显式建模了 CAN 时序偏斜,按读取顺序给三组关节不同的观测延迟。… 注入 0.4–2 ms 的均匀分布延迟。你们是 6:6 双 CAN 分总线,这个问题一模一样”
    Experience.md § 执行器 + 时序建模 (line 6)
  • A joint frozen at the action clamp pays zero action_rate forever - penalize pre-clip saturation to make the cheat cost moneysaturation-cheating-zero-rate-cost
    Mechanism understoodwalkreward-shapingreward-shapingaction-rategate-battery

    Whenever actions are clipped and any smoothness/rate penalty exists, add a pre-clip saturation penalty so living at the clamp costs more than oscillating - and audit for frozen-at-clamp joints (action std ~0, |a| at exactly the clip value) as a standing acceptance row.

    Symptom

    With action_rate_l2 raised to -0.2, walk_v7's hip_pitch actions froze at exactly +/-1.000 (the clamp), reproduced bit-for-bit on hardware (splits frozen at +/-0.35 rad); the gait-shaping term joint_pos_ref collapsed to 0.026-0.035. The repo had died in the same trap once before (walk_v0: four joints pinned at +/-1.0).

    Context

    Mechanism: a joint pinned at the clamp has action-rate cost exactly zero and forever zero - under a strong smoothness tax, "push to the clamp and freeze" becomes the dominant optimum. Lowering the weight (-0.2 -> -0.1) only reduces temptation; the frozen state still costs nothing, so the structural fix adds action_saturation = sum(relu( |a_raw| - 0.9)) at weight -1.0, computed on the PRE-clip network output - post-clip, |a|=1.01 and |a|=3 punish identically and the out-of-range gradient dies (v0's old disease: mean |a| 1.71 soaked in saturation). Economics: freezing at |a|=1.0 now pays 0.1/joint/step (two hips = 40% of alive) vs ~0.0004/step for the healthy reference oscillation - the cheat flips from free to ~250x negative. Honest limits were recorded: A1 does not forbid freezing at 0.89 (the anti-freeze pressure must come from the oscillation demand of joint_pos_ref), and the alternative "rate on post-clip target" was rejected as 换汤不换药 - a pinned target also has zero rate.

    Change

    v8-A: add action_saturation (-1.0, thresh 0.9, pre-clip) AND halve action_rate_l2 (-0.2 -> -0.1, still 3.3x the v5 value); success criterion pre-declared (joint_pos_ref telemetry returns to v6 scale).

    Outcome

    Booked as the structural repair of the v7 freeze; also fixed a config hygiene trap discovered on the way - action_rate was assigned twice in __post_init__ (v5 comment line then v7 line), merged to one assignment "别再留两处赋值给下次审计埋雷".

    Mechanism

    Clipping creates a zero-gradient, zero-cost absorbing region in action space; any penalty on action derivatives makes that region strictly optimal once entered. Only a penalty on clamp proximity itself (measured pre-clip so depth of violation is visible) restores a slope out of the absorbing region.

    Applies when

    • joints sit at exactly the action clip with near-zero variance
    • raising a smoothness penalty degrades gait amplitude
    • shaped-oscillation terms collapse after a rate-weight increase
    “钉死在钳位的关节 action_rate 代价精确为零且永远为零;−0.2 之下"推到钳位冻起来"成了压倒性最优 … 本仓第二次栽在同一坑(walk_v0 死于四关节钉死 ±1.0)。回调权重(−0.2→−0.1)只降低诱惑不消除作弊 … 算在 clip 前的原始网络输出上 … 作弊收支从"白赚"变成"倒贴 ~250 倍"。”
    train/WALK_V8_SPEC.md § 1. 改动 A — 治饱和作弊
  • Push-test protocol - positive side first, fragile side spotted, axes aligned in the log, and cross-machine push counts stay qualitativepush-test-chirality-protocol
    Mechanism understoodomnireal-acceptancereal-acceptancegate-batteryprocess

    Order disturbance tests from the robust side to the fragile side with protection scaled to sim-measured asymmetry, align and log frame conventions before testing, and treat cross-domain disturbance counts as qualitative evidence only.

    Symptom

    Hand-push testing on hardware risked falls on a side sim had already flagged as fragile, and push counts invited apples-to-oranges comparison with sim numbers.

    Context

    Sim chirality was explicit: descendants were far more fragile in -y (fric-3000@kd1.2: +6 N*s survived 15/20 vs -6 N*s only 3-9/20) while the s1e control was perfectly symmetric (40/40). The protocol therefore: push the positive direction first, keep a spotter for the negative side; before any push, record which real-robot side corresponds to sim's +y in the log ("上机前对一次坐标"); and - citing the chaos lesson ("混沌课文") - real push results are used only as qualitative corroboration, never compared numerically with sim survival counts across machines.

    Change

    Push testing became a scripted, chirality-aware protocol with frame alignment as a logged precondition and an explicit epistemic limit on cross-domain count comparison.

    Outcome

    The fragile side was tested with protection informed by sim's quantified asymmetry; logs stayed interpretable because the frame correspondence was recorded before the first push.

    Mechanism

    Disturbance-response chirality is a real, quantifiable lineage property, so test order should follow measured fragility; and perturbation outcomes are chaotic in the details (divergent trajectories from tiny differences), so counts do not transfer across domains even when qualitative rankings do.

    Applies when

    • planning push/disturbance tests on hardware
    • sim shows directional asymmetry in disturbance survival
    • someone proposes comparing real push counts to sim counts
    “先正向后负向, 负向留人扶 —— sim 手性明确: 后代在负 y 向显著更脆 (fric-3000 @kd1.2: +6 N·s 15/20 vs −6 N·s 3~9/20), 而 s1e@0.8 两向 40/40 完全对称。上机前对一次坐标 … 跨机不做二值结论 (混沌课文): 真机推力只作定性对照, 不与 sim 计数对比。”
    train/REAL_RUN_S2.md § 3. 抗推 (可选, 人手推; 做则按此协议)
  • A suspended (no-load) test acquits or convicts the actuator before you blame authoritysuspended-test-isolates-actuator-authority
    Mechanism understoodomnireal-acceptancehardwarereal-acceptanceattribution

    Before attributing a failure to actuator authority, measure no-load tracking error and steady-state torque fraction; blame authority only if the task fails while the error grows with demanded force - and then fix gains or targets, not training.

    Symptom

    hip_roll sagged 0.21 rad on the ground and saturation questions loomed over the sidewalk plan - was the roll axis physically too weak (authority), or was something else limiting it?

    Context

    Before C4, the roll-authority question was settled by measurement triage: suspended test (--suspend, feet off ground) showed hip_roll tracking error 0.0008 rad - actuator acquitted; the entire 0.21 rad ground sag is load-induced. Steady-state torque was 25% of limit - 75% margin remains. Since sidewalk needs lateral force, not exact angles, authority was ruled "not a hard limit", with a pre-registered criterion for when it WOULD become one: sidewalk fails to track AND roll error keeps growing - then the fix is raising hip_roll kp or lowering the vy target, not more training.

    Change

    Hypothesis "roll authority insufficient" demoted from blocker to a monitored branch with an explicit trigger condition; C4 proceeded.

    Outcome

    Later open-loop probes confirmed the actuator could produce the behavior (sidewalk feed-forward ran at full amplitude, 5/5 survival), and the eventual C4 failure causes were measurement and reward, never authority.

    Mechanism

    Suspended vs loaded comparison separates the actuator's closed-loop competence from the load path: tiny no-load tracking error means the motor/controller is fine and any loaded deviation is statics (gravity / stiffness budget, kp trading error for force). Torque-fraction measurement then bounds how much force headroom actually remains.

    Applies when

    • suspecting an axis is "too weak" for a new skill
    • large position sag on a loaded joint
    • deciding between hardware fix, gain change, and more training
    “吊挂(--suspend)实测 hip_roll 跟踪误差 0.0008 rad → 执行器无罪,地面下垂 0.21 rad 全是负载所致;稳态占限扭 25% → 仍有 75% 扭矩余量。… 判据:若 C4 出现「侧走跟不动且 roll 误差继续变大」,那才是权限账 … 解法是提 hip_roll 的 kp 或降 vy 目标,不是硬训。”
    train/C_LADDER_RUN.md § 3d. roll 权限:已部分澄清,不是硬上限
  • Three same-shaped judging errors - task metrics (survival, tracking, displacement) cannot stand in for posture metricstask-metrics-vs-posture-metrics
    Replicatedomnireal-acceptancereal-acceptancegate-batteryattribution

    Keep validated posture-class rows (tilt, per-joint L/R asymmetry, temperature) in every acceptance battery alongside task rows; when operator feel contradicts the gates, suspect the metric class before the operator - and never build a new skill on what is actually an asymmetry defect.

    Symptom

    The C2 product judged "full pass" on task metrics (A800: turn-gap 7 pp, vx+0.30 19/20) felt WORSE in the operator's hands than the half-pass 700: A800 tilted up to 12.90 deg (700: 6.64), drifted left while standing, showed larger per-joint asymmetries, and ran its hip_rolls 5 degC hotter.

    Context

    The re-judgment catalogued three same-type metric errors in one campaign: (1) stand judged by SURVIVAL - missed 0.5-1.4 m wandering; (2) stand ranked by DISPLACEMENT - ordering was opposite to real feel (tilt ordering matched); (3) chirality judged by wz-tracking GAP - measured turning symmetry while the robot's actual disease was postural left/right asymmetry, "两个不同的东西,且结论相反". Common pattern named: "我一直用「任务指标」当判据,而真机手感对应的是「姿态 指标」… 任务类指标不能替代它". The fix was already in the data: the per-joint left/right asymmetry table (printed identically by sim2sim and deploy) agreed with hardware in direction on every row - "判据可用、有预测力,我只是没把它写进 PASS 条件". Shipping decision followed the posture read: product reverted to 700 ("又一次「买到 精度、卖掉别的」"), and the C4 root moved to 700 as well, with the sharpest line of the episode: A800's left-drift "like sidewalking" is probably its frontal-plane asymmetry defect, not a capability - "在缺陷上建能力是危险的".

    Change

    Two posture quantities with demonstrated real-robot predictive power promoted into every PASS battery: tilt-max median and per-joint left/right asymmetry (both sim-computable, deploy-homologous); motor-temperature readout added to session close-out.

    Outcome

    Deployment flipped to the posture-better checkpoint; the hip_roll temperature table (43-48 degC vs 25-28) confirmed the earlier 90%-of-heat account; the run-line acceptance battery inherited the posture rows from birth ("任务类替代不了姿态类").

    Mechanism

    Task metrics measure goal attainment under the evaluator's episode definition; posture metrics measure the body state trajectory that operators, motors, and downstream skills actually experience. The two can rank candidates oppositely because task success tolerates postural pathology - so a battery without posture rows is blind to exactly what hardware feel reports first.

    Applies when

    • hardware feel disagrees with a green acceptance table
    • choosing between checkpoints that split task vs posture metrics
    • selecting the root for a skill that resembles an existing defect
    “共同模式:我一直用「任务指标」(存活 / 跟踪率 / 位移)当判据,而真机手感对应的是「姿态指标」(倾角、逐关节左右不对称)。→ 验收判据里必须有姿态类指标,任务类指标不能替代它。… A800 的「左飘像 side walk」很可能 … 是它更大的额平面不对称的表现 —— 在缺陷上建能力是危险的。”
    train/README.md § C2 选点改判 (2026-08-09): 手性判据第三次选错指标
  • When hardware underperforms, audit deployment knobs before prescribing retrainingdeploy-knob-attribution-before-retraining
    Mechanism understoodomniattributionattributionreal-acceptanceprocess

    Before any "retrain it" decision, reproduce the symptom in sim under the exact deployment configuration; if the symptom follows the deployment knob rather than the checkpoint, fix the knob or randomize it in training - never top-up-train the skill.

    Symptom

    Real-robot feedback after the C4 deployment - "turning is weak" - with two retraining options on the table: top up turn training, or restart from the s1e root.

    Context

    The sim account showed the policy turned well (75-81% at pw1.0); the robot was deployed at power-scale 0.8. The 3-6 pp difference between C2 and C4 policies at the same power was noise; the 40-50 pp difference between power levels was the entire effect. Both proposed retraining paths would have burned budget on a non-existent training gap, and restarting from s1e would additionally have discarded the sidewalk skill that took four rungs and a coordinate-bug hunt to obtain.

    Change

    Decision: retrain nothing. (1) Try pw1.0 on hardware first - sim says net gain; (2) only if 1.0 is unacceptable (heat/feel), the correct training fix is power/torque randomization in the S2 plant line (one variable, fixes turn and backward together) - not skill top-up; (3) restart-from-root explicitly ranked worst.

    Outcome

    The "weakness" was fully explained by the deployment knob; the sim/real signatures matched the earlier power-derating law verbatim ("与 C2 时代 power 衰减主要伤非前进轴 逐字吻合").

    Mechanism

    The policy's competence is defined under its training plant; deployment knobs (power scale, teleop mapping, command bands) silently define a different plant. Attributing a deploy-plant effect to a training gap produces exactly the wrong fix - more training on the wrong variable.

    Applies when

    • real robot underperforms a skill that sim says is fine
    • proposals on the table include retraining or re-rooting
    • deployment uses any override the trainer never saw (power scale, remapped commands, different control rate)
    “正确的训练修法不是补训转向,而是训练时加 power/力矩随机化让策略在 0.8 下自己补偿 —— 单变量,属 S2 plant 线,一次同时修好转向与后退;从 s1e 重训是最差选项:丢掉四轮 + 一个指标 bug 才换来的侧走,而 C2 的转向本来就没问题。”
    train/C_LADDER_RUN.md § 3p. 三 处置顺序(回答「补训转向 还是 回 s1e 重训」:都不该)
  • Ungated phase shaping made standing 42x more expensive than stepping - and the stepping was cooking the hip motorsmoving-gate-42x-stand-tax
    Mechanism understoodomnireward-shapingreward-shapinghardwarereal-acceptanceprocess

    Gate every phase/clock-driven shaping term on the command that justifies motion, price the cmd=0 case explicitly during design - and when a reward flaw is sim-only, book it with trigger conditions instead of operating immediately on a working lineage.

    Symptom

    At cmd = 0 the sim policy never stood still - it stepped in place and crept 0.98 m per 20 s; on the real robot the same lineage's stepping made hip_roll motors run 20 degC hotter than every other joint (43-48 degC vs 25-28).

    Context

    The arithmetic closed it: the three phase-shaping terms (joint_pos_ref 1.6, feet_contact_number 1.2, feet_clearance_swing 1.6) are driven by the gait clock with NO command gating, so standing at cmd=0 forfeits 3.32/step of shaping while honest standing earns only 0.078 of tracking - stepping wins 42x. The heat chain: perpetual stepping = perpetual single support = one hip_roll stalled at ~4.3 N*m (25% of torque limit) carrying the torso's frontal-plane moment - two hip_rolls = 90% of whole-machine steady-state I2R; measured stand-vs-step comparison: total heat -71% when actually standing. The suspended test acquitted the actuator (0.21 rad sag -> 0.0008 rad in air) and mechanics vetoed the easy fix ("降 kp 救不了热" - equilibrium torque equals the external load regardless of kp). The fix (moving_gate: hard-gate the three shaping terms on |cmd| > eps) was designed - then DEFERRED by the user because the real robot at the time stood fine: "真机不表现该问题, 为真机不存在的病改奖励表不划算", with written trigger conditions (C-ladder stand row persistently failing, or real robot starting to step/drift at cmd=0) and the known hazard tag (this is exactly the reward-change class that triggers the B-arm signature). When the real robot later DID step and cook, the booked trigger fired and moving_gate moved from debt to to-do with its benefit re-priced: "停止空烧 hip_roll,稳态发热降 ~71%".

    Change

    moving_gate designed with the gate_by_cmd convention; deferral, triggers, and expected heat recovery all pre-registered instead of patching the reward for a then-sim-only symptom.

    Outcome

    The cmd=0 stepping went from mystery to closed arithmetic; the thermal measurement (hip_roll +20 degC) quantitatively confirmed the 90%-of-heat prediction; the reward change waited for real-world justification instead of spending a risky revision early.

    Mechanism

    Clock-driven shaping terms define a perpetual-motion bounty unless gated by command; the resulting idle gait is not a training bug but the table's optimum. Its cost surfaces on hardware as stall-torque heating set by statics (mass x lateral offset), which no gain change can remove - only removing the motion (gating) or widening the stance can.

    Applies when

    • the policy steps in place or creeps at zero command
    • specific joints run hot in idle behaviors
    • deciding when a known reward flaw justifies a risky mid-lineage fix
    “塑形合计 −3.32/步 … 净: 站定亏 42 倍 … 两颗 hip_roll 4.29 / 4.09 N·m(各占限扭 25%),占全机稳态 I²R 的 90% … 吊挂实测 hip_roll 跟踪误差 0.21 rad → 0.0008 rad … 降 kp 救不了热 … 真机不表现该问题, 为真机不存在的病改奖励表不划算”
    train/README.md § C1 FAIL 节 (cmd=0 的 sim/真机分歧记账) / C2 真机 A/B 三b 发热定性
  • When a contract default changes, old policies must run under a pinned legacy profile - a silent clock swap is out-of-distribution on hardwarelegacy-profile-pinning
    Mechanism understoodwalkreal-deploycontract-freezereal-acceptanceprocess

    Treat every trained policy as bound to the contract values of its training era: version the deployment profiles, pin old policies to their era's profile in every command, and never let a changed default silently apply to an old artifact.

    Symptom

    The walk profile's gait clock moved from 0.40 s to 0.50 s for new training, but versions v5-v9 were all trained at 0.40 s - running them under the updated default would silently feed a 25% slower phase clock to policies that never saw one.

    Context

    The re-test runbook hard-codes --policy-profile legacy_walk_040 into every command for the old versions, with the warning not to omit the flag: the mismatch is invisible (no error, no crash) but puts the policy out of distribution on hardware, where the same file had already documented that off-clock operation collapses gait quality.

    Change

    Deployment profiles versioned per training era; historical policies permanently associated with their era's profile; runbooks write the profile flag explicitly rather than relying on defaults.

    Outcome

    Old policies stayed runnable and comparable after the contract moved on; the silent-mismatch failure mode was closed by convention.

    Mechanism

    Changing a shared default rebinds every old artifact to a contract it was not trained under; unlike a schema break, a value change produces no error - only degraded, unexplainable behavior. Version-pinned profiles make the binding explicit and permanent.

    Applies when

    • changing any default in the deployment contract (clock, scales, gains) while old policies remain in use
    • writing runbooks that mix policy generations
    • a re-tested old policy behaves worse than its era's records
    “2026-08-02 起 walk profile 的时钟改为 0.50(WALK_V10_SPEC §3)。v5~v9 全是 0.40 训的,本文件所有命令已改带 --policy-profile legacy_walk_040 ——不要省掉这个 flag,否则是拿慢 25% 的相位时钟静默喂旧策略(分布外,真机危险)。”
    train/REAL_SWEEP_V5_V8.md § 1. 预检 ⚠️ 时钟改为 0.50
  • A pull-assist curriculum keyed to a global success share was satisfied by the categories that already worked and withdrew before prone learned anything - conditioning the criterion on prone took it from 2.5% to 98.7%curriculum-criterion-conditioned-on-lagging-category
    Mechanism understoodrecoverytraining-runcurriculumgate-battery

    Measure a curriculum's advancement criterion on the population the scaffold is meant to help; a global success share is met by whatever already works, and the help is withdrawn before the lagging case learns.

    Symptom

    Under the beta-anchored action, supine and side stood reliably while prone still sat (1.3%). A pull-assist curriculum added to help it was withdrawn completely within ~790 iterations and prone moved only to 2.5% (noise).

    Context

    The pull assist follows HoST: an upward force on the base, active only when the torso is within 30 deg of vertical, scaled by body weight (HoST's 200 N on G1 = 0.583 BW -> 56 N here, steps of 5.6 N, ten levels to zero); the product must pass with no assist. It had already taught sit-to-stand in V1. In V2.1 its advancement criterion was the standing-time share over all envs (threshold raised to 0.55 because the share was already ~0.53).

    Change

    V2.2: PullAssistForce with gate_category = "prone" - only envs whose first step classifies them as prone count toward the criterion - and the threshold back at 0.35. A feasibility signal was pre-registered: if prone's share stayed near zero under the full 56 N, return to the roll-over path instead of adding force.

    Outcome

    The prone-conditioned curriculum withdrew level by level only as prone itself passed: prone 98.7%, and the four-category gate passed for the first time on the line (98.6% overall, re-falls 0%, torque gate PASS).

    Mechanism

    A pooled success share is filled by the categories that already succeed (supine/side ~53%), so the scaffold is removed on their account before the lagging category has used it.

    Applies when

    • an assist, guide force or easier setting is withdrawn by a success threshold
    • one task category lags while the pooled metric looks healthy
    • a curriculum ran to completion without changing the lagging category
    “**教训:全局站立占比阈会被存量类别(supine/side ~53%)凑够,拉力在 prone 学会前就撤光了 —— metric 设计失误,不是拉力机制失效**(它在 V1 教会过 坐→站)。 … **V2.2(已启动)**:`PullAssistForce` 加 `gate_category="prone"` —— 达标判据 只统计 prone 类 env”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §34 V2.1 判决(2026-08-10)
  • Gate a new reward term by its command so all old modes score pointwise identicalgate-new-reward-terms-by-command
    Mechanism understoodomnireward-shapingreward-shapingprocess

    When a reward term must be added mid-lineage, gate it on the condition that defines the new task so every pre-existing situation scores exactly as before - and still watch for value-rescale pathologies inside the new mode.

    Symptom

    Adding a lateral tracking reward (track_lin_vel_y_exp) ungated would have paid 0-2.0 per step even in modes with cmd_vy = 0 (healthy gait sway of vy ~0.1 already earns 1.28), shifting the whole reward table by a large bias and rescaling the value function - no longer "just adding one mode".

    Context

    C4 was the C ladder's only true reward surgery. Single-variable discipline required that the change be invisible to every existing mode. The chosen construction: gate_by_cmd=True - the term pays only when |cmd_vy| > 0.02, so for all modes with cmd_vy == 0 the term is pointwise zero, i.e. the reward is pointwise identical to before the change. The same trick appeared earlier in C1: replacing the vy L2 tax with a command-error version that is "对 cmd_vy≡0 逐点同值" (pointwise equal when cmd_vy is 0), explicitly classified as not-a-reward-change.

    Change

    track_lin_vel_y_exp added with gate_by_cmd=True (weight +2.0, std 0.15); the residual acknowledged honestly - inside the side bucket the values DO change, so the rung still watched the known reward-reshuffle pathology signature (s1c B-arm: scatter -> half-recover -> collapse) as a stop criterion.

    Outcome

    Old modes provably unaffected (pointwise-equal argument); attribution for any change in old-skill metrics stayed clean through the C4 redo series.

    Mechanism

    PPO's critic normalizes to the reward scale it sees; an ungated additive term shifts returns in every state and re-scales advantages globally, entangling the new skill with all old ones. Command-gating confines the new term's support to the new mode's state distribution, making "pointwise identical elsewhere" a provable property rather than a hope.

    Applies when

    • adding a tracking/shaping term for a new command or skill to a lineage that must not regress
    • reward change proposed while other skills are still being gated
    • reviewing whether a config diff counts as a reward change
    “只在 |cmd_vy| > 0.02 时付。不门控的话它对 cmd_vy≡0 的老模式也给 0~2.0 分(健康摇摆 vy≈0.1 → 1.28),等于给整张奖励表加一个大偏置、值函数尺度全变 … 门控后老模式逐点得 0 = 与加项前逐点同值,单变量纪律成立。… 但 side 桶内的值确实变了 —— 这仍是奖励表改版,开级盯 s1c B 臂签名”
    train/C_LADDER_RUN.md § 3d. gate_by_cmd=True(重要)
  • Ideal PD is not enough - add a delay buffer and fit armature/friction/delay per jointactuator-delay-buffer-fitting
    Observed oncewalkactuator-modelingactuator-modelingplant-calibration

    Never ship ideal PD to hardware: add a measured delay (in control steps) and per-joint armature/friction fitted from step and sine responses, and treat remaining actuator mismatch as your standing largest sim2real residual.

    Symptom

    Standard ideal PD actuator model transfers poorly; sim assumes targets take effect instantly and joints reach arbitrary acceleration.

    Context

    A developer with a successful on-hardware Isaac Lab biped modified the actuator model in two ways and calibrated it against the real robot: step-response plus sine-sweep tests (positive step, negative step, sine tracking), overlaying sim curves on measured curves and hand-tuning.

    Change

    (1) Delay buffer: action targets take effect after a uniform 6 time-step delay on all joints; (2) acceleration limiting so the actuator cannot reach arbitrary acceleration; (3) per-joint fit of armature / friction / delay - different joints genuinely needed different values.

    Outcome

    Hip joints fit worst, knee best; the developer rated the result "not perfect, the best I could do" and still listed actuator-model improvement as next work - i.e. even the fitted model remained the dominant residual.

    Mechanism

    Real actuation is a lagged, bandwidth-limited system; a delay buffer and acceleration cap are the two cheapest structures that reproduce its phase and magnitude response. Per-joint differences come from differing load, wiring, and friction states, so a single global constant underfits.

    Applies when

    • actuator model in sim is ideal PD with no delay
    • step-response of real joint visibly lags or overshoots the sim's
    • budgeting which sim2real gap to attack first
    “标准 ideal PD actuator 不够用,他改了两处:延迟缓冲:目标不是立即生效,全部关节统一 6 个 time step 延迟 / 加速度曲线:执行器不能瞬间达到任意加速度 … 用 armature / friction / delay 三个参数逐关节拟合,标定方法是阶跃响应 + 正弦扫描 … 髋部关节偏差最大,膝关节最好。”
    Experience.md § 执行器建模 —— 最值得抄的一条 (lines 50-59)
  • Run acceptance under measured contact parameters - honest condim/torsional-friction flipped a false PASS into a real-matching FAILeval-plant-honesty-contact-params
    Mechanism understoodwalksim2sim-gatesim2simplant-calibrationgate-battery

    Pin the evaluation plant's contact model to measured values (contact dimension, torsional/rolling friction, mu) before trusting any gate that involves slip, impact, or drift - a gate can only fail a policy for physics its simulator contains.

    Symptom

    walk_v5 passed the old acceptance battery yet failed on the real robot (footfall force, drift, kicking) - the evaluation plant was flattering the policy.

    Context

    The battery was re-run under "honest contact parameters" - condim 4 (adding torsional contact) with measured torsional friction 0.035 - and v5 then FAILED exactly the rows corresponding to its real problems: heading 185 deg (limit 30), support-foot yaw slip 284 deg (limit 80), landing force 1.72x (limit 1.5x), suspended tilt 45.9 deg (limit 10). The slip physics depends on torsional friction, which the default contact model (condim 3) does not even simulate - a slip problem is invisible to an evaluator that cannot represent yaw friction at the foot. Term-sizing measurements for the new rewards were likewise taken under the same honest parameters (cmd 0.45, skipping the 5 s start transient).

    Change

    Acceptance harness pinned to condim 4 / torsion 0.035 (measured); verdicts issued under defaults declared non-citable for these rows.

    Outcome

    Sim acceptance verdicts began agreeing with hardware ("现在失败, 与真机一致"); the v6 fixes could be developed and validated against an evaluator that could actually see the disease.

    Mechanism

    An evaluator is a plant model too: contact dimensionality and friction values decide which failure modes exist in the simulation at all. Evaluating under default contact parameters tests the policy in a world where its real failure is physically impossible, producing structurally false PASSes.

    Applies when

    • sim acceptance passes policies that fail on hardware
    • slip/drift/impact gates run under default simulator contact settings
    • setting up a cross-simulator evaluation harness
    “accept_v2.py 已加三条判据, walk_v5 在诚实的接触参数下(--condim 4 --torsion 0.035)现在失败, 与真机一致:直行 15s 航向累计 <30° | 185° ✗ … 落脚力峰值 <1.5× 体重 | 1.72× ✗”
    train/WALK_V6_MINIMAL.md § 5. 验收
  • After installing the measured plant, re-run all generations paired on old and new plant - identity metrics carry verdicts, physics metrics re-baselineplant-swap-invariants-vs-shifts
    Mechanism understoodwalksim2sim-gatesim2simplant-calibrationgate-batteryattribution

    Treat every plant upgrade as an era boundary: re-run the retained policy set paired (same seeds/flags) on both plants, carry forward only verdicts whose metrics proved plant-invariant, re-baseline the rest - and mine the systematic shifts as measurements of the old plant's biases.

    Symptom

    With the plant finally fully measured (weighed masses 9.792 kg, bench-identified armature, in-situ friction), no historical sim number was comparable to new runs - "历史 sim 数字跨纪元不可比" - and it was unknown which historical verdicts still held.

    Context

    The era-2c cross-test ran all seven walk generations on the complete plant under one harness (14/14 survived), then re-ran the same 14 configurations on the OLD plant retrieved from git, same flags and seeds, with a self-check (one historical record reproduced digit-for-digit). The split was clean. Policy-identity metrics moved essentially zero across the plant swap - dominant frequency (v7's period-doubling 1.20 -> 1.20), knee amplitude (9.9 -> 10.0), foot distance (+/-2 mm), saturation (100 -> 100, 0 -> 0) - so all seven cross-generation verdicts (freeze signature, saturation-line closure, knee-collapse location, slip-penalty accounting, N2 lineage, v6's balance, v11's triad) were re-confirmed on the honest plant. Plant-physics metrics shifted systematically with ordering preserved: slip down 10-25% (measured friction makes ground-twisting costlier), landing vertical velocity down 15-50% ("旧 plant 高估落地 凶度"第二次独立证实), landing force mixed (mass up 2.2% vs friction braking the swing - two effects fighting). Exactly ONE behavior-level change: v8's low-speed period-doubling vanished (1.30 -> 2.50, lift normalizing) - confirming it had been machine-dependent bifurcation-edge behavior that armature+friction push off the knife edge, while v7's period-doubling stood untouched: saturation-freeze-driven, a policy property, not a numerical accident.

    Change

    Era re-baselining protocol: after any plant upgrade, one paired same-seed sweep of all retained generations on old and new plant; verdicts keyed to identity metrics carry over, thresholds re-read against the new-plant table, and the differences themselves become plant-physics findings.

    Outcome

    Seven verdicts survived with evidence rather than assumption; two causes of the period-doubling family were separated with plant-side proof; and the sim's landing-violence overestimate was independently confirmed a second time.

    Mechanism

    A policy's structural properties (frequencies, amplitudes, frozen joints) are functions of its weights and survive plant changes; contact-mediated quantities are joint properties of policy and plant and shift when the plant becomes honest. Pairing seeds across plants isolates the plant's contribution exactly, so the sweep both validates history and measures what the old plant had been lying about.

    Applies when

    • installing measured masses/armature/friction into the sim
    • historical thresholds are cited across a plant change
    • a hardware-only behavior might be bifurcation-edge sensitivity
    “策略身份指标逐位不动:主频(v7 1.20→1.20)、膝摆 … 这些是策略属性,plant 换代携带无损,历史定论因此全部成立。… 唯一行为级变化:v8@0.15 的倍周期消失(主频 1.30→2.50…)——印证当时"分岔边缘、机器相关"的判定:armature+摩擦把 v8 推离刀锋;v7 的倍周期纹丝不动(1.20→1.20),它是饱和冻结驱动的深层属性,不是数值巧合。”
    train/README.md § 纪元 2c 全代同机横测 (2026-08-04): 完全体 plant 上历史结论全部存活
  • Fine-tuning through a reward-table change was falsified (scatter, half-recover, collapse) - continuation training is legal only with the reward frozenfine-tune-reward-change-falsified
    Replicatedomnicurriculumcurriculumfork-selectionreward-shapingprocess

    Never fine-tune through a reward-table change - retrain from zero; reserve checkpoint continuation for frozen-reward plant/DR widening, reset noise_std when branching, and watch for the scatter/half-recover/collapse signature as the abort trigger.

    Symptom

    The s1c A/B experiment: arm B fine-tuned from an existing checkpoint under the revised reward (same contract, same network, changed reward + small DR) and failed with a characteristic signature - scatter, half-recover, fall back ("打散→半恢复→摔回"); arm A trained from zero under the same config won decisively (full shaping lifted swing to 21.6 mm within 500 iters; shipped at 5500).

    Context

    Verdict recorded: "从零 + 强塑形是本机唯一验证过的发育路径" (from-zero plus strong shaping is this machine's only validated development path). The signature became a standing stop criterion in every later rung that touched a reward ("s1c B 臂签名,出现即停"). Crucially the boundary of the law was drawn explicitly when S2 continuation training was proposed: "当年证伪的是「奖励表中途改版的 fine-tune」… S2 续训奖励表全程冻结、只逐级加宽 plant/DR,属另一类" - continuing a checkpoint with the reward FROZEN while widening plant/DR one rung at a time is a different class and was allowed (and then worked, powering the whole S2/C lineage) - with the honest fallback that if frozen-reward continuation ever collapses, that rung retrains from zero and the doctrine gets re-examined with data. Fine-tune arms also need mechanical care: reset the checkpoint's collapsed noise_std (terminal 0.033 "会杀死探索") and account for iteration counters re-zeroing (curriculum gates fire immediately).

    Change

    Reward changes and lineage continuation permanently separated: reward revisions -> from-zero retrain; plant/DR widening -> frozen-reward continuation with per-rung gates; the B-arm signature promoted to a universal tripwire.

    Outcome

    No later reward revision was attempted by fine-tune; frozen-reward continuation carried S2 (PD/COM/friction rungs) and the C command ladder successfully from the s1e root.

    Mechanism

    A trained policy sits in an optimum of its reward's geometry; changing the reward moves the optimum but leaves the policy's exploration noise near-zero and its value function calibrated to the old returns - it disassembles the old solution faster than it can assemble the new one. Widening DR under a frozen reward instead keeps the optimum's identity and asks only for local robustification.

    Applies when

    • proposing to fine-tune an existing policy under a revised reward
    • planning a robustification ladder from a validated checkpoint
    • a continued run scatters then partially recovers then collapses
    “B 臂 fine-tune 证伪(打散→半恢复→摔回——从零 + 强塑形是本机唯一验证过的发育路径)。… 当年证伪的是「奖励表中途改版的 fine-tune」(B 臂,塑形突变致终盘摔回);S2 续训奖励表全程冻结、只逐级加宽 plant/DR,属另一类;若 s2_lag1 续训本身塌方,回退方案 = 该级从零重训,续训教义再议(拿数据说话)。”
    train/OMNI_V0_SPEC.md § 3. S1.3 / 4. 与 s1c fine-tune 证伪的关系
  • Add a termination that makes the degenerate strategy fatal - no height cut-off meant crouch-shuffling could live forevertermination-closes-degenerate-basin
    Observed oncewalkreward-shapingterminationreward-shaping

    For each known degenerate strategy, check whether the termination set makes it fatal; if the robot can live indefinitely inside the degenerate posture, add a termination just past the intended operating envelope rather than escalating penalties.

    Symptom

    Crouched foot-dragging survived indefinitely because the termination set contained only bad_orientation (40 deg) and base contact - there was no height termination at all, so a deep squat was a viable long-term strategy.

    Context

    The hypothesis audit found the missing termination (hypothesis 7); the cross-check against published configs found the field practice: Booster terminates at 0.45 m (38% of body height) and the research warning is that the termination height must not be so low that crouching survives it. The proposed value: 0.32 m, just below the walk crouch base height 0.3739 - a deep squat terminates immediately, "断掉蹲着蹭的活路" (cutting off the crouch-shuffle's livelihood).

    Change

    Add height termination at 0.32 m as a second-priority item of the walk fix package, alongside restoring base_height_l2 to -10.

    Outcome

    Entered the v5/v6 fix package under which the crouch-shuffle optimum disappeared (34 mm clearance, 87% tracking by v6).

    Mechanism

    Termination conditions define which strategies exist at all: a reward penalty prices a behavior, but a termination deletes its future returns entirely. Degenerate basins that are merely penalized can remain optimal under enough tracking pressure; a termination placed between the degenerate posture and the intended one makes the basin unreachable as a steady state.

    Applies when

    • a degenerate but stable behavior persists across reward tunings
    • auditing termination conditions for a locomotion task
    • a policy exploits the gap between penalized and terminated states
    “加终止高度:研究第 6 条"终止高度不能低到让蹲着也能活"。我们完全没有高度终止。建议 0.32 m(略低于 walk 蹲姿基座高 0.3739,深蹲即终止)。… 加终止高度 0.32 m(深蹲即终止,断掉蹲着蹭的活路)”
    train/WALK_DIAGNOSIS.md § 修正 ④ / 最终改动清单 第二优先
  • The +/-50 mm lateral COM randomization meant to spread the legs coincided with legs pulling IN - rolled back per its own pre-registered contractcom-dr-rollback-on-symptom
    Observed oncewalkdr-tuningdomain-randomizationattributiongate-battery

    When adopting a DR value that covers no local measurement, write its intent and rollback trigger into the config at adoption time; roll it back as the control arm the moment the symptom contradicts the intent, and promote the symptom's metric into the acceptance battery.

    Symptom

    After v7 adopted the reference developer's oversized lateral COM randomization (+/-50 mm) explicitly to force leg spread, the real robot's legs narrowed instead - lateral mean 154 mm / closest 107 mm in sim (nominal 214.5), narrower still on hardware with occasional leg contact.

    Context

    The rollback was clean because the adoption had been honest: the robot.yaml comment recorded the intent AND that the +/-50 value covered no local measurement (only a 16/7 mm measured offset existed; even the prior widening to +/-20 was subjective), plus the reference's own reported side effect (base sway) and the note "这一项要单独跑、 单独归因". When the opposite symptom appeared, v8 returned y to +/-20 mm as the control arm ("要么没起作用、要么帮了倒忙 … 按约定退回做 对照"), kept x/z untouched (a noise-level difference not worth another variable), and named the second suspect: the landing penalty itself, via the reference's own three-link chain (landing penalty -> stance narrows -> spacing penalty needed). A gate lesson was booked in the same table: v7's sim numbers had ALREADY crossed the line (154/107 vs v5's 182/147) - "这个指标本可拦下 v7" - so foot-distance became a standing acceptance row (min >120 mm, zero leg-leg contacts).

    Change

    base_com_offset_m y: 0.050 -> 0.020 (x/z kept), regenerated through the export tool rather than hand-editing derived files; foot-distance acceptance row added.

    Outcome

    A borrowed DR lever with no local measurement basis was retired the moment its symptom contradicted its purpose, at single-variable cost; the metric that would have caught it pre-hardware entered the gate.

    Mechanism

    DR ranges shape behavior through the policy's robustness strategy, which is jointly determined with every reward term; a lever that forces stance width on one robot can be dominated by a stronger narrowing pressure (landing softness) on another. Levers adopted without local measurement must carry their own rollback trigger, because there is no nominal to argue from when they misbehave.

    Conflicts

    Causality is not fully closed in the source: the narrowing may come from the landing penalty rather than the COM lever ("腿距的第二嫌疑人是 ④ 本身"); the rollback is the pre-agreed control experiment, not a verdict that the lever caused the narrowing.

    Applies when

    • importing DR ranges or behavioral-forcing randomizations from references
    • a DR lever's observed effect contradicts its documented purpose
    • a sim metric existed that would have caught a shipped regression
    “⑥ 的本意 … 是逼策略把脚分开;真机结果是脚向内收且偶发相碰——要么没起作用、要么帮了倒忙。… 注释当时就写了"这一项要单独跑、单独归因"。现在症状出现了,按约定退回做对照。”
    train/WALK_V8_SPEC.md § 3. 改动 C — 质心随机化退回(撤销 v7-⑥ 的 y 项)
  • Isaac splits Coulomb friction into static and dynamic columns - wiring only static means zero loss during motion, silently discarding the identified valuesim-api-friction-columns
    Mechanism understoodinfraplant-calibrationplant-calibrationactuator-modelingdomain-randomization

    When installing identified actuator parameters, map each measured quantity to the simulator's exact API column for the operative regime (dynamic for moving loss, viscous for damping), verify per joint after landing, and audit how randomization intervals fall on each column's nominal.

    Symptom

    The hardware-identified Coulomb friction (tau_c) was about to be installed into the trainer through the friction= field alone - which in Isaac 5 populates only STATIC friction, so during motion the joints would lose no torque at all: "只给 static 则运动中不损耗, 辨识的 τ_c 走路时等于没接" (the identified tau_c would effectively not be connected while walking).

    Context

    The v12 integration wired all three columns deliberately: armature= and friction= from the 2026-08-04 hardware identification, PLUS dynamic_friction= (Isaac 5 splits Coulomb into static/dynamic; the moving-loss column is dynamic) and viscous_friction= (= the measured joint damping 0.02, aligned to MJCF's damping). Each value was re-checked per joint after landing. A DR interaction was audited and booked rather than hidden: randomize_joint_parameters jitters ALL friction columns with ONE interval - the [-0.05, +0.10] band was calibrated against the Coulomb nominal, and landing on the viscous nominal 0.02 it becomes [0, 0.12], "偏宽但保守" (wide but conservative), accepted with the note that pre-viscous behavior was already [0, 0.10] on a base of 0.

    Change

    Measured actuator parameters installed across all applicable API columns (armature, static, dynamic, viscous), with the DR side effect on shared randomization intervals audited and recorded.

    Outcome

    The first generation where the identified plant actually acts during motion in the trainer; the silent-column failure mode documented before it cost a training run.

    Mechanism

    Physics engines decompose "friction" differently (single coefficient vs static/dynamic/viscous columns); a measured parameter is only installed when it reaches the column the solver reads in the regime that matters (motion, not stiction). Randomizers that share one interval across columns rescale the band by each column's nominal - a hidden unit change.

    Applies when

    • installing identified friction/armature into any trainer
    • porting plant parameters between simulators or engine versions
    • joint losses in sim do not match bench measurements during motion
    “并额外传 dynamic_friction=(Isaac 5 把库仑拆 static/dynamic 两列,只给 static 则运动中不损耗,辨识的 τ_c 走路时等于没接)与 viscous_friction=(= joint_damping 0.02,对齐 MJCF damping)。… randomize_joint_parameters 用同一个 friction 区间抖三列, [-0.05,+0.10] 是按库仑标称标的,落到粘滞标称 0.02 上成了 [0,0.12]”
    train/WALK_V12_SPEC.md § 7. 核查单 (Isaac 接 V.ACTUATORS 新字段)
  • The deploy-side walk/recovery switch - into recovery at tilt > 65 deg held 0.3 s, back at tilt < 15 deg with angular rate < 1 rad/s and straight knees held 1 s, a 15 s timeout, last action cleared both ways - and the handoff steps the design required are only partly implementedwalk-recovery-fsm-handoff
    Observed oncerecoveryreal-deployreal-acceptancecontract-freezeprocess

    Specify a deploy-time controller switch as hysteretic, time-filtered predicates the robot can measure (proxy what it cannot, e.g. straight knees for height), a timeout that ends in a safe stop, and a complete handoff (history, clock, last action, command ramp) - then test that the code performs every handoff step, because the design document is not the implementation.

    Symptom

    With a recovery policy and a locomotion policy as separate networks, the robot needs a switch: when is it "fallen", when is it "up", and what state must be reset so the next policy does not act on the previous one's history.

    Context

    The 08-09 design: enter recovery when fallen (tilt > 55 deg or height < 0.60 x 0.384 m) for 150 ms, leave for a stand-hold when upright (tilt < 12 deg, height > 0.85 x 0.384 m, feet steady, |omega| < 0.8) for 400 ms - wide entry, strict exit, hysteresis - then a mandatory handoff trio before walking resumes (reset the walking policy's observation history, restart its phase clock at 0, clear its previous action and latency buffer) and a command ramp instead of a jump. The 08-14 implementation in deploy_policy (--recovery-policy): each policy under its own manifest contract (walk: nominal + scale; recovery: beta-anchored), the same rl_default gains, the tilt cutoff disabled; RECOVERY at tilt > 65 deg for 0.3 s; LOCO again at tilt < 15 deg and |omega| < 1 and knees straight (< 0.35 rad) for 1 s - the robot's computer has no height estimate, and straight knees stand in for height so a V3.0-style upright kneel cannot pass as standing; RECOVERY longer than 15 s ends in a safe stop; last_action cleared on both switches; command forced to 0 during RECOVERY; power scaling applies to LOCO only.

    Change

    An open account was written down with it: the recovery end state (0.355 m stance, hip yaw -/+27 deg) is outside the walking policy's training start distribution, so the first test must use stand / zero command as LOCO, and the long-term fix is to widen the walking policy's initial states rather than bend recovery's stance to suit walking.

    Outcome

    The spec records only a successful compile (py_compile); the hang test was left to be done on site. The operator runbook carries the three-step procedure (hang with stand as LOCO, mat and push, then omni walk as LOCO) but no outcome.

    Mechanism

    Two policies trained separately each assume their own history, clock and last action; a switch that carries any of them across feeds the next policy a state it never saw - the design called this "the walking policy seeing a ghost history".

    Conflicts

    The 08-09 design requires resetting history, clock and previous action plus a command ramp; the 08-14 implementation records last_action clearing and a zero command during recovery; the one-leg spec of 2026-09-14 lists the handoff hygiene as specified but not implemented - reset_history() is called by nothing (a 215-dim policy would carry four frames of pre-fall history), the phase clock is not zeroed, and there is no command ramp back to LOCO. No hardware run of the FSM is recorded in either source.

    Applies when

    • switching between separately trained policies on hardware
    • a policy with history or phase observations is re-enabled mid-run
    • the robot lacks a sensor the switching criterion was designed around
    “**判据**:进 RECOVERY = 倾角 >65°(`--fall-tilt-deg`)持续 0.3 s;回 LOCO = §5 真机可测子集:倾角 <15° ∧ |ω|<1 ∧ **膝直 <0.35 rad(NX 无高度观测, 高度门用膝直代理 —— 防 V3.0 型"跪坐但直立"误判)** 持续 1 s (`--recover-hold`);RECOVERY 单次 >15 s(`--recovery-max-time`)安全停。 … **切换卫生**:两向切换 last_action 清零;RECOVERY 态 cmd 强制 0”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §50 FSM 双策略调度(2026-08-14,用户令):deploy_policy --recovery-policy
  • Push DR helped one lineage and hurt another at the same dose - robustness budget is conserved and gets borrowed, not createdpush-dr-conditional-budget-conservation
    Replicatedomnidr-tuningdomain-randomizationcurriculumattribution

    Before opening a disturbance-DR rung, measure whether the untrained policy already meets the spec; if training it anyway, expect the benefit to be conditional on the lineage's existing DR load, grade the intensity, and audit retained margins - budget spent elsewhere will be borrowed back.

    Symptom

    The push rung's outcome flipped with the lineage: direct +/-0.6 m/s push failed outright on first attempt (base walking collapsed - kd1.2 scan 0/3 from iter 3300, sim2sim self-falls with pushes OFF - no PASS point existed); staged +/-0.3 then gave the narrow-kd single-working-point lineage real gains (push survival 1/5 -> 4/5) while the SAME dose made the dual-working-point balanced-band lineage WORSE (20-seed survival 18 -> 12/20 plus across-the-board push regression).

    Context

    The four-ladder verdict ("四梯定案", s2e/s2f at both intensities) named the pattern: "push DR 收益条件性" - the benefit is conditional on how much robustness budget the lineage has already spent. The law candidate: "DR 总预算守恒, 平衡带鲁棒性从抗扰余量借" - total DR budget is conserved; a lineage already covering a wide plant band pays for push tolerance out of its disturbance margin. Both S2 ladders therefore closed at the friction rung, with the decisive numerator: untrained push tolerance already met the 4-6 N*s requirement, so the rung was not needed at all ("⑥ push 不训(收益条件性,免训 ±0.6 已达 标)"). The same accounting later justified the C-before-S2 ordering ("push/μ 两轮已实证 DR 预算有限且会被重分配") and trimmed the second S2 pass to three rungs.

    Change

    Push removed from the standing ladder; graded intensity retained as the method IF a lineage ever needs push training; "does the untrained policy already meet the disturbance spec" instituted as the first check before opening any disturbance rung.

    Outcome

    Two rungs (push, ground mu) deleted from the second S2 pass on measured grounds; the ladder's real yield was re-stated honestly as precision, not robustness (speed gate 0 -> 20/20, zero-command drift 0.98 -> 0.06 m, but push 159 -> 125/160).

    Mechanism

    A fixed-capacity policy allocates representation and margin across the training distribution; adding a disturbance axis to a lineage that already spans a wide plant family forces reallocation - the new tolerance is bought with existing margins. Lineages with narrow plant coverage have free budget, so the identical DR dose lands as gain. Benefit is a property of (dose x lineage state), never of the dose alone.

    Applies when

    • proposing push/perturbation training on a hardened lineage
    • the same DR rung helped one lineage and hurt another
    • accounting where a ladder's robustness gains actually came from
    “push DR 收益条件性 —— s2e⑥a (单工作点血统 kd 窄带) ±0.3 得抗推 1/5→4/5; s2f⑥ (双工作点平衡带血统) 同档反而 20-seed 存活 18→12/20 且抗推全面倒退。规律候选: DR 总预算守恒, 平衡带鲁棒性从抗扰余量借。两阶梯均以 ⑤ 摩擦级收官 … 抗推 4~6 N·s 免训已达标。”
    train/OMNI_V0_SPEC.md § 4. ⑥ push 四梯定案 (2026-08-07)
  • A stand gate judged by survival passes a robot that wanders a meter - judge posture insteadstand-gate-posture-not-survival
    Mechanism understoodomnireal-acceptancegate-batteryreal-acceptance

    For every gate, ask what behavior the metric is a proxy for and validate its ordering against real observations; replace metrics whose ordering disagrees with reality, and demote them explicitly rather than silently.

    Symptom

    Real robot "standing" drifted 0.5-1.4 m across the floor while the sim stand gate scored a clean 20/20 - because the gate's metric was episode survival, which wandering does not violate.

    Context

    C2's stand condition was originally written as "survival regression <=2/20". A real-robot counter-example on 2026-08-08 forced the re-judgment: wandering robots survive. Cross-checking candidate sim metrics against real-robot feel showed max-tilt median ordering agreed with hands-on ranking, while displacement ordering was actually OPPOSITE to real impressions - so displacement was demoted to a reference quantity, not a gate.

    Change

    Stand PASS criterion rewritten from survival to posture: "stand tilt max median <= root baseline +1.5 deg"; displacement kept only as reference. Applied to all subsequent rungs (C4 and redo levels inherit it).

    Outcome

    Later rungs gated stand on tilt (e.g. C4 product: 7.7 deg vs parent 7.1 deg, +0.6 deg PASS); the wandering failure mode became visible to the battery instead of hidden by survival.

    Mechanism

    A gate metric is a proxy for an intended behavior; survival is a proxy for "did not fall", not "stood still". Metric choice must be validated against ground truth (real-robot feel/measurement), and a proxy whose ordering disagrees with reality on real data is worse than no metric - it steers selection backwards.

    Applies when

    • writing PASS conditions for stand/idle/hold behaviors
    • a gate passes policies that visibly misbehave on hardware
    • choosing between candidate metrics for an acceptance battery
    “stand 必须用位姿判,不能用存活判(2026-08-08 真机反证改判):站着乱走 0.5~1.4 m 时存活照样 20/20 —— C2 的 stand 条件原写「存活退化 ≤2/20」,选错了指标。改为: stand 倾角 max 中位 ≤ 根基线 +1.5°(倾角与真机手感排序一致;位移排序与真机相反,降为参考量)”
    train/C_LADDER_RUN.md § 3b. PASS 条件 ⚠️ stand 必须用位姿判
  • Slowing the gait clock at deployment is out-of-distribution and backfires - lower the commanded speed instead, or train the knobcycle-time-override-is-ood
    Mechanism understoodwalkreal-deployreal-acceptanceattributioncontract-freeze

    Any deployment override must correspond to a dimension the policy was trained to handle; to make a parameter field-adjustable, randomize it in training and observe it - otherwise use the levers inside the trained envelope (commands) and leave the knob alone.

    Symptom

    Real-robot feedback "walks very fast and unstable" suggested slowing the gait; a deploy-side --cycle-time override existed, making "just slow the clock" a one-flag temptation.

    Context

    A sim sweep of the override on walk_v6 @cmd 0.3 showed monotone degradation away from the trained 0.40 s cycle: at 0.50 s tilt jumped 7.9 -> 13.2 deg and landing force 1.52x -> 2.24x; at 0.80 s (half speed) clearance collapsed to 3 mm - dragging again - with 20 deg tilt. Meanwhile the legitimate lever, lowering the commanded speed with the clock untouched, improved everything monotonically: cmd 0.1 gave 104% tracking, 6.7 deg tilt, minimum slip - the most stable operating point. The file distinguishes the two "slows" explicitly: lower command = smaller steps at the same 2.5 Hz rhythm; a slower rhythm itself requires retraining - randomize cycle_time (e.g. 0.40-0.65 s) during training and expose it as an observation, and only then does --cycle-time become a field-adjustable knob.

    Change

    Deployment guidance: never ship a cycle-time override the policy was not trained under; respond to "too fast/unstable" with lower commands; schedule clock variability as a training-time (contract-level) change if a field knob is wanted.

    Outcome

    The sweep quantified the trap before hardware paid for it (dragging and 2.2x landing force at slowed clocks); cmd 0.1 documented as the stable demo point.

    Mechanism

    The policy is a function fitted around the training distribution; a deploy-side override moves an input (phase rate) to values never seen, so behavior degrades unpredictably - the knob LOOKS like a capability because it exists in the code, but capability lives in the training distribution, not the interface.

    Applies when

    • a deploy tool exposes overrides (clock, scale, gains) beyond the training distribution
    • hardware feels "too fast/aggressive" and a quick knob exists
    • deciding between a deploy-side tweak and a retrain
    “0.80s | 1.25Hz | 0.165 | 3mm(拖地) | 20.0° … 慢一半直接崩 … 策略按 0.40 训练, 别的周期属分布外。… 降指令速度才是有效杠杆 … cmd 0.1 是最稳的工作点。… 要节奏本身变慢必须重训 —— 训练期把 cycle_time 随机化(如 0.40~0.65s)并作为观测的一维, 部署时 --cycle-time 就成了现场可调的旋钮。”
    train/WALK_DIAGNOSIS.md § 2026-08-01 追加: 调慢步态时钟(--cycle-time)在仿真里是反效果
  • When training fails repeatedly, inject the target behavior open-loop - stop tuning rewards for an unverified behavioropen-loop-probe-before-reward-tuning
    Mechanism understoodomniattributionattributionprocesscurriculum

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop on the real plant/sim first, and only resume training once you hold a measured, safe, sign-verified target trajectory.

    Symptom

    Three sidewalk training rounds failed; hypotheses multiplied (exploration failure / wrong reference waveform / insufficient authority) with no way to pick between them by running more training.

    Context

    Instead of a fourth reward guess, the team wrote probe_side_ref.py: the candidate reference is injected open-loop on top of a frozen policy's output (bypassing PPO entirely), directly measuring "what happens if the robot literally does this waveform" - separating all three hypotheses in one experiment (5 seeds x 8 s per condition, several waveform families and gains). The probe immediately eliminated the authority hypothesis (full-amplitude execution, 5/5 survival) and localized the problem to the waveform/measurement side. The closing principle was written down after the saga: without a verified target behavior, tuning rewards is "在黑暗里试钥匙" (trying keys in the dark).

    Change

    Standing method: before opening another training rung on a failing skill, build an open-loop (or task-space) generator of the intended behavior, measure whether the physical system can express it and what it looks like - then train toward a verified, quantified target.

    Outcome

    The probe chain produced the verified waveform (reversed-sign triangle, half gain), quantified safe amplitude (tilt 8.2 deg at band top, foot distance clear of the wall), exposed the metric bug when probe and training disagreed, and ultimately supplied the feed-forward that made C4 pass in +100 iters.

    Mechanism

    Training couples exploration, reward design, and feasibility into one opaque outcome; open-loop injection cuts the loop and tests feasibility and waveform alone. A behavior demonstrated open-loop converts the remaining failure into a pure credit-assignment/reward question - and its measured trajectory becomes the reference itself.

    Applies when

    • repeated training failures on one skill with multiple live hypotheses
    • uncertainty whether the platform can physically express the behavior
    • a reference trajectory's shape/sign/amplitude is guessed, not measured
    “三轮 FAIL 之后不再猜,写 train/probe_side_ref.py 把参考开环注入到策略输出之上(绕过 PPO),直接量「照这个波形做会怎样」,一次分开三个假说:甲 探索 / 乙 波形 / 丙 权限。”
    train/C_LADDER_RUN.md § 3f. C4 真因定谳(开环探针) / 3k. 建议的下一步
  • Guessed joint friction was 2.5x low and damping 5x high - measure, then DR around nominalfriction-measured-not-guessed
    Mechanism understoodwalkplant-calibrationplant-calibrationdomain-randomization

    Measure frictionloss and damping separately (they need different rigs), put the measured value at DR center, and express DR as an additive band around that nominal - a DR range around a guessed value can exclude the real robot entirely.

    Symptom

    Old MJCF friction values were invented, not measured; when finally measured, every guessed value was wrong by a large factor in some direction.

    Context

    Joint friction split into Coulomb (frictionloss, tau_c) and viscous (damping b). Measured with the robot hung from a crane (吊机测) while armature was measured no-load; the two measurements are deliberately separated. Old MJCF: frictionloss 0.05, damping 0.1, DR joint_friction range [0, 0.1] "凭空拍的" (made up out of thin air).

    Change

    Replace guessed values with measured ones - frictionloss: RS06 0.15 / RS02 0.12 / RS00 0.13 N*m (old 0.05, i.e. 2.5x too low); damping: 0.02 N*m*s/rad on all three motor types (old 0.1, i.e. 5x too high). DR reshaped from an absolute made-up range [0, 0.1] to an additive band around measured nominal: joint_friction_add [-0.05, +0.10].

    Outcome

    "摩擦定稿(与 armature 一起, plant 参数第一次全部来自实测)" - friction frozen as part of the first fully-measured plant; DR now brackets a measured truth instead of spanning an invented interval.

    Mechanism

    Coulomb friction and viscous damping have opposite behavioral signatures (constant-torque threshold vs velocity-proportional drag); guessing both wrong in opposite directions gives a plant that is simultaneously too easy to start moving and too hard to move fast. DR centered on a wrong nominal makes the policy robust to a family of plants that does not contain the real one.

    Applies when

    • plant friction/damping values have no measurement provenance
    • DR ranges are absolute intervals rather than bands around a nominal
    • policy is over- or under-damped on hardware relative to sim
    “测关节摩擦。吊机测,而armature应该空机测试。… frictionloss τ_c (N·m) │ 0.15 │ 0.12 │ 0.13 │ 0.05(低 2.5×) … damping b (N·m·s/rad) │ 0.02 │ 0.02 │ 0.02 │ 0.1(高 5×) … DR │ joint_friction_add: [−0.05, +0.10] 叠标称 │ 旧 [0, 0.1] 凭空拍的”
    Experience.md § 摩擦定稿表 (lines 12-25)
  • Tightening the bridge's rate limiter under an unchanged policy cut torque peaks 30-50% and made other things worse - the policy cannot see the limiter, keeps commanding and winds up; a deploy-side limiter is a safety net, not a curedeploy-rate-limiter-windup
    Mechanism understoodrecoverysim2sim-gateactuator-modelingsim2simreal-acceptance

    A rate or torque limiter added at deployment lowers peaks but the policy still commands as if unconstrained (saturation, windup, new contacts); use it as a safety net mirrored in evaluation, and put the constraint where the policy can learn around it.

    Symptom

    After the violent first real-robot get-up, the cheapest candidate fix was to tighten the bridge's slew (rate) limit for the recovery policy without retraining.

    Context

    Probe on R3.1 in MuJoCo (5 categories x 3 seeds, mu 1.0), monkeypatching the limiter with no repository change: TIGHT = RS06 4.0 / RS02 3.0 / RS00 2.0 rad/s (about 0.08/0.06/0.04 rad per policy step) against the current vel_limit setting.

    Change

    The probe decided the role of the limiter rather than a deployment.

    Outcome

    Success 14/15 -> 12/15; get-up median 2.35 -> 3.53 s (max 9.30); torque demand peak median hip_pitch 164% -> 111%, knee 166% -> 86%; action saturation still 100%; leg-leg contact 558 -> 860 frames. The limiter was kept only as a real-robot safety net (mirrored into sim2sim evaluation); the cure moved into training - where the next lesson was that a limiter anchored on the last command is itself an integrator (slew-anchor-is-an-integrator).

    Mechanism

    A policy that never trained with the limiter keeps issuing the targets it learned; the limiter clips them, the target window runs ahead (windup), and the robot follows a trajectory the policy never evaluated.

    Applies when

    • a trained policy is too violent on hardware and a quick deploy-side fix is tempting
    • adding slew, torque or velocity limits in a bridge or firmware
    • evaluation and deployment use different limiter settings
    “判读:**链路侧收紧立等可取地把 τ 峰值砍 30~50%,但成功率掉、饱和率仍 100%、 腿-腿接触反升** —— 策略感知不到限速器,目标窗口继续狂奔。⇒ 收紧 slew 只配当 **真机侧安全网**(必须同步进 sim2sim 口径,基础设施现成),**不配当治法; 治法必须进训练**。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 探针:收紧桥层 slew,r3_1 不重训直接测
  • Deployment power derating damages non-forward axes far more than forward - sweep it in sim before deployingpower-scale-hurts-nonforward-axes
    Replicatedomnireal-deployreal-acceptanceactuator-modelingattribution

    Treat deployment power/torque scaling as a plant parameter: evaluate the policy in sim at the exact deployment scale, expect non-dominant axes to degrade first under derating, and either deploy at the training power or train with power randomization.

    Symptom

    Policies deployed at power-scale 0.8 (a safety derating of commanded torque) looked fine walking forward but were weak at backward and turning, inviting the wrong diagnosis "the skill was not trained well".

    Context

    Measured repeatedly: on s1e, going 1.0 -> 0.8 cost forward 18% but backward 58%; on C4-ff800, turn tracking was +25%/+40% at pw0.8 vs +75%/+58% at pw1.0, backward 51-52% vs 97-103%, while forward stayed 96-98% at both. Sim evaluation numbers in the plan were all pw1.0, but the robot was being run at 0.8.

    Change

    Pre-deploy protocol added: sweep the exported policy across power in sim (for PW in 0.8 0.9 1.0: eval_c_matrix --power $PW --seeds 20) and deploy at the first level where both turn directions reach >=50%. For C4 the recommendation was raise the robot to pw1.0 - the sweep showed it nearly free (saturation 47%->33%, left foot-clipping danger zone 25%->6%, cost only tilt 6.7->8.3 deg).

    Outcome

    Turning "weakness" resolved without any retraining; the sim sweep correctly predicted the real-robot signature at both power levels.

    Mechanism

    Forward walking is the reward-dominant, torque-cheapest skill with the most margin; backward/turn/sidewalk live closer to the torque envelope, so a uniform torque derating consumes their margin first. Training ran at power 1.0 (the trainer does no power scaling), so deploying at 0.8 is a systematic underactuation the policy never experienced.

    Applies when

    • deploying with any torque/power derating or safety scale
    • secondary skills (backward, turn, lateral) underperform on hardware while forward walking looks fine
    • choosing the deployment power level for a new policy
    “power 衰减对非前进轴的伤害远大于前进轴(s1e:前进 1.0→0.8 掉 18%,后退掉 58%)。转向是非前进轴,0.8 下很可能明显跟不动。”
    train/C_LADDER_RUN.md § 3c. A-2 上机前先定部署力度档 / 3p. 二
  • A quadratic clearance reward stalled for 3500 iters near target - switch to an indicator on accumulated heightindicator-reward-avoids-gradient-decay
    Mechanism understoodwalkreward-shapingreward-shaping

    When a shaped term plateaus near its target, check the gradient profile: replace vanishing-gradient forms with threshold/indicator forms for the final approach, and prefer delta-accumulation over absolute positions to immunize against frame offsets.

    Symptom

    The quadratic-error clearance term froze at -0.002 from iteration 2000 to 5500 - thousands of iterations with no progress on foot lift.

    Context

    Diagnosis: a quadratic penalty's gradient vanishes as the error approaches target, so exactly where the last millimeters must be earned the incentive fades to nothing. The replacement (Humanoid-Gym form): accumulate the swing-phase height climb per foot, reward a BINARY indicator |accumulated - target| < 0.01 masked to the planned swing window, reset on contact, weight +1.6 as a positive reward. Two properties: the indicator's incentive is constant until the threshold is crossed (no decay zone), and accumulating height DELTAS makes any constant sole-frame offset cancel automatically - which structurally sidesteps the earlier 0.0585 m zero-point bug ("顺带绕开我先前那个'忘了减 0.0585 导致惩罚恒为 0'的坑").

    Change

    Clearance reformulated from quadratic penalty on instantaneous height to indicator on per-swing accumulated climb (target 0.03 m by leg-length scaling, weight +1.6).

    Outcome

    Part of the v5 package under which lift finally moved (v5 29 mm, v6 34 mm vs the stalled 18-24 mm era); the offset-cancellation property removed one whole bug class from the term.

    Mechanism

    Policy-gradient learning follows the reward's local slope; quadratic shaping concentrates slope far from target and starves it near target, so convergence stalls precisely at the finish line. An indicator pays a constant bounty until the goal is met; formulating on deltas rather than absolutes removes sensitivity to reference- frame constants.

    Applies when

    • a reward term's value freezes short of target for thousands of iters
    • designing clearance/height/precision terms
    • reward code depends on absolute link positions
    “现行二次型在接近 target 时梯度趋零 —— 这正是 clearance 从 iter 2000 到 5500 卡在 −0.002 不动的原因。… 二值指示在跨过阈值前梯度恒定,没有衰减区 … 累积 delta 让 SOLE_OFFSET 自动抵消”
    train/WALK_V5_SPEC.md § 3. clearance 改峰值型(去掉二次型的梯度衰减)
  • Every acceptance run records video of the very rollout that produced the numbers - the seated basin, edge-standing feet, tangled legs and the narrow stance were all seen on video before, or instead of, a metric catching themvideo-as-acceptance-record
    Replicatedrecoverysim-evalgate-batteryreal-acceptancemeasurement

    Make video a default output of every acceptance run, rendered from the same rollout the metrics come from (fixed views including the feet), and watch it - posture failures are visible before any gate row exists for them; never let video replace or override the numeric gate.

    Symptom

    Posture failures in the recovery line kept arriving as things the numbers had no row for: a kneeling W-sit, feet standing on their outer edges, crossed legs, a stance too narrow to hold on hardware.

    Context

    From 2026-08-09 (user decision) accept_recovery renders offscreen by default, following one env for the whole episode and archiving the clip. The MuJoCo gate's --video renders three views (side, front, feet) of the same rollout the metrics come from, with all plant modelling (delay, push); the older replay-based renderer produced an independent trajectory without delay and was not used for acceptance. Video never gates: if rendering breaks, --no-video keeps the numeric gate running.

    Change

    Video as a default acceptance artifact, named per policy, category and view, reviewed by the user.

    Outcome

    The R0.1 prone clip showed the same kneel-sit as R0; the R3.1 failure clip showed the crossed legs; the v2_5 feet view showed edge standing and led to V2.6; the v2_6 videos led the user to order a real-robot A/B between v2_5b and v2_6. Once, the recorder did not start (P1c final acceptance) and the visual material had to be produced separately.

    Mechanism

    Metrics exist only for failure modes someone anticipated; video shows the unanticipated ones, and rendering the metric rollout itself guarantees the picture and the numbers describe the same episode.

    Applies when

    • setting up an acceptance pipeline for posture-sensitive skills
    • numbers pass but a human reviewer is uneasy
    • sim videos are rendered by a separate replay tool
    “**验收存视频(用户定 2026-08-09)**:`accept_recovery.py` 默认开 Isaac 离屏 渲染,跟拍一个 env 的整局并归档 … 跪坐这类盆地在数字表出现前肉眼先看见,视频是验收的 定性存档,数字门不受它影响”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §5 R0 验收门(预注册)验收存视频(用户定 2026-08-09)
  • When the eval proxy has a known systematic bias, gate on within-proxy differences, not absolutesrelative-metrics-survive-proxy-bias
    Mechanism understoodwalkgate-batterygate-batterysim2simattribution

    Where the evaluator is known-biased, design gates as within-evaluator contrasts (left vs right, A vs B, pre vs post) that cancel common-mode error; reserve absolute thresholds for quantities whose proxy calibration has been checked.

    Symptom

    MuJoCo systematically overestimated yaw turn gain (1.58/2.45 where Isaac measured 1.04/1.02), and friction alignment recovered only part of the gap (mu 1.0 -> 0.6 pulled it to 1.43/2.27) - so any absolute turn-gain acceptance threshold in MuJoCo would be judging the proxy's bias, not the policy.

    Context

    The acceptance criterion was rewritten to use only the left/right difference of the turn gain (standard: <20%): both directions pass through the same biased proxy, so the bias largely cancels in the difference while the policy's chirality - the thing being gated - survives. The absolute-gain row was dropped: "转向只判左右差,不判绝对值". A parallel task was still opened to align MuJoCo contact parameters to Isaac's (mu, restitution), since drift proved highly friction-sensitive (straight-line drift -65 deg at mu 1.0 vs +1 deg at mu 0.4) - bias reduction and bias-robust metrics proceeded together.

    Change

    Gate metric changed from absolute turn gain to left-right gain difference; proxy-alignment work scheduled separately rather than blocking acceptance.

    Outcome

    Turn acceptance became meaningful across proxy versions (v5 70% -> v6 19% difference measured the real improvement) while the absolute bias question was pursued without holding the ladder hostage.

    Mechanism

    A systematic multiplicative or additive proxy bias applies to both arms of a mirrored measurement; differencing (or ratioing) mirrored conditions cancels the common-mode bias to first order, leaving the asymmetry signal. Metrics built this way remain valid while the proxy is imperfect - which it always is somewhere.

    Applies when

    • a sim proxy disagrees with the trainer or hardware in absolute terms
    • writing acceptance thresholds for direction-paired skills
    • proxy calibration work would otherwise block a ladder
    “转向只判左右差,不判绝对值 —— MuJoCo 的偏航增益系统性高估(实测 1.58/2.45 vs Isaac 1.04/1.02),摩擦只能解释一部分(μ 1.0→0.6 仅拉到 1.43/2.27)。差值是相对量。”
    train/WALK_V5_SPEC.md § 6. 验收
  • Prove an armless get-up exists before training it - a connected static domain, 25% torque on the cheapest path, an 8 mm hand-over gap, a static roll-over - and write down what each scan cannot representget-up-feasibility-accounts-before-training
    Mechanism understoodrecoveryplant-calibrationplant-calibrationhardwareprocess

    Before training a get-up or any multi-contact skill, compute the quasi-static accounts - connectivity of the static domain, torque along the cheapest path, hand-over gaps, COM shift available for rolling - and state which configurations each scan cannot represent; when a policy gets stuck in one of those, extend the scan before blaming the reward.

    Symptom

    A torso-and-legs robot has no arms to push off the ground; whether it can get up from the floor at all was unknown when the line opened.

    Context

    recovery_feasibility.py ran three accounts before any training (the run line's "hard accounts first" discipline): a sagittal quasi-static scan (0.05 rad grid, 44,520 configurations, MuJoCo FK, flat-foot assumption). (1) The static standing domain (COM over the feet, torques in limit) has 29,586 cells, flood-fill connected with no islands, from a 0.097 m deepest squat to the 0.384 m stand. (2) The minimum-torque path peaks at 25% of the limits (knee 2.9/12, ankle 3.7/17 N*m) - a 4x margin. (3) All 508 ground-contact configurations have the contact behind the COM; the smallest gap to pure foot support is 8 mm. A roll-over account: swinging both straight legs to one side shifts the COM 96 mm against a 62 mm torso half-width - 1.6x, so rolling needs no momentum. Three conclusions were written down for later attribution: the legs are 80% of the mass (swinging them moves the COM), prone has no flat-foot hand-over face (merge into a supine/side sit first), and supine needs no sit-up (hip flexion is limited to 75 deg).

    Change

    The accounts gated opening the line and were cited in every later argument about what the robot can physically do.

    Outcome

    They held where they applied: in V1.0 every fall category was righted under a hard rate limit, which the spec records as the quasi-static roll-over account verified by training, and the 25% torque path was the basis for pursuing a slow get-up. They also misled once: account (3) is sagittal, and on 08-09 the spec corrected its scope - it cannot represent the splayed W-sit where the policy actually stalled. A follow-up prone hip-ROM scan (471,625 cells) found 3,912 two-foot-contact cells and none with both soles within 25 deg of level (best 40.2 deg): a flat-footed push-up from prone is infeasible on this robot, so the fix became where the feet go after sitting up.

    Mechanism

    A get-up needs a connected path through statically feasible configurations and enough torque along it; quasi-static accounts bound both cheaply, and momentum can only make the real problem easier. A reduced-dimensional scan, though, only speaks for the configurations it can express.

    Conflicts

    In R0.1-R0.2 the spec read account (3)'s "prone has no front hand-over" as "prone lacks the roll-over skill"; R0.3's confusion matrix showed prone had righted its torso 159/159, and the spec then restricted account (3) to the sagittal configurations it models.

    Applies when

    • opening a get-up, recovery or climbing skill on a new robot
    • a robot lacks arms or other obvious contact options
    • a policy stalls in a configuration a feasibility scan never modelled
    “本机 **torso + legs、无手臂可撑地**,开训前先证明存在不依赖手臂的物理解 … 双腿同侧直腿摆最大横移 **96 mm = 1.6×** … **静态摆腿即可翻身,无需动量**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §1 机械可行性判决(2026-08-09,recovery_feasibility.py 三笔账)
  • Every gate on a penalty is an exit - to stop paying a stance tax the policy parked 2 deg outside a 30 deg uprightness gate (lunging), and, from scratch, just under a height gate (crouching); a positive band and an always-on guard fixed bothpenalty-gate-is-an-escape-hatch
    Replicatedrecoveryreward-shapingreward-shapinggate-battery

    Never gate a penalty on a state the policy can leave by getting worse; use always-on guards for what must never happen and positive, gated bands for what you want, and count every gate on a penalty as one more escape route to check in the logs.

    Symptom

    V2.9 added stance_width_task = relu(0.34 m - foot spacing) x standing gates, weight -10. Three checkpoints scored 0% on acceptance: standing height reached, feet on the ground, angular rate low, but the torso leaned 32.5/32.3/31.9 deg in a fore-aft lunge. The training dashboard read "tax paid off, base_height at full value".

    Context

    The penalty was gated by uprightness (tilt < 30 deg) and standing height; its tax had no time gate (500 steps x -1.65) while the standing income sat behind a 3 s zero gate (~325 steps), so leaning just past 30 deg lost a little gated income and saved the whole tax. base_height has no upright gate, so the lunge still collected it. The first metric also measured full horizontal spacing, so a staggered lunge counted as "wide".

    Change

    Two laws written down: a penalty may only carry gates the policy cannot escape by getting worse (make it an always-on guard) or it becomes a positive band ("not earned" is not "escaped"); and width is measured laterally in the base yaw frame. From scratch (V3.1 P1) with the lateral metric but the same gates, the policy parked just under the height gate instead (base_height 1.176/1.5, h ~ 0.30 m against a 0.3264 gate), the beta curriculum never advanced in 1,700 iterations, and the run was stopped early. P1b flipped the penalty into a positive band +2.0 x clamp(lateral / 0.34) x standing gates; P1c added yaw_guard = -5 x relu(|hip_yaw| - 30 deg), always on, no gate, no exemption.

    Outcome

    P1b: lateral stance 0.364 m, all four categories 100%, MuJoCo mu 1.0/0.4 both 100%. P1c: all six acceptance criteria passed for the first time on the line (hip-yaw saturation 1.2%), the guard's tax converging to -0.016 (almost never touched).

    Mechanism

    A gated tax that is not paid is saved, so the policy moves to the cheapest state just outside the gate; a gated income that is not earned is simply lost, so a positive band has no exit. HoST's style penalties are ungated or binary - the same law seen from the other side.

    Applies when

    • adding a penalty multiplied by an uprightness, height, phase or contact gate
    • a policy settles just beyond a gate threshold
    • training metrics look paid-up while acceptance collapses
    “**罚项的门 = 策略的逃生门**。带直立门的负项可以靠"变得更差"(倾出门外) 全时免税;HoST 的 style 罚全部无门控/二值恰是同一律的反面实证。修律: 负项只许挂"变差逃不掉"的门(上限护栏 always-on),或改正向 band (收不到 ≠ 逃掉)。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §46 定案(血统内第四败 + 两条新律)
  • Before training a one-leg stand the spec named the cheapest cheats - hopping on the support foot, a raised foot resting unloaded, a leg tripod - and gave each a countermeasure and a gate; one still appeared and was caught by exactly those gatesenumerate-cheapest-cheats-before-training
    Observed onceonelegreward-shapingreward-shapinggate-battery

    Before training, list the cheapest behaviours that would satisfy each reward term without doing the task, give each a countermeasure in the reward and a gate in acceptance, and prove the intended behaviour is reachable with a probe - then treat any gate the policy games as evidence about the reward, not the gate.

    Symptom

    The literature's single-leg benchmark reports eight state-of-the-art general policies holding a clean one-leg stand 0 times out of 90 - they survive by sneaking steps and hops - so the task's first adversary was the policy's own cheating.

    Context

    The spec's shape self-check ("what is the zero-cost option?") listed, for the one-foot bucket: the cheapest cheat, a foot resting on the ground without load, countered by a 5 N contact threshold plus positive swing income; the second cheapest, small hops on the support foot to reset balance, countered by a continuous support-air penalty plus a gate of zero support-foot flight segments. The probe that preceded training had already seen a third: early low-lift postures "survived" by pressing the swing foot at 78-95 N, a leg tripod, removed by folding the shank back. The two-foot bucket was checked too: its zero-cost behaviour is ordinary standing, with no odd base state.

    Change

    Countermeasures and gates written before training: swing-contact and support-air penalties, gate 2 (zero swing-foot contact frames above 5 N), gate 3 (zero support-foot flight segments).

    Outcome

    The first run still found the unloaded-foot cheat (a binary reward band gave it no gradient to lift) - and it was caught, by the contact gates and the cross-simulator comparison, not discovered on hardware. The retrained V0 passed all gates 40/40, including zero support-foot flight after the flight detector was corrected.

    Mechanism

    A policy optimizes the reward, not the intent; the cheapest behaviours that satisfy the reward are predictable from the reward's structure, and a gate written for each before training turns a silent cheat into a failed row.

    Applies when

    • designing rewards for balance, contact or "hold still" tasks
    • benchmark policies are known to cheat the task
    • writing acceptance gates for a new skill
    “文献里 8 个 SOTA 通用策略在单脚站基准上 0/90 干净保持, 全靠偷步偷跳活命,这是本任务的第一反作弊对象 … 单脚桶下最便宜的作弊是"脚虚放地上不受力"——接触判定 >5 N 力阈(沿用),配 swing_height_band 正收入拉开。 … 第二便宜是"支撑脚小跳重置"——support_air_penalty 连续罚 + 验收门支撑脚腾空段=0 双保险。”
    git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §0 目标口径 / §5 形状自检(零成本选项是什么)
  • The latency DR range must cover the measured deployment pipeline - 0-20 ms could not even reach the real 1-2 control stepslatency-dr-covers-measured-pipeline
    Mechanism understoodwalkactuator-modelingactuator-modelingdomain-randomizationhardware

    Measure end-to-end action latency in control steps on your own stack (including cross-process queue boundaries), set the DR range to cover it with margin, and never import a delay count without its control frequency.

    Symptom

    Action latency was randomized over 0-20 ms (0-1 control step at 50 Hz), but the measured deployment path is 1-2 steps: the deploy process writes the target, an independently running BusWorker picks it up on its NEXT cycle, plus CAN round-trip - the training range could not cover the robot's actual latency at all.

    Context

    Fix: widen action_latency_s to 0-0.06 (0-3 steps). The external reference's "uniform 6 steps" was explicitly NOT copied - that number depends on his unknown control frequency; locally, a sweep at 0/1/2/3 steps showed walk_v5 survives all with insensitive metrics, so 6 steps "在我们这里没有依据" (has no local basis). The range was set from the measured pipeline with margin, not from a foreign constant.

    Change

    action_latency_s (0, 0.02) -> (0, 0.06), justified by pipeline analysis (writer/worker cycle boundary + bus time) and bounded by the local latency sweep.

    Outcome

    The DR band now brackets the true deployment latency; the policy trains against the delay it will actually face instead of a fictional sub-step world.

    Mechanism

    Latency DR only immunizes against delays inside its support; a range below the physical pipeline guarantees an untrained distribution shift at deployment. The correct range comes from tracing the pipeline's worst case (queueing boundaries + transport), and foreign step-counts are meaningless without the control rate they were measured at.

    Applies when

    • setting or auditing action-delay randomization
    • deployment uses a separate bus/worker process from the policy loop
    • importing delay-modeling numbers from other projects
    “现行 0~20 ms = 0~1 个 50Hz 控制步, 而实测部署链路是 1~2 步(deploy 写 STATE.target 后, 独立跑的 BusWorker 下一轮才取走下发, 再加 CAN 往返)——现在的区间覆盖不到真机的实际延迟。… 不照抄参考来源的"统一 6 步": 那取决于他的控制频率(未知), 而我们扫过 0/1/2/3 步 … 6 步在我们这里没有依据。”
    train/WALK_V7_SPEC.md § ⑤ action_latency_s 0~0.02 → 0~0.06
  • IMU observation age cut 52-68 ms to ~4 ms by moving AHRS onto the MCU - as a single variableimu-age-move-fusion-downstream
    Observed oncewalkreal-deployhardwarereal-acceptanceprocess

    Audit observation age end-to-end and move time-critical fusion as close to the sensor as possible - and when you fix a latency, change only that one variable so the gain is attributable.

    Symptom

    IMU-derived observations reaching the policy were 52-68 ms old because attitude fusion ran in Python on the loaded host computer - stale attitude is a direct feedback-loop delay the policy was not trained with.

    Context

    The fix was scoped deliberately narrowly: move the AHRS computation from Python to the STM32 H7 (MC02). CAN topology explicitly unchanged, so the change is a clean single variable.

    Change

    AHRS fusion relocated Python -> H7. Before/after - IMU age: 52-68 ms -> ~4 ms; CAN timing: unchanged; Python load: high -> ~0.

    Outcome

    IMU age reduced by an order of magnitude with no confound; host CPU headroom recovered ("把计算单元搬在stm32上, 这样imu有剩余").

    Mechanism

    Sensor age is pipeline latency, not sensor quality: fusing on the MCU next to the sensor removes host scheduling jitter and interpreter overhead from the critical path. Keeping the bus topology fixed makes the improvement attributable to the relocation alone.

    Applies when

    • measured sensor-to-policy age far exceeds sensor sample period
    • attitude fusion or filtering runs on a loaded host CPU in an interpreted runtime
    • planning infrastructure changes during a sim2real campaign
    “AHRS 搬到 H7——这个不改 CAN 拓扑,只是把一段计算从 Python 挪到 MC02,单变量:IMU age 52–68 ms → ~4 ms / CAN 时序 不变 / Python 负载 高 → ≈0”
    Experience.md § AHRS 搬到 H7 (lines 28-35)
  • Freeze the deployment contract, stamp every export, and let an automated checker catch wiring bugscontract-freeze-and-checker
    Replicatedomniprocesscontract-freezeprocesssim2sim

    Freeze and fingerprint the policy I/O contract; ship contract changes as new versioned profiles that leave old artifacts bit-identical; and extend the automated contract checker with every pipeline change, forcing the new path to execute in the check.

    Symptom

    Contract-level changes (observation layout, action pipeline) are where silent sim/real divergence is born; two real wiring bugs appeared the one time the action pipeline was extended.

    Context

    The 215-dim observation contract was frozen ("纪元 3,三机 digest" - an era number plus a digest agreed across three machines); proposals that would break it (e.g. a GRU memory) were rejected on contract grounds. Every exported ONNX is stamped and verified with a manifest (onnx_manifest --stamp / --verify), and deployment refuses mismatched combinations. When C4 added the lateral feed-forward, it went in as a NEW profile (omni_ff) leaving the existing omni profile's behavior bit-identical; the checker (check_contract) was extended to force the feed-forward path to actually execute (cmd_vy=0.13) and promptly caught two genuine bugs: (1) re-clamping with soft_joint_pos_limits after the feed-forward (0.23 rad deviation) instead of reusing the parent's clip; (2) indexing processed actions by asset.joint_names instead of the action term's own contract-ordered _joint_names, which landed the feed-forward on the wrong joints (l_hip_yaw / r_ankle_pitch).

    Change

    Contract discipline as implemented: frozen dims + digest; manifest stamping and refusal; contract changes only via new versioned profiles; checker updated in the same commit as any pipeline change, with inputs chosen so new code paths are exercised.

    Outcome

    Both wiring bugs caught before any training or deployment ("两个都是 check_contract 当场抓出来的 —— 这次它值回票价"); old deployments provably unaffected by the new profile.

    Mechanism

    The contract is the only interface the policy and robot share; freezing plus fingerprinting makes divergence detectable, and an executable checker turns "the contract holds" from a belief into a test - but only if its inputs actually drive the new code path.

    Applies when

    • modifying the action or observation pipeline of a deployed policy
    • exporting policies for hardware
    • proposals that would change observation dims or history structure
    “契约校验抓到的两个真错误(记账,别再犯):1. 前馈后误用 soft_joint_pos_limits(URDF 限位 ×0.9)重钳 → 0.23 rad 偏差 … 2. 用 asset.joint_names 索引 _processed_actions → 前馈落到 l_hip_yaw/r_ankle_pitch 上 … 两个都是 check_contract 当场抓出来的 —— 这次它值回票价。”
    train/C_LADDER_RUN.md § 3j. 契约级改动 / 契约校验抓到的两个真错误
  • Symmetrizing the config made the gait MORE asymmetric - the asymmetry lived in the policy weightsasymmetry-in-weights-not-config
    Mechanism understoodwalkattributionattributionreward-shapingcurriculum

    Localize a persistent asymmetry by intervening at the config layer first: if the symptom survives (or worsens), it is in the weights - fix it with symmetry-constrained training, not with trims or offsets.

    Symptom

    walk_v1 on hardware: straight-line command curved 149 deg in 15 s (9.9 deg/s) with 3.06 m lateral runout; turn gain +31% one way vs +129% the other (75% difference); knee asymmetry 4.4 deg in sim, 9.6 deg on the robot.

    Context

    The obvious suspect was the asymmetric default pose in the config. The decisive test: symmetrize standing_pose and run the SAME policy in sim - the asymmetry got LARGER (hip_pitch 6.8 -> 9.3 deg). Root cause therefore not in the config but baked into the policy weights: PPO without a symmetry constraint commonly converges one-sided, because splitting the work 50/50 and loading one side yield the same return, and the gradient falls randomly into one of the equivalent optima.

    Change

    Fix redirected from config trimming to retraining with mirror data augmentation (walk_v2 spec) - a weights-level fix for a weights-level disease.

    Outcome

    With augmentation (and the symmetric-default precondition), stand_v1 reached 0.0 deg asymmetry on all six joint pairs (from 4.4-7.7 deg), height fluctuation 7 mm -> 1 mm, mean |action| down 33%.

    Mechanism

    Reward-equivalent solution families (who carries the load) leave the symmetric solution unpreferred; SGD picks an arbitrary member and entrenches it. Config changes move the coordinate frame around the entrenched asymmetric function - they cannot move the function. The counterfactual test (change config, watch symptom) localizes the layer the disease lives in.

    Applies when

    • a robot veers or loads one side despite a symmetric-looking config
    • deciding between config trims and retraining for an asymmetry
    • mirrored-turn gains differ by tens of percent
    “根因不在配置里:把 standing_pose 对称化后在 sim 里跑同一策略,不对称反而变大(hip_pitch 6.8°→9.3°)—— 说明不对称烙在策略权重里。这是无对称约束的 PPO 的常见收敛结果(左右各担一半与一边多担的回报相同,梯度会随机落进其中一个)。”
    train/RETRAIN_v2.md § 1. 为什么是对称增强(证据)
  • A time gate that had cured one lineage's rushing made the from-scratch lineage trade away its stance width twice (0.364 -> 0.235 m, 0.355 -> 0.251 m) - its disease was absent there, so the fix was retired and the pre-gate checkpoint shippedtime-gate-vs-wide-stance-retire-the-fix
    Replicatedrecoveryreward-shapingreward-shapingcurriculumfork-selection

    Carry a fix into a new lineage only if its disease is present there; a mechanism that cured one lineage can be net negative in another, and when doubling a term's weight recovers almost nothing, treat the two objectives as structurally in conflict and remove the one whose purpose is gone.

    Symptom

    V3.1's phase 2 (the 3 s zero gate on standing income, continued from P1b) kept 100% success on every friction level and slowed the get-up, but the lateral stance drifted 0.364 -> 0.235 m and hip yaw crept to 57 deg against its 60 deg limit. With the width band's weight doubled (P2c, after P1c) it drifted again, 0.355 -> 0.251 m, below the pre-registered 0.30 m failure line.

    Context

    The zero gate had been introduced in V2.5/V2.5b to slow the old lineage's get-up. In V3.1 the rushing was already absent: P1c got up in 0.90-1.06 s with a worst torque ratio of 73.3%, better than the stamped v2_6c, because the full beta curriculum, second-difference smoothing and pull curriculum had cured the violence inside training.

    Change

    Recorded as a candidate law with two data points - the zero gate and a wide stance are mutually exclusive here - and the zero gate was removed from the V3.1 recipe. P1c (the pre-gate checkpoint) went through the full stamp-level acceptance instead.

    Outcome

    P1c passed everything: all six criteria, lateral stance 0.355 m, foot tilt P75 2.0 deg, mu {1.0, 0.8, 0.6, 0.4} x 10 seeds all 100%. recovery_v3_1p1c.onnx was stamped and pushed to the robot channel.

    Mechanism

    The zero gate moves the income toward "stay stable until the end", and under low-friction DR a wide stance has a slip tail, so survival outbids the width band; doubling the band's price bought back only 0.016 m - an auction that does not converge signals structural conflict, not an under-priced term.

    Applies when

    • porting reward mechanisms from an old lineage into a fresh recipe
    • a width, margin or posture metric erodes during a late training phase
    • a weight increase produces a negligible change in its target
    “**定律候选(二实证):归零门 × 宽站互斥**。 … 加价翻倍只挽回 0.016,竞拍不收敛)。 … **归零门是 v2_5 血统的历史包袱,对 V3.1 配方是净负资产,P2 阶段除名 —— P1c 即终点形态**。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §49 终验(2026-08-14)
  • Training the final recipe from scratch in one run - every mechanism the lineage had accumulated - produced 0% and a seated robot; the order in which the lineage acquired those mechanisms was part of why it workedcurriculum-history-is-part-of-the-product
    Observed oncerecoverytraining-runcurriculumprocess

    A recipe that ends a lineage is not a recipe for a from-scratch run: consolidate it as ordered curriculum phases matching how the lineage acquired its mechanisms, check that every curriculum criterion is reachable from the starting policy, and read the run's raw term values, not the total reward, before calling it green.

    Symptom

    V3.0 trained the lineage's whole final recipe from scratch in one 9,000 iteration run - full beta curriculum, prone-conditioned pull assist, friction DR, the 3 s zero gate on standing income, flat_feet - testing the proposition "the product is defined by its configuration, not by its training history". Training looked all green (reward 26.33, episode length 500, 100% time-outs).

    Context

    Read in raw units against v2_6c at the same weights, the green was a seated equilibrium: base_height 0.377 vs 0.640, stand_pose 0.205 vs 0.516, flat_feet 0.0000 (zero because it sits outside its height gate, not because the feet were flat). Acceptance: 0.0% in Isaac at the deployed authority, 0.0% on every MuJoCo friction level, and still 0.0% at the training-time authority (100% seated at 0.222 m, upright and still).

    Change

    The full stdout (316k lines) was read: the beta curriculum's criterion (standing share over 0.35) was met zero times, so beta never left the wide setting and the policy had no experience at the deployed authority; the pull curriculum was stuck on the same criterion. The lineage had escaped the seated basin with immediate income (V2.0-V2.2) and only then added the zero gate to cure rushing (V2.5); from scratch, the zero gate removed the early "stand fast, earn more" gradient needed to escape. Verdict "the curriculum history is part of the product", limited to n = 1. v2_6c stayed the product.

    Outcome

    V3.1 kept the order as explicit phases: P1 from scratch with immediate income (zero gate off) until the curricula advance, P2 adding the zero gate. P1b/P1c escaped the seated basin and passed; P2 was later judged net negative and dropped (time-gate-vs-wide-stance-retire-the-fix).

    Mechanism

    Mechanisms that refine a competent policy (time gates, tight authority) can delete the gradient a naive policy needs, and a curriculum whose advancement criterion the naive policy never meets freezes at its first level.

    Conflicts

    The spec limits the falsification to "this recipe + this curriculum criterion" (n = 1, no seed sweep, no criterion tuning). V3.1's phased run succeeding is consistent with the ordering reading but changed other terms too.

    Applies when

    • consolidating a long lineage of continuation fixes into one clean recipe
    • a from-scratch run with all mechanisms enabled plateaus early
    • curriculum state is not logged or never advances
    “命题:产物由配置定义,而非训练史定义。 … 训练 log(完整 stdout 316k 行)`[beta_anchor]` 仅初始 1 行,**达标 0 次** … V2.0~V2.2 靠**即时计酬**爬出坐姿盆地(§33),站立巩固后 V2.5 才装归零门 治"过快"(§39)。**课程史是产品的一部分。**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §42 V3.0 判决(2026-08-11):从零单 run 全机制 FAIL 于坐姿盆地
  • Decompose the offending quantity by channel first - then penalize the failure event, not the jointspenalize-the-slip-not-the-joint
    Mechanism understoodwalkreward-shapingreward-shapingattribution

    Before penalizing motion to fix a side effect, measure which channels actually carry the offending quantity; prefer penalties conditioned on the failure event that are exactly zero for healthy behavior - and do not medicate behaviors that measurement shows are not sick.

    Symptom

    Heading drift with support-foot yaw slip (v5: 212-284 deg accumulated over 15 s); the previous v6 draft had attacked it by penalizing lateral joints (a roll 4.0 / yaw 2.0 "home" group) - which collapsed training into the standing basin.

    Context

    Before choosing the penalty target, the yaw angular momentum was decomposed by joint group with MuJoCo subtree_angmom weighted by real walking joint velocities: pitch-class joints (hip_pitch + knee) carry 95.3%, hip_roll 3.3%, hip_yaw 1.4%. The failed "home" group had been taxing 2.7/step to manage a 4.7% channel. The replacement, feet_yaw_slip (-0.2, |support-foot yaw rate| while in contact), targets the failure event itself and - decisively - costs a non-slipping gait exactly zero, which "横向回家组做不到". The same rung's do-not-do table applied the complementary principle to foot spacing: measured 196-214 mm, stable, no crossing - "没病不吃药" (no disease, no medicine).

    Change

    Removed joint-usage penalties for the drift problem; added the event-conditional slip penalty (-0.2, realized tax 0.141/step = 12% of tracking) alongside the existing linear-slip term.

    Outcome

    Turn-gain left/right difference improved 70% -> 19% and heading 185 -> 60.3 deg by v6 without a standing-basin collapse; the 2.7/step lateral tax never returned.

    Mechanism

    Penalizing joints taxes every use of a channel including healthy use, and if the channel carries little of the offending quantity the tax buys nothing while pushing the optimum toward immobility. An event-conditional penalty (slip while in contact) prices only the failure, leaving the healthy gait's cost surface untouched - and the channel decomposition tells you in advance whether a joint-side fix can even work.

    Applies when

    • choosing a penalty target for drift/slip/impact problems
    • a proposed penalty taxes joints or motions rather than failure events
    • a previous joint-penalty attempt collapsed the gait
    “pitch 类 (hip_pitch + knee) 占偏航角动量 95.3% … hip_yaw 1.4% … 压 hip_yaw 是管 1.4% 的通道收 2.7/步 的税 —— 上一轮正是这样把策略推进了站立盆地。滑移项不惩罚走路: 不打滑的步态代价为零, 这是横向"回家"组做不到的。”
    train/WALK_V6_MINIMAL.md § ① / ② 新增 feet_yaw_slip
  • Cutting swing amplitude 40% raised yaw-momentum demand 53% - the falsified fix is recorded so nobody walks that road againamplitude-cut-falsified-yaw-fix
    Mechanism understoodwalksim-evalmeasurementattributionreward-shaping

    Test gait fixes against the quantity the ground must actually supply (torque/force rates vs friction ceilings), not against kinematic proxies; record falsified fixes with their mechanism so the search space shrinks permanently.

    Symptom

    Support-foot yaw slip stayed at ~223 deg (vs 225 deg) after walk_v6 cut joint swing amplitudes by ~40% (hip_pitch 20.6->10.4 deg, knee 31.3->18.8 deg) - the change built on the theory "smaller swing = less yaw momentum to dump into the ground".

    Context

    Direct measurement inverted the theory: yaw-momentum amplitude ROSE 16% (+/-0.1000 -> +/-0.1159) and its rate of change rose 53% (1.86 -> 2.84 N*m demanded from the ground), pinned exactly at the foot's supply ceiling (2.8-3.3 N*m at mu 0.6-0.7) - so slip could not drop. The extra demand lives in higher harmonics: v6's crisper foot placement (better clearance 34 mm, lower landing force) shortens the momentum- exchange window, concentrating the same exchange into less time. The verdict was written as a closed road: "下一轮不要再走这个方向". An honest residue was also booked: WHY amplitude down but momentum up 16% remained unresolved, with the named next step (per-rigid-body decomposition of H_z, since per-joint RMS sensitivity ignores phase correlations).

    Change

    The "reduce amplitude to reduce yaw momentum" lever was removed from the planning space; future yaw-slip work redirected toward the supply side (friction) and momentum-rate mechanics.

    Outcome

    Slip unchanged (225 -> 223 deg); the falsification and its mechanism became a permanent constraint on the fix search space.

    Mechanism

    Ground yaw torque demand scales with the rate of change of angular momentum, not its amplitude; kinematic amplitude cuts that also sharpen contact timing can raise dH/dt while lowering range. When demand exceeds the friction-limited supply ceiling, slip is set by the ceiling, so demand-side changes below the ceiling do nothing visible.

    Applies when

    • attacking foot slip or yaw drift via gait shape changes
    • a fix targets an amplitude while the constraint is a rate
    • documenting a failed intervention after a version comparison
    “walk_v6 | ±0.1159 (+16%) | 2.50 Hz | 2.84 N·m (+53%) … 脚的供给上限 2.8~3.3 N·m(μ 0.6~0.7)—— v6 正好顶在天花板上, 所以滑移一点没降。… 结论: "减小摆动幅度以降低偏航动量"这条被 v6 证伪 —— 砍 39% 幅度, 需求反升 53%。下一轮不要再走这个方向。”
    train/WALK_DIAGNOSIS.md § ② 未生效的机理: 减小摆动幅度反而让偏航需求上升

Next page