Skip to content

Training Coach

The coach reads a training run and returns a diagnosis, in-whitelist proposals and an experiment plan — and every claim cites one of the cards below. This is that corpus: 22 methodology rules distilled from a real sim2real programme, and 171 episode cards behind them, each with the sentence in the war history it came from.

Doctrine

A report may cite any of these as doctrine-N.

  1. doctrine-1Contract freeze and fingerprint discipline

    The policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.

    Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).

    Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.

    contract-freeze-and-checkerlegacy-profile-pinningderived-asset-staleness-checkgain-profile-belongs-in-the-stamprecovery-two-policies-and-a-state-machinewalk-recovery-fsm-handoff

  2. doctrine-2Attribution by resolved training params - never eval-override knobs

    Capability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.

    Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).

    Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.

    kd-bandwidth-mu-law-attributiondeploy-knob-attribution-before-retrainingcycle-time-override-is-oodresolved-config-is-source-of-truth

  3. doctrine-3PASS gates become constraints; FAIL gates become objectives

    Once a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.

    Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.

    Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".

    fixed-acceptance-matrix-per-rungpreregistered-stop-criteria-per-rung

  4. doctrine-4One variable per ladder rung - counted against what the checkpoint saw

    A rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.

    Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.

    Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.

    resume-state-dr-auditorthogonal-batch-with-ablation-orderfill-the-missing-factorial-cell

  5. doctrine-5Pre-register risks, readings, and stop criteria before the ladder

    Before a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.

    Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).

    Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.

    preregister-risks-and-fork-readingspreregistered-stop-criteria-per-rungpreregistered-real-expectationsfeasibility-accounts-lock-design-point

  6. doctrine-6Plant parameters are measured, never invented

    Every plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.

    Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).

    Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.

    friction-measured-not-guessedarmature-n2-rotor-inertiatorque-limit-shape-by-measured-peakslatency-lerp-reverse-extrapolationfeasibility-accounts-lock-design-pointsim-api-friction-columnsget-up-feasibility-accounts-before-trainingsingle-support-gain-authority-probe

  7. doctrine-7Sim2sim gate before sim2real - under deployment conditions

    Every checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).

    Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).

    Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.

    sim2sim-gate-before-sim2realeval-plant-honesty-contact-paramspipeline-latency-is-plant-not-drbody-frame-velocity-api-audittorque-penalty-bought-by-leg-bracingtorque-disagreement-between-simulators-unresolved

  8. doctrine-8Observation honesty - the actor's inputs are a hardware contract

    The actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.

    Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).

    Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.

    observation-honesty-critic-onlyhistory-obs-needs-plant-variationreward-observability-limitdeploy-heading-loop-and-align-training

  9. doctrine-9Reward economics are audited in realized currency

    Reward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.

    Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).

    Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.

    realized-contribution-auditreward-cost-of-ignoring-auditgate-new-reward-terms-by-commandignore-floor-diagnosiscalibrate-threshold-between-healthy-and-sickcalibration-threshold-with-withdrawal-clauseinert-reward-term-auditseated-basin-dead-exp-kerneltail-torque-needs-hinge-on-computed-demand

  10. doctrine-10The zero-cost option must be the desired behavior

    For every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.

    Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).

    Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.

    penalize-the-slip-not-the-jointsaturation-cheating-zero-rate-costmoving-gate-42x-stand-taxtermination-closes-degenerate-basinpenalty-gate-is-an-escape-hatchsoft-limit-penalty-charges-nominal-poseunpriced-foot-attitude-is-a-free-variableenumerate-cheapest-cheats-before-trainingbinary-band-reward-fake-touchdown

  11. doctrine-11Measurement discipline: independent referees, signs, distributions

    A disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.

    Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).

    Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.

    independent-referee-for-metric-disputesbody-frame-velocity-api-auditsame-sign-response-is-yaw-biasmedian-hides-bimodal-distributionratio-metrics-need-absolute-checkheading-integral-not-body-ratesame-distribution-reward-comparisonmultiseed-sign-test-for-driftsingle-impulse-recovery-is-chaotic

  12. doctrine-12The deployment pipeline is plant

    Irreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.

    Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).

    Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.

    pipeline-latency-is-plant-not-drpower-scale-hurts-nonforward-axesdeploy-scaling-not-training-equivalentteleop-command-band-per-axislatency-dr-covers-measured-pipelinedeploy-rate-limiter-windupslew-anchor-is-an-integratorbeta-anchored-action-targetpower-derating-cuts-full-range-contract

  13. doctrine-13DR budget is finite; its distribution is the measured support

    Robustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.

    Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).

    Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.

    push-dr-conditional-budget-conservationdr-tail-plant-continuationconstant-value-dr-overfits-margintask-shaping-before-plant-hardeningcom-randomization-forces-leg-spreadcom-dr-rollback-on-symptomthin-dr-judged-by-channel-coveragefriction-priority-re-measured-after-plant-change

  14. doctrine-14Gates measure what hardware feels: posture, margins, stripped assists

    Acceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.

    Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).

    Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.

    task-metrics-vs-posture-metricsstand-gate-posture-not-survivalconstant-value-dr-overfits-marginsuspension-probe-removes-free-stabilizerhip-roll-sum-predicts-lateral-driftchirality-scored-separatelylow-speed-commands-reward-draggingend-state-confusion-matrixfrozen-acceptance-distribution-and-pinned-seedvideo-as-acceptance-recordepisode-length-bounds-what-a-gate-seesremoved-wall-returns-on-hardware

  15. doctrine-15Fork and root selection: recoverability, maturity, frozen rewards

    Choose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.

    Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).

    Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.

    fork-root-recoverable-shortfallroot-maturity-vs-product-qualityfine-tune-reward-change-falsified

  16. doctrine-16Curricula: verified engagement, lineage counters, disease-phase gating

    Automatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.

    Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).

    Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.

    auto-curriculum-engagement-checkcurriculum-counter-lineage-stepsgate-penalties-to-the-disease-phaseaggregate-metrics-mask-subgroup-failurebucket-share-is-not-a-gradient-levercurriculum-criterion-conditioned-on-lagging-categoryper-step-income-drives-speed-time-gatetime-gate-vs-wide-stance-retire-the-fix

  17. doctrine-17Probe before training: feasibility first, hypotheses in tables

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.

    Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).

    Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).

    open-loop-probe-before-reward-tuninghypothesis-table-code-auditperiod-doubling-evidence-racesuspended-test-isolates-actuator-authorityconfiguration-probe-wall-not-slopeprone-dead-end-is-foot-placementdof-vel-penalty-is-not-a-pacing-knobamplitude-cut-falsified-yaw-fix

  18. doctrine-18External advice is recomputed locally; values transfer as ratios

    Every external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.

    Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).

    Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.

    external-advice-audit-against-own-arithmetictransfer-ratios-not-absoluteslatency-dr-covers-measured-pipelinereference-structure-fk-amplitude-divisioncycle-average-tracking-for-gait-quantitiesadvisor-paraphrase-vs-paper

  19. doctrine-19Hardware sessions are scripted experiments, not tuning sessions

    Real-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.

    Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).

    Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.

    risk-ordered-real-deploymentbattery-bracketed-real-abknow-zero-command-behaviorpush-test-chirality-protocolno-field-tuning-protocolreversible-single-variable-field-experimentssim-veto-needs-real-confirmationfirst-real-get-up-violent-stage-one-policystaged-hang-mat-floor-for-get-uppower-cycle-preflighttwo-machine-config-disciplinefall-guard-becomes-a-statehardware-log-is-the-attribution-input

  20. doctrine-20Close questions in writing; restart when the debt is structural

    Audited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.

    Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).

    Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.

    frozen-verdicts-semantic-boundariesstale-verdicts-under-old-stackzero-offset-calibration-shifts-envelopeplant-swap-invariants-vs-shiftsfreeze-lineage-fix-structure-restartminimal-reward-table-with-provenancewrite-hardware-verdicts-back

  21. doctrine-21Name the quantity in the space it lives in

    A goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.

    Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).

    Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.

    joint-space-proxy-for-task-space-quantitycontact-detector-single-signal-lieszero-partial-credit-is-not-an-iteration-problemheading-integral-not-body-rate

  22. doctrine-22Continuation needs a live gradient; a release is chosen by a scan

    Continue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.

    Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).

    Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.

    converged-continuation-is-poisoncheckpoint-choice-is-a-full-gate-scanstance-decided-by-get-up-pathstop-stacking-roll-back-and-auditcurriculum-history-is-part-of-the-productcontinuation-budget-not-from-zeroroot-maturity-vs-product-quality

Experience cards

115 cards matching “walk-recovery-fsm-handoff”.

  • An auto-curriculum ratchet capped out at iter 248 and never engaged - stage difficulty manually or verify engagementauto-curriculum-engagement-check
    Observed oncewalkcurriculumcurriculumdomain-randomizationprocess

    Prefer manually staged difficulty with gated transitions; if you use an automatic curriculum, instrument its internal state and alarm when it stops engaging - a saturated curriculum is constant DR wearing a curriculum's name.

    Symptom

    A curriculum mechanism intended to grow difficulty adaptively (s1f's ratchet) hit its cap at iteration 248 and never bit again - for 96% of the run its effect was equivalent to constant DR, i.e. the curriculum existed in name only.

    Context

    When external advice suggested graded wz bands (start ±0.15, then ±0.30), the team agreed with grading but explicitly rejected automatic curriculum, citing the s1f episode. The same logic had already been paid for with push grading: ±0.6 failed twice, ±0.3 was feasible - grading matters, but the grade transitions were made by hand at verified checkpoints.

    Change

    Ladder policy: difficulty staged manually, one band per rung, each transition gated by the acceptance battery; automatic ratchets not used unless their engagement is monitored and demonstrated.

    Outcome

    Every C-ladder band change (wz ±0.12-0.25 first, wider later) was an explicit, attributable rung; no silent constant-DR-in-disguise runs recurred.

    Mechanism

    Adaptive curricula couple their own state machine to noisy training metrics; a ratchet that saturates early stops adapting but keeps its name, so the operator believes difficulty is progressing when it is frozen. Manual staging costs more decisions but each decision is observable and reversible.

    Applies when

    • choosing between auto-curriculum and staged bands for a new skill
    • a curriculum's difficulty parameter plateaus early in training
    • post-hoc attribution of what difficulty a lineage actually saw
    “C2 的 wz 分级(±0.15 → ±0.30,不要一上来 ±0.6)—— 与我们付过学费的 push 分级同形(±0.6 两轮 FAIL,±0.3 才可行)。但必须手动分级,不做自动课程 —— s1f 的课程化棘轮 iter 248 封顶未咬合,96% 时长等价常量 DR”
    train/C_LADDER_RUN.md § 1. 采纳 3 条 (C2 的 wz 分级)
  • Every new penalty ships with a pre-registered withdrawal clause - if healthy gait must pay above the cap, the term stands downcalibration-threshold-with-withdrawal-clause
    Replicatedwalkreward-shapingreward-shapingprocess

    Introduce every new penalty with: the zero-cost-option audit, a replay-calibrated weight formula (healthy pays a fixed small fraction of tracking), and a pre-registered withdrawal condition - and let the clause fire without argument when the calibration says the term cannot be afforded.

    Symptom

    Three same-shaped crashes had established a failure archetype: v4's clearance, v8a's landing window (weight off by 58x uncalibrated), and v6a's bare hip_yaw suppression all combined a zero-cost "don't move" option with a fee on any motion - a reverse barrier that pushes policies toward standing still.

    Context

    The v11 landing-window penalty was therefore introduced under a calibration-threshold protocol: (1) shape chosen with the window tightened (h_gate 0.03 -> 0.02, because 0.03 equaled the clearance target and priced the entire descent); (2) weight from a FORMULA, not judgment: measure the term's raw value on healthy replays (v5/v10b), set w = -(0.10-0.15 x tracking reward) / raw_healthy; (3) withdrawal clause pre-registered: if healthy gait must pay >15% of tracking no matter the tuning, the term is withdrawn to the next version rather than forced in - "不硬上". The companion hip_yaw quieting term ran the same protocol (calibrate on replays, healthy pays <=5%) and was later retired entirely when a structural fix (zero action scale) made its shaping tax unnecessary.

    Change

    Penalty introduction protocol: shape audit (what is the zero-cost option?), replay-based weight formula, healthy-pay cap with a written stand-down condition - all before training.

    Outcome

    The landing term was in fact withdrawn under its clause (v12 records "P5 落地窗口罚 已撤 … 维持撤下"), demonstrating the protocol firing as designed instead of the fourth same-type crash.

    Mechanism

    A penalty's damage mode is mispricing healthy behavior; since the healthy price is measurable in advance on replays, both the weight and the go/no-go decision can be computed rather than discovered by a ruined training run. The withdrawal clause converts "make it work" pressure into a clean deferral.

    Applies when

    • adding any motion-taxing penalty to a working gait
    • a proposed term's weight has no measurement behind it
    • a previous same-shaped term crashed training
    “权重公式而非拍脑袋:先在 v5/v10b 回放上量 h_gate=0.02 的原始值,w = −(0.10~0.15 × 跟踪奖励) / raw_健康;标定门槛:若健康步态无论如何要付 >15% 跟踪,本项撤下留 v12,不硬上 (v4 clearance/v8a-B/v6a 三次同型翻车的教训:代价为零的"不动"选项 + 一动就收费 = 反向壁垒)。”
    train/WALK_V11_SPEC.md § 6. P5 —— 落地窗口罚(三代欠账,标定门槛制)
  • Model CAN polling skew - joint observations are 6-9 ms stale by read ordercan-timing-skew-modeling
    Observed oncewalkactuator-modelingactuator-modelingsim2simhardware

    If joints are read sequentially over a shared bus, reproduce the per-group observation staleness in sim (or randomize it over the measured range) - synchronous observations are a privileged fiction.

    Symptom

    Policies trained with synchronous joint observations degrade on hardware where motors are polled sequentially over CAN - hip data is already 6-9 ms old by the time ankle data arrives.

    Context

    Menlo's core sim2real finding on a leg platform of almost identical mass to Lucen's. CAN is a sequential bus: one poll cycle reads motors in a fixed order, so the observation vector mixes timestamps. Lucen runs 6:6 motors on a dual CAN split, so the problem transfers one-to-one.

    Change

    Explicitly model CAN timing skew in sim: give the joint groups different observation delays matching physical read order. Menlo went further - running real firmware in the loop with a motor simulator between MuJoCo and firmware injecting 0.4-2 ms uniformly distributed delay.

    Outcome

    Reported by the reference team as their core sim2real enabler on a same-scale platform; recorded in Lucen's experience log as directly applicable ("这个问题一模一样").

    Mechanism

    A policy exploits any cross-joint temporal coherence present in training observations; when hardware breaks that coherence per bus position, the learned feedback acts on inconsistent state estimates, injecting phase error exactly at the control bandwidth.

    Conflicts

    Second-hand episode: outcome numbers are the reference team's report, not a Lucen-run experiment; Lucen adopted the requirement but the corpus has no Lucen A/B of skew-on vs skew-off.

    Applies when

    • robot polls actuators sequentially over CAN/RS485 or similar shared bus
    • sim2real degradation appears as jitter or oscillation not seen in sim
    • designing the observation/delay model before a training run
    “电机走 CAN 是顺序轮询的,髋部电机的数据到踝部电机上报时已经陈旧了 6-9 ms,他们直接在仿真里显式建模了 CAN 时序偏斜,按读取顺序给三组关节不同的观测延迟。… 注入 0.4–2 ms 的均匀分布延迟。你们是 6:6 双 CAN 分总线,这个问题一模一样”
    Experience.md § 执行器 + 时序建模 (line 6)
  • The +/-50 mm lateral COM randomization meant to spread the legs coincided with legs pulling IN - rolled back per its own pre-registered contractcom-dr-rollback-on-symptom
    Observed oncewalkdr-tuningdomain-randomizationattributiongate-battery

    When adopting a DR value that covers no local measurement, write its intent and rollback trigger into the config at adoption time; roll it back as the control arm the moment the symptom contradicts the intent, and promote the symptom's metric into the acceptance battery.

    Symptom

    After v7 adopted the reference developer's oversized lateral COM randomization (+/-50 mm) explicitly to force leg spread, the real robot's legs narrowed instead - lateral mean 154 mm / closest 107 mm in sim (nominal 214.5), narrower still on hardware with occasional leg contact.

    Context

    The rollback was clean because the adoption had been honest: the robot.yaml comment recorded the intent AND that the +/-50 value covered no local measurement (only a 16/7 mm measured offset existed; even the prior widening to +/-20 was subjective), plus the reference's own reported side effect (base sway) and the note "这一项要单独跑、 单独归因". When the opposite symptom appeared, v8 returned y to +/-20 mm as the control arm ("要么没起作用、要么帮了倒忙 … 按约定退回做 对照"), kept x/z untouched (a noise-level difference not worth another variable), and named the second suspect: the landing penalty itself, via the reference's own three-link chain (landing penalty -> stance narrows -> spacing penalty needed). A gate lesson was booked in the same table: v7's sim numbers had ALREADY crossed the line (154/107 vs v5's 182/147) - "这个指标本可拦下 v7" - so foot-distance became a standing acceptance row (min >120 mm, zero leg-leg contacts).

    Change

    base_com_offset_m y: 0.050 -> 0.020 (x/z kept), regenerated through the export tool rather than hand-editing derived files; foot-distance acceptance row added.

    Outcome

    A borrowed DR lever with no local measurement basis was retired the moment its symptom contradicted its purpose, at single-variable cost; the metric that would have caught it pre-hardware entered the gate.

    Mechanism

    DR ranges shape behavior through the policy's robustness strategy, which is jointly determined with every reward term; a lever that forces stance width on one robot can be dominated by a stronger narrowing pressure (landing softness) on another. Levers adopted without local measurement must carry their own rollback trigger, because there is no nominal to argue from when they misbehave.

    Conflicts

    Causality is not fully closed in the source: the narrowing may come from the landing penalty rather than the COM lever ("腿距的第二嫌疑人是 ④ 本身"); the rollback is the pre-agreed control experiment, not a verdict that the lever caused the narrowing.

    Applies when

    • importing DR ranges or behavioral-forcing randomizations from references
    • a DR lever's observed effect contradicts its documented purpose
    • a sim metric existed that would have caught a shipped regression
    “⑥ 的本意 … 是逼策略把脚分开;真机结果是脚向内收且偶发相碰——要么没起作用、要么帮了倒忙。… 注释当时就写了"这一项要单独跑、单独归因"。现在症状出现了,按约定退回做对照。”
    train/WALK_V8_SPEC.md § 3. 改动 C — 质心随机化退回(撤销 v7-⑥ 的 y 项)
  • Oversized lateral COM randomization (+/-5 cm) deliberately forces leg spreadcom-randomization-forces-leg-spread
    Observed oncewalkdr-tuningdomain-randomizationreward-shaping

    DR ranges can be behavior-shaping tools, not just robustness padding: oversize a randomization axis to force a strategy the reward struggles to express - and expect a compensating behavior to appear as the cost.

    Symptom

    Feet drift toward the centerline and even collide; policy has no incentive to keep a lateral support base.

    Context

    COM randomization ranges were chosen asymmetrically by axis: lateral +/-5 cm ("比常规大,故意的" - larger than usual, on purpose), fore-aft +/-2 cm, vertical +/-2 cm. The oversized lateral range is not robustness padding but a behavioral forcing function. Lucen logged it as directly relevant to its own roll-channel / sideways leg-kick symptom.

    Change

    Set COM randomization to lateral +/-5 cm, fore-aft +/-2 cm, vertical +/-2 cm, with the lateral band intentionally oversized to make narrow stances fail during training.

    Outcome

    Effective at separating the feet on the reference robot; side effect - the base began swaying left-right, which then required a foot-centerline distance penalty (see reward-chain-foot-height-landing-spacing).

    Mechanism

    Randomizing COM laterally makes narrow-stance policies fall for some draws, so PPO discovers wide stances as the only strategy robust across the band - DR used as an implicit reward. The sway side effect appears because the policy hedges against unknown COM by active lateral correction.

    Applies when

    • feet too close / self-collision in a learned gait
    • roll-axis instability suspected to come from narrow stance
    • choosing COM or mass-offset DR ranges
    “两脚太近甚至互撞 → 先试质心横向随机化 ±5 cm,逼迫策略把脚分开;有效但引发新问题——基座开始左右摇摆 … 横向 ±5 cm(比常规大,故意的,用来逼出分腿)/ 前后 ±2 cm / 垂直 ±2 cm”
    Experience.md § 质心随机化范围 (lines 75, 84-86)
  • Changing the gait clock silently flipped a hardwired threshold's meaning - write derived constants as expressionsderived-constants-must-track-their-base
    Mechanism understoodwalkprocessprocessreward-shapingcontract-freeze

    Before changing any base parameter (clock, control rate, scale), enumerate every constant derived from it and every constant that must NOT change; convert derived literals into expressions of the base so the next change cannot silently flip a term's meaning.

    Symptom

    Slowing the clock 0.40 -> 0.50 s would have silently inverted the feet_air_time threshold's semantics: the 0.25 s threshold was hardwired, so at ct 0.40 the swing window (~0.20 s) sat below it (constant pressure to lengthen strides), while at ct 0.50 the window (~0.25 s) equals it - the term's meaning flips from "push longer" to "neutral" with no code error anywhere.

    Context

    The clock change audit walked every dependent quantity: most followed automatically (joint_pos_ref / clearance / contact_number cycle_time params, gait_phase observation, deploy/sim2sim/policy_io, export) - wiring confirmed, zero hand edits; the air_time threshold was the one hardwired constant, fixed by preserving the RATIO: 0.25 -> 0.3125 = 0.625 x ct, with the recommendation to commit it as the expression 0.625*ct "一劳永逸" (solved once and forever). The same audit also listed what must NOT follow the clock (50 Hz control rate, physics dt/decimation, 47-dim contract, action_latency absolute seconds, PD/torque limits) - the change's blast radius stated in both directions.

    Change

    feet_air_time threshold re-expressed as a fraction of cycle_time; auto-following vs must-not-change lists written into the spec for the clock migration.

    Outcome

    The clock migration (v10, repeated in v11) carried no silent semantic flips; the expression form removed the trap for every future clock change.

    Mechanism

    Constants derived from a base parameter encode a ratio at their birth; storing the evaluated number severs the dependency, so changing the base leaves stale semantics with no failing test. Expressions preserve the intent; and an explicit both-directions dependency list (follows / must-not-follow) is what makes a base-parameter change reviewable.

    Applies when

    • changing gait clock, control frequency, or units
    • a reward threshold interacts with a phase/window duration
    • config audit finds literals that encode ratios
    “feet_air_time 阈值 0.25 是写死的,不跟 ct 走——0.40 时摆动窗 ~0.20s<0.25(恒拉长压力),0.50 时摆动窗 ~0.25s≈阈值(语义翻转)。按比例保原压力:0.25 → 0.3125(=0.625×ct;建议直接写成 0.625 * ct 表达式,一劳永逸)。”
    train/WALK_V10_SPEC.md § 3. T —— 慢时钟 (训练侧必做一件)
  • Exponential tracking kernels go flat exactly when the error is largest - pair them with an L2 term for the far fieldexp-kernel-needs-l2-far-field
    Mechanism understoodwalkreward-shapingreward-shaping

    Never let an exp/Gaussian kernel be the only tracking pressure on a quantity that can drift far from target: pair it with an unbounded (L2) term sized as the "don't diverge" floor, and check which frame the kernel reads.

    Symptom

    With only an exp-type yaw tracking term (exp(-err/std^2), std 0.25), a robot whose heading had drifted badly received almost no corrective gradient: at error 0.6 rad/s the term evaluates to exp(-0.36/0.0625) = 0.003 - near zero AND flat.

    Context

    The exp kernel is excellent for fine tracking near zero error but its gradient vanishes at large error - precisely when correction matters most. Fix: add track_ang_vel_z_err_l2 (-0.5), a plain quadratic on the same quantity: "exp 管精细跟踪、L2 管'别发散', 互补". Both terms deliberately read WORLD-frame wz (matching the exp term's source), because this torso sways enough that body-frame wz means are systematically off (measured -0.039 while actually turning +0.152). The same far-field-gradient argument reappears in the v8 risk list: frozen joints could not climb back because their huge error put them on the exp plateau ("远端梯度消失是冻结自锁的帮凶").

    Change

    Added the L2 companion term at -0.5 alongside the existing exp term (a term that had been in an earlier draft and was lost in a rewrite - itself worth noticing).

    Outcome

    Corrective pressure restored across the whole error range; the exp+L2 pairing became the house pattern for tracking terms.

    Mechanism

    d/de[exp(-e^2/s^2)] -> 0 as e grows: the kernel saturates and cannot distinguish bad from terrible. A quadratic's gradient grows with error, covering the far field; summing the two yields monotone corrective pressure with fine shaping near the target.

    Applies when

    • tracking rewards use exp/Gaussian kernels alone
    • a drifted or frozen state fails to recover during training
    • designing tracking terms for quantities with large transient errors
    “exp 在误差大时梯度趋零, 恰好在最需要纠正的时候失灵。… 误差 0.6 → exp(-0.36/0.0625) = 0.003, 接近零且平坦。… exp 管精细跟踪、L2 管"别发散", 互补。”
    train/WALK_V7_SPEC.md § ② track_ang_vel_z_err_l2 −0.5 —— 补 exp 的梯度洞
  • Design the next run to complete the 2x2 - either outcome then convicts or acquits a factor cleanlyfill-the-missing-factorial-cell
    Mechanism understoodwalkprocessprocessattributionreward-shaping

    When two config factors are jointly suspected, lay out the factorial of existing evidence, spend one run on the missing cell with both interpretations and an early-abort tripwire written in advance - and treat either outcome as a verdict, not a disappointment.

    Symptom

    Hip joints froze at the action clamp in v7, but the history could not say whether the culprit was the raised action_rate (-0.2) or the halved reference amplitude (scale 0.15): existing versions covered only three corners of the (rate x scale) space - v5 (-0.03, 0.30) healthy, v6 (-0.03, 0.15) healthy, v7 (-0.2, 0.15) frozen.

    Context

    v9 was designed explicitly as the missing cell (-0.2, 0.30), with the readings pre-registered: v9 not frozen -> the real anti-freeze force was always the reference amplitude and -0.2 may stay; v9 frozen -> -0.2 is convicted beyond appeal (freezes at both amplitudes) and the next version goes straight to a structural fix. "两个结局都是干净的信息" - both endings are clean information.

    Change

    One training run allocated purely to complete the factorial, with freeze tripwires (joint_pos_ref telemetry <0.1 at iter 1000-1500 -> abort, do not run to 6000) so a conviction costs the minimum compute.

    Outcome

    v9 froze - the rate weight was convicted at both amplitudes ("−0.2 铁案定罪"), and v10 moved to the structural saturation fix with the weight question closed instead of re-litigated.

    Mechanism

    Three corners of a 2x2 leave the two factors confounded in the failure corner; the fourth observation makes each factor's marginal effect identifiable. Pre-registering both readings turns the run into a guaranteed-informative experiment regardless of outcome.

    Applies when

    • two config changes are confounded in a failure
    • version history already covers some corners of a factor grid
    • deciding what single experiment buys the most attribution
    “这恰好补齐一个 2×2 实验矩阵的缺格 … v9 不冻 → 真正的抗冻结主力一直是参考摆幅,−0.2 可以留;v9 仍冻 → −0.2 铁案定罪(两种摆幅下都冻),v10 直接上结构修复 … 两个结局都是干净的信息。”
    train/WALK_V9_SPEC.md § 0. 设计原则 (2×2 实验矩阵)
  • Guessed joint friction was 2.5x low and damping 5x high - measure, then DR around nominalfriction-measured-not-guessed
    Mechanism understoodwalkplant-calibrationplant-calibrationdomain-randomization

    Measure frictionloss and damping separately (they need different rigs), put the measured value at DR center, and express DR as an additive band around that nominal - a DR range around a guessed value can exclude the real robot entirely.

    Symptom

    Old MJCF friction values were invented, not measured; when finally measured, every guessed value was wrong by a large factor in some direction.

    Context

    Joint friction split into Coulomb (frictionloss, tau_c) and viscous (damping b). Measured with the robot hung from a crane (吊机测) while armature was measured no-load; the two measurements are deliberately separated. Old MJCF: frictionloss 0.05, damping 0.1, DR joint_friction range [0, 0.1] "凭空拍的" (made up out of thin air).

    Change

    Replace guessed values with measured ones - frictionloss: RS06 0.15 / RS02 0.12 / RS00 0.13 N*m (old 0.05, i.e. 2.5x too low); damping: 0.02 N*m*s/rad on all three motor types (old 0.1, i.e. 5x too high). DR reshaped from an absolute made-up range [0, 0.1] to an additive band around measured nominal: joint_friction_add [-0.05, +0.10].

    Outcome

    "摩擦定稿(与 armature 一起, plant 参数第一次全部来自实测)" - friction frozen as part of the first fully-measured plant; DR now brackets a measured truth instead of spanning an invented interval.

    Mechanism

    Coulomb friction and viscous damping have opposite behavioral signatures (constant-torque threshold vs velocity-proportional drag); guessing both wrong in opposite directions gives a plant that is simultaneously too easy to start moving and too hard to move fast. DR centered on a wrong nominal makes the policy robust to a family of plants that does not contain the real one.

    Applies when

    • plant friction/damping values have no measurement provenance
    • DR ranges are absolute intervals rather than bands around a nominal
    • policy is over- or under-damped on hardware relative to sim
    “测关节摩擦。吊机测,而armature应该空机测试。… frictionloss τ_c (N·m) │ 0.15 │ 0.12 │ 0.13 │ 0.05(低 2.5×) … damping b (N·m·s/rad) │ 0.02 │ 0.02 │ 0.02 │ 0.1(高 5×) … DR │ joint_friction_add: [−0.05, +0.10] 叠标称 │ 旧 [0, 0.1] 凭空拍的”
    Experience.md § 摩擦定稿表 (lines 12-25)
  • Curriculum-gate a penalty to the phase where its disease occurs - early on it only taxes explorationgate-penalties-to-the-disease-phase
    Replicatedwalkcurriculumcurriculumreward-shapingaction-rate

    For penalties aimed at late-stage pathologies (freezing, saturation, degenerate attractors), ramp the weight in only after exploration noise has decayed; anchor the terminal weight to measured healthy-vs-sick raw values, and shift all related tripwires to after the ramp completes.

    Symptom

    The action_saturation penalty, applied from iteration 0 in v8a, taxed exploration itself: with init_noise_std 1.2 the sampled actions paid ~-2.45/step before any policy had formed - while the disease it targets (clamp freezing) is a LATE pathology (v9 froze at iteration ~2624).

    Context

    v10 re-introduced the same penalty behind a curriculum gate: weight 0 until iter 1000, ramping linearly to -1.0 by iter 2000 - present only when the disease can occur, absent while exploration noise dominates. The trust argument was evidence, not hope: in v8a the term, while active, had pulled joint_pos_ref from 0.041 up to 0.155 and climbing - proof it can extract a policy from the frozen pit. Weight magnitudes were anchored to measured raw values (healthy v5 0.310 / v6 0.106 vs frozen v7 1.145 / v9 1.22 per step: at -1.0 healthy pays 6-18% of tracking, frozen pays 60%+, standing ~0). v10c then isolated the gated term as THE anti-freeze mechanism by single variable, upgraded to untouchable status in v11: "S 的门控机制(v10c 单变量铁案:任何情况下 不许撤,只许调终值)" - and v11 dared to relax other penalties only because S stood guard.

    Change

    action_saturation gated 0 -> -1.0 over iters 1000-2000 (later terminal value tuned -1.0 -> -0.5 with the gate mechanism itself frozen); tripwires adjusted to respect the gate's timing (freeze check moved to iter 2500-3000 to give the ramped term its effect window).

    Outcome

    Freezing stopped recurring while early training kept full exploration; the mechanism graduated from experiment to invariant within two versions.

    Mechanism

    A penalty's incidence depends on who occupies its support: early in training that is exploration noise (whose suppression starves learning), late it is the converged pathology. Time-gating aligns the penalty's presence with its target's presence, buying the constraint without the exploration tax - and tripwire timing must then be computed from the gate schedule, not from ungated precedents.

    Applies when

    • a structural penalty punishes exploration in early training
    • a late-onset pathology (freeze/saturation) needs a standing guard
    • deciding when a curriculum ramp should engage
    “v8a 实锤它的病根是"罚在采样动作上"——init_noise_std 1.2 的早期等于罚探索(~−2.45/步);而冻结是晚期病(v9 速率 2624 才死平)… 门控让它只在病发期在场。… v8a 里它在场时 joint_pos_ref 从 0.041 爬到 0.155 且仍在升——有从低谷爬出的实证。”
    train/WALK_V10_SPEC.md § 2. S 保险 —— action_saturation 课程门控
  • Measure yaw rate by integrating heading, not by averaging body-frame angular velocity - the two differed 15xheading-integral-not-body-rate
    Mechanism understoodwalksim-evalmeasurementsim2simattribution

    For any secular rate (turn gain, drift), integrate the world-frame angle over the window; never average instantaneous body-frame rates during oscillatory motion - and when code comments warn about a measurement, believe them before re-measuring.

    Symptom

    Two measurements of the same turn gain disagreed by a factor of ~15: time-averaged body-frame omega_z gave -0.05 while the sim2sim harness's heading-angle integration gave +0.473.

    Context

    The harness code comment had already documented and predicted the failure: during gait the torso oscillates (body-frame omega_z std up to 0.7); projecting world angular velocity onto a swaying body axis and then averaging biases the estimate systematically - "实测体系均值 −0.04 而实际在以 +0.15 转" (measured body-frame mean -0.04 while actually turning at +0.15). The author's own -0.05 measurement was declared void and the training machine's 1.58/2.45 turn gains confirmed valid.

    Change

    Measurement doctrine fixed: yaw rate for evaluation = net heading change by integration over the window; instantaneous body-frame rates are unusable for averaged directional statistics during legged gait.

    Outcome

    Subsequent friction sweeps and turn-gain accounting were all conducted in the heading-integral currency, making cross-simulator comparisons (MuJoCo vs Isaac 1.04/1.02) meaningful.

    Mechanism

    Averaging a vector quantity expressed in an oscillating frame couples the frame's oscillation into the mean (a rectification bias); the heading integral is computed in the world frame where the gait oscillation integrates to ~zero, leaving the secular component.

    Applies when

    • measuring turn gain, heading drift, or any secular angular rate
    • a body-frame-averaged statistic disagrees with trajectory-level truth
    • writing evaluation code for oscillating platforms
    “我用体坐标系 ωz 的时间均值测,得 −0.05;sim2sim 用航向角积分,得 +0.473。差 15 倍。… 步态中躯干摇晃(体系 ωz std 可达 0.7),把世界角速度投到摇摆的体轴上再取均值会系统性偏掉 … 结论:偏航率必须用航向积分,体系瞬时角速度取均值不可用。”
    train/WALK_DIAGNOSIS.md § ③ 转向增益 —— 我的测法是错的,训练机的 1.58/2.45 成立
  • IMU observation age cut 52-68 ms to ~4 ms by moving AHRS onto the MCU - as a single variableimu-age-move-fusion-downstream
    Observed oncewalkreal-deployhardwarereal-acceptanceprocess

    Audit observation age end-to-end and move time-critical fusion as close to the sensor as possible - and when you fix a latency, change only that one variable so the gain is attributable.

    Symptom

    IMU-derived observations reaching the policy were 52-68 ms old because attitude fusion ran in Python on the loaded host computer - stale attitude is a direct feedback-loop delay the policy was not trained with.

    Context

    The fix was scoped deliberately narrowly: move the AHRS computation from Python to the STM32 H7 (MC02). CAN topology explicitly unchanged, so the change is a clean single variable.

    Change

    AHRS fusion relocated Python -> H7. Before/after - IMU age: 52-68 ms -> ~4 ms; CAN timing: unchanged; Python load: high -> ~0.

    Outcome

    IMU age reduced by an order of magnitude with no confound; host CPU headroom recovered ("把计算单元搬在stm32上, 这样imu有剩余").

    Mechanism

    Sensor age is pipeline latency, not sensor quality: fusing on the MCU next to the sensor removes host scheduling jitter and interpreter overhead from the critical path. Keeping the bus topology fixed makes the improvement attributable to the relocation alone.

    Applies when

    • measured sensor-to-policy age far exceeds sensor sample period
    • attitude fusion or filtering runs on a loaded host CPU in an interpreted runtime
    • planning infrastructure changes during a sim2real campaign
    “AHRS 搬到 H7——这个不改 CAN 拓扑,只是把一段计算从 Python 挪到 MC02,单变量:IMU age 52–68 ms → ~4 ms / CAN 时序 不变 / Python 负载 高 → ≈0”
    Experience.md § AHRS 搬到 H7 (lines 28-35)
  • A quadratic clearance reward stalled for 3500 iters near target - switch to an indicator on accumulated heightindicator-reward-avoids-gradient-decay
    Mechanism understoodwalkreward-shapingreward-shaping

    When a shaped term plateaus near its target, check the gradient profile: replace vanishing-gradient forms with threshold/indicator forms for the final approach, and prefer delta-accumulation over absolute positions to immunize against frame offsets.

    Symptom

    The quadratic-error clearance term froze at -0.002 from iteration 2000 to 5500 - thousands of iterations with no progress on foot lift.

    Context

    Diagnosis: a quadratic penalty's gradient vanishes as the error approaches target, so exactly where the last millimeters must be earned the incentive fades to nothing. The replacement (Humanoid-Gym form): accumulate the swing-phase height climb per foot, reward a BINARY indicator |accumulated - target| < 0.01 masked to the planned swing window, reset on contact, weight +1.6 as a positive reward. Two properties: the indicator's incentive is constant until the threshold is crossed (no decay zone), and accumulating height DELTAS makes any constant sole-frame offset cancel automatically - which structurally sidesteps the earlier 0.0585 m zero-point bug ("顺带绕开我先前那个'忘了减 0.0585 导致惩罚恒为 0'的坑").

    Change

    Clearance reformulated from quadratic penalty on instantaneous height to indicator on per-swing accumulated climb (target 0.03 m by leg-length scaling, weight +1.6).

    Outcome

    Part of the v5 package under which lift finally moved (v5 29 mm, v6 34 mm vs the stalled 18-24 mm era); the offset-cancellation property removed one whole bug class from the term.

    Mechanism

    Policy-gradient learning follows the reward's local slope; quadratic shaping concentrates slope far from target and starves it near target, so convergence stalls precisely at the finish line. An indicator pays a constant bounty until the goal is met; formulating on deltas rather than absolutes removes sensitivity to reference- frame constants.

    Applies when

    • a reward term's value freezes short of target for thousands of iters
    • designing clearance/height/precision terms
    • reward code depends on absolute link positions
    “现行二次型在接近 target 时梯度趋零 —— 这正是 clearance 从 iter 2000 到 5500 卡在 −0.002 不动的原因。… 二值指示在跨过阈值前梯度恒定,没有衰减区 … 累积 delta 让 SOLE_OFFSET 自动抵消”
    train/WALK_V5_SPEC.md § 3. clearance 改峰值型(去掉二次型的梯度衰减)
  • Prove a new penalty actually fires - two ways a clearance term silently did nothinginert-reward-term-audit
    Mechanism understoodwalkreward-shapingreward-shapingplant-calibration

    Before training with a new reward term, log its realized per-step value under the current policy and confirm it is nonzero where intended - check coordinate zero-points against FK and check who occupies the term's gate; and never weaken the term that creates the states your new term needs.

    Symptom

    A newly designed swing-height clearance penalty could have trained as a no-op twice over, and the companion advice to lower feet_air_time actively backfired when tried.

    Context

    Instance 1 (zero-point offset): the proposed code used body_pos_w of the foot link, but that is the ankle_roll_link frame origin, which sits 0.0585 m above the ground even with the foot flat on it - so (0.03 - 0.0585) is always negative and the penalty is永远 0; the 0.0585 offset must be subtracted (verified identical in MuJoCo FK and Isaac). Instance 2 (gate occupancy): the clearance penalty fires only in swing phase; a dragging policy keeps both feet in contact, so the penalty is constantly 0 for exactly the policy it was meant to fix - and worse, any slight lift immediately incurs it, a reverse threshold. Lowering feet_air_time to 0.5 on that advice measurably collapsed air time to 0.0002 (below v3). Corrected understanding: "clearance 是把已有的摆动相抬高, 造出摆动相仍要靠 air_time" - air_time creates the swing phase, clearance raises it.

    Change

    Fixed the height zero-point; kept feet_air_time as the swing-phase creator with clearance layered on top; both errors documented as corrections to the team's own earlier advice.

    Outcome

    With both fixed, swing height rose from 22-23 mm (v2) to 29 mm (v5) to 34 mm (v6); the inert-term failure class entered the standing checklist.

    Mechanism

    A penalty's gradient exists only where its gate is occupied and its argument crosses its threshold; frame offsets shift the threshold out of reach, and phase gates can have zero occupancy under exactly the policy being treated. Terms interact as an ecology - one term must create the states in which another can act.

    Applies when

    • adding any gated or thresholded penalty (clearance, impact, slip)
    • a new term produces no behavioral change at any weight
    • body-frame positions are used in reward code
    “body_pos_w 是 ankle_roll_link 坐标系原点,平放触地时仍高出地面 0.0585 m。… (0.03 − 0.0585) 恒为负 → 惩罚永远是 0 … clearance 惩罚只在摆动相生效,拖地时两脚始终触地 → 惩罚恒 0;而一旦轻微抬脚就立刻扣分,对正在拖地的策略是反向门槛。… 正确认识:clearance 是"把已有的摆动相抬高",造出摆动相仍要靠 air_time。”
    train/WALK_DIAGNOSIS.md § walk_v4 独立验收 — 本文档给的两处代码/建议是错的
  • Torque caps cannot soften footfalls - impact is falling-mass momentum, only the reward can treat itlanding-impact-not-fixed-by-torque-caps
    Mechanism understoodwalkreward-shapingreward-shapinghardwareactuator-modeling

    Classify each hardware symptom by the physics that sets it: quantities fixed by ballistic momentum at contact must be treated through the policy's trajectory (reward terms on approach velocity/force), never through actuator caps - and size such penalty weights against your own tracking reward, not a lighter robot's.

    Symptom

    Footfalls slammed at 1.78x body weight in sim baseline (human walking: 1.2-1.5x); the tempting hardware-side fix was cutting actuator torque limits.

    Context

    Measured directly: scaling torque limits from x1.0 down to x0.4 left peak landing force essentially unchanged (1.75 -> 1.78x body weight) - the impact force comes from the momentum of the falling mass at touchdown, not from motor effort. The fix has to change the trajectory, i.e. the policy, i.e. the reward: feet_contact_forces penalty above a threshold of 113 N (= 1.2x the 9.58 kg robot's weight), clipped, weight -0.005. The weight was sized locally, not copied: the reference robot's -0.001 would amount to 0.9% of tracking reward on this robot ("策略不会理它" - the policy would ignore it); -0.005 gives 4.4%.

    Change

    Added threshold-type contact-force penalty (-0.005, threshold 1.2x body weight) as one of v6-minimal's three changes; hardware torque cuts explicitly rejected as a footfall treatment.

    Outcome

    Landing force 1.72x -> 1.55x by v6 (target <1.5x, missed by 3% - progress booked honestly); the torque-cap dead end was documented so it would not be retried.

    Mechanism

    At touchdown the ground stops a ballistic mass; the impulse is set by approach velocity and effective inertia, which motors can no longer influence in the final instant. Only earlier trajectory choices (approach velocity, timing) reduce it - and those are selected by the reward, not by actuator limits.

    Applies when

    • footfall impact or landing noise on hardware
    • proposals to derate torque as a softness fix
    • importing contact-force penalty weights from another robot
    “⚠️ 硬件限扭降不了落脚力 —— 砸地力来自下落质量的动量: 实测 tau ×1.0→×0.4, 落脚力 1.75→1.78× 体重纹丝不动。只有这条奖励能治。… ⚠️ 权重不能用 Pi 的 −0.001 —— 实测在我们身上只占跟踪奖励的 0.9%, 策略不会理它 (Pi 6.94 kg 更轻)。−0.005 给到 4.4%。”
    train/WALK_V6_MINIMAL.md § ③ 新增 feet_contact_forces
  • A single run's drift direction may be a limit cycle, not a policy bias - check the sign distribution across seedsmultiseed-sign-test-for-drift
    Mechanism understoodwalksim-evalmeasurementattributiongate-battery

    Distinguish "bias" from "broken symmetry limit cycle" by the sign distribution over many seeds; report drift as (mean, sign split), and never compare single-run drift magnitudes across versions.

    Symptom

    Net yaw over 15 s appeared to worsen from -41 deg (v2) to -84 deg (v4), inviting the conclusion that the new version drifted more.

    Context

    The Isaac-side view across 32 environments told a different story: per-env yaw was mixed-sign (20 negative / 12 positive) with mean ~0 - the drift is a limit cycle whose direction depends on initial conditions, not a systematic policy bias. The single MuJoCo run had sampled one draw from that distribution, so its magnitude could not be compared across versions as if it were a property.

    Change

    Evaluation rule: before classifying drift as systematic, run multiple seeds and examine the sign distribution; single-trajectory drift magnitudes are samples, not properties.

    Outcome

    The v2-vs-v4 drift "regression" was reclassified as not-established; later drift work (hip_roll l+r bias) used cross-policy, cross-seed evidence instead.

    Mechanism

    Symmetric dynamical systems can settle into either of two mirrored limit cycles; the selected cycle is decided by noise and initial state. A statistic whose sign is initial-condition-dependent has no meaning as a single sample - only its distribution does.

    Applies when

    • comparing heading drift or lateral drift across policy versions
    • a symmetric-looking behavior shows a consistent direction in one run
    • deciding whether to fix "drift" in reward or calibration
    “偏航反而变差(−41° → −84°):注意 Isaac 侧 32 env 的逐 env 偏航是正负混合(20/12)、均值 ≈0,说明这是极限环性质(方向随初值)而非策略偏置 —— MuJoCo 单次跑测到的是分布里的一个样本,不能当作系统性偏差。要判断需多种子统计。”
    train/WALK_DIAGNOSIS.md § walk_v4 独立验收 读法 (偏航)
  • Privileged signals (true velocity, foot force, foot height) go to the critic onlyobservation-honesty-critic-only
    Mechanism understoodwalkobservation-designobservation-honesty

    Treat the actor observation vector as a hardware contract: every element must exist on the real robot with realistic noise; privileged simulator truths belong in the critic only.

    Symptom

    Policies trained on ground-truth base linear velocity work in sim and fail on hardware, where only a drifting IMU and encoders exist - the policy has learned to depend on a signal that does not survive deployment.

    Context

    Many open-source locomotion stacks feed simulator ground-truth linear velocity to the actor. The reference team refused: the real robot has no ground-truth velocity. Asymmetric actor-critic keeps the training benefit of privileged information without deploying the dependency.

    Change

    Route ground-truth velocity, foot contact forces, and foot heights to the critic only; the actor observes exclusively signals that exist on hardware (IMU-derived quantities, encoders, commands, previous actions).

    Outcome

    Recorded as adopted doctrine in Lucen's experience log; the trained actor's input contract matches what the real robot can actually produce.

    Mechanism

    The critic is discarded at deployment, so it may consume any privileged state to reduce value-estimation variance; the actor's observation set is a deployment contract - anything in it that hardware cannot supply (or supplies with different noise/drift) becomes a train/deploy distribution shift the policy was never trained to handle.

    Applies when

    • designing actor/critic observation spaces
    • reviewing a config where the actor sees base_lin_vel or contact forces
    • sim policy is strong but real robot drifts, oscillates, or falls without obvious actuator cause
    “很多开源代码库把真值线速度喂给策略,Asimov 团队没有,因为真机上没有真值速度,只有会漂的 IMU 和编码器;用完美速度训练出来的策略会依赖它,然后在硬件上失效。真值速度、足底力、足高统统只给 critic”
    Experience.md § 观测空间的诚实性 (line 7)
  • Multiple changes may share one rung only if their symptom spaces are orthogonal - with the ablation order written in advanceorthogonal-batch-with-ablation-order
    Observed oncewalkprocessprocessattribution

    Batch changes into one rung only when you can name each change's private symptom space in writing; pre-register the ablation order (numeric before structural) and per-change escalation plans, so a mixed outcome decomposes without new decisions.

    Symptom

    v8 needed four repairs at once (saturation cheating, landing impact, leg narrowing, heading alignment) - strict one-variable laddering would have cost four training cycles for wounds that were all already diagnosed.

    Context

    The batch was allowed because each change owns a disjoint symptom space, stated explicitly: "A→形态/饱和, B→落地/步态高度, C→腿距/roll 摇摆, D→偏航/转向" - so a single run can attribute each outcome to its change by which symptom moved. For the failure case, the ablation order was pre-registered (A2 -> A1 -> B -> D -> C, "先撤数 值改动" - retract numeric tweaks before structural ones), and every change carried its own escalation/rollback plan (e.g. A insufficient: joint_pos_ref 1.6->2.0 or widen the exp kernel; B overdone - robot afraid to land: v_ok 0.30->0.45; D unstable: rel 0.5->0.25, not back to 0).

    Change

    Four-change rung executed as one run with per-change symptom ownership, per-change contingency plans, and a pre-registered global ablation order for unattributable regressions.

    Outcome

    The rung retained single-run attributability without paying 4x training cost; the contingency table meant no failure mode would require improvising an ablation under pressure.

    Mechanism

    The one-variable rule exists to keep attribution possible, not as an end in itself; attribution survives batching exactly when the changes' observable effects are separable. Orthogonality is a claim that must be argued per pair in advance - and the pre-registered ablation order is the escape hatch for the case the claim fails.

    Applies when

    • several diagnosed fixes are queued and ladder time is scarce
    • deciding between strict laddering and a combined rung
    • a combined rung shows a regression no single change explains
    “四个改动症状空间基本正交,可单 run 归因:A→形态/饱和,B→落地/步态高度,C→腿距/roll 摇摆,D→偏航/转向。出现无法归因的整体退化时消融顺序 A2→A1→B→D→C(先撤数值改动)。”
    train/WALK_V8_SPEC.md § 8. 风险与归因
  • Decompose the offending quantity by channel first - then penalize the failure event, not the jointspenalize-the-slip-not-the-joint
    Mechanism understoodwalkreward-shapingreward-shapingattribution

    Before penalizing motion to fix a side effect, measure which channels actually carry the offending quantity; prefer penalties conditioned on the failure event that are exactly zero for healthy behavior - and do not medicate behaviors that measurement shows are not sick.

    Symptom

    Heading drift with support-foot yaw slip (v5: 212-284 deg accumulated over 15 s); the previous v6 draft had attacked it by penalizing lateral joints (a roll 4.0 / yaw 2.0 "home" group) - which collapsed training into the standing basin.

    Context

    Before choosing the penalty target, the yaw angular momentum was decomposed by joint group with MuJoCo subtree_angmom weighted by real walking joint velocities: pitch-class joints (hip_pitch + knee) carry 95.3%, hip_roll 3.3%, hip_yaw 1.4%. The failed "home" group had been taxing 2.7/step to manage a 4.7% channel. The replacement, feet_yaw_slip (-0.2, |support-foot yaw rate| while in contact), targets the failure event itself and - decisively - costs a non-slipping gait exactly zero, which "横向回家组做不到". The same rung's do-not-do table applied the complementary principle to foot spacing: measured 196-214 mm, stable, no crossing - "没病不吃药" (no disease, no medicine).

    Change

    Removed joint-usage penalties for the drift problem; added the event-conditional slip penalty (-0.2, realized tax 0.141/step = 12% of tracking) alongside the existing linear-slip term.

    Outcome

    Turn-gain left/right difference improved 70% -> 19% and heading 185 -> 60.3 deg by v6 without a standing-basin collapse; the 2.7/step lateral tax never returned.

    Mechanism

    Penalizing joints taxes every use of a channel including healthy use, and if the channel carries little of the offending quantity the tax buys nothing while pushing the optimum toward immobility. An event-conditional penalty (slip while in contact) prices only the failure, leaving the healthy gait's cost surface untouched - and the channel decomposition tells you in advance whether a joint-side fix can even work.

    Applies when

    • choosing a penalty target for drift/slip/impact problems
    • a proposed penalty taxes joints or motions rather than failure events
    • a previous joint-penalty attempt collapsed the gait
    “pitch 类 (hip_pitch + knee) 占偏航角动量 95.3% … hip_yaw 1.4% … 压 hip_yaw 是管 1.4% 的通道收 2.7/步 的税 —— 上一轮正是这样把策略推进了站立盆地。滑移项不惩罚走路: 不打滑的步态代价为零, 这是横向"回家"组做不到的。”
    train/WALK_V6_MINIMAL.md § ① / ② 新增 feet_yaw_slip
  • Size each joint's action authority to its measured working range - lock a channel to zero only when its job is provably elsewhereper-joint-action-scale-lockdown
    Observed oncewalkobservation-designobservation-honestyreward-shapingcontract-freeze

    Set per-joint action scales from measured target ranges and momentum decompositions: full authority for working channels, working-range authority for balance channels, zero for channels whose contribution is measured negligible - and prefer this structural quieting over perpetual reward penalties, keeping the contract dimensions intact.

    Symptom

    "Walks crooked" - roll and yaw channels wandered; hip_yaw peak-to-peak reached 22.3 deg in v11 while contributing essentially nothing to locomotion; a uniform action scale of 0.5 gave every joint the same authority regardless of its actual job.

    Context

    The v12 design replaced the scalar action scale with per-joint scales justified by measurements: pitch-class 0.5 (the gait's entire working channel - untouched); hip_yaw 0.0 - lossless because it carries only 1.4% of yaw momentum (v6 decomposition) and straight-line targets use merely +/-0.006-0.03 ("乱动纯属浪费"); roll 0.2 - NOT zero, because lateral balance and weight transfer are roll's unique job (locking it would degenerate into foot-edge rocking, "比现在更歪"), and 0.2 covers the measured working range +/-0.06-0.19 while "0.5 的另一半全是 '歪歪扭扭'的来源". Costs were accepted consciously: turning demoted to an observation row with a fallback (yaw 0 -> 0.1 in v13). Contract preserved: the 12-dim action interface unchanged, yaw values simply neutralized. The structural lockdown also RETIRED the reward-side hip_yaw_quiet penalty - "A 的 yaw=0 结构性取代,不再付奖励塑形成本".

    Change

    action_scale_joint pitch 0.5 / roll 0.2 / yaw 0.0 wired through robot.yaml -> policy_io (verified: all-ones action gives hip_yaw target exactly 0) -> Isaac action term, guarded by the contract checker ("它就是抓这种双侧不一致的").

    Outcome

    Designed and verified on the shared side before the lineage freeze; stands as the pattern for authority sizing: structure replaces reward shaping wherever a channel should simply not act.

    Mechanism

    Action scale is a per-channel authority budget; uniform budgets give noise channels the same voice as working channels, and reward-side quieting then pays a permanent shaping tax for what a zero scale provides for free. But zeroing is only lossless when decomposition proves the channel's contribution negligible AND no unique function (balance) lives there.

    Conflicts

    Wired and verified on the config/deploy side but never trained - the 2026-08-05 reset suspended v12 before the Isaac-side run.

    Applies when

    • some joints wander without contributing to the task
    • a quieting penalty (deviation/L1) taxes every step forever
    • deciding action-space authority for a new task or robot
    “yaw=0 是无损的:实测它只贡献 1.4% 偏航动量、直行目标只 ±0.006~0.03,乱动纯属浪费。… roll 不能为 0:横向平衡/重心换脚是它的独有职责,锁死会退化成脚缘摇摆(比现在更歪)。0.2 的依据:各代实测 roll 目标只用 ±0.06~0.19,0.5 的另一半全是"歪歪扭扭"的来源。”
    train/WALK_V12_SPEC.md § 3. A —— 逐关节动作幅度(用户"只动 pitch"的安全版)
  • Real robot walked at half the sim clock for two generations - resolved by racing a reward-side and a plant-side evidence line, not by guessingperiod-doubling-evidence-race
    Observed oncewalkattributionattributionactuator-modelingplant-calibrationprocess

    For a hardware-only pathology, refuse to guess: pre-register one probe per side of the sim2real boundary (can the reward mechanism change it on hardware? can fitted plant parameters reproduce it in sim?) and let the first positive result direct the next version.

    Symptom

    The number-one sim2real gap: on hardware v6/v7 stepped at 1.23-1.32 Hz - almost exactly half the 2.50 Hz gait clock they were trained and simulated at; sim never reproduced it, two generations running.

    Context

    Instead of committing training budget to a guess, v8 pre-registered two mutually controlled evidence lines and kept the clock OUT of the training variables: (a) reward-side - if the v8 saturation fix revives joint_pos_ref (the term that pins the gait to the clock), re-run hardware and see whether frequency returns to 2.5 Hz (hypothesis: v7's frozen actions meant NO reward was pinning the gait to the clock, and the real plant - with armature and friction making high frequencies expensive - slid down to the leg's pendulum natural frequency ~1.1 Hz); (b) plant-side - record suspended joint data (fit_actuator), fit armature/friction, load the fitted values into sim2sim and see whether the 1.25 Hz reproduces IN SIM. Decision rule fixed in advance: "谁先给出阳性结果谁定 v9 的方向 (奖励侧 vs plant 侧)" - whichever line goes positive first sets the next version's direction.

    Change

    Period-doubling excluded from the v8 change set; both diagnostic lines scheduled in parallel as non-blocking work; frequency reported factually in acceptance with no pass/fail attached ("倍周期是否消失 不设判定,它是 §9 的关键证据").

    Outcome

    The gap was routed into a decisive-experiment structure rather than a speculative retrain; the plant-side line pointed at exactly the unmodeled armature/friction that were later measured and installed as the plant baseline. Resolution (era-2c full-plant retest): the family had TWO causes - v8's low-speed period-doubling vanished once measured armature+friction were installed (1.30 -> 2.50 Hz, bifurcation-edge machine sensitivity), while v7's stood untouched at 1.20 Hz (saturation-freeze-driven policy property) - both evidence lines paid off, one per case.

    Mechanism

    A behavior appearing only on hardware has candidate causes on both sides of the sim2real boundary; changing training to fix it tests only one side per expensive cycle. Two cheap parallel probes - one intervening on the reward mechanism, one making sim reproduce the real behavior - localize the cause to a side before any training money is spent, and sim-reproduction of a real pathology is itself the strongest form of plant validation.

    Applies when

    • a gait pathology appears on hardware but never in any simulator
    • deciding whether a sim2real gap is reward-side or plant-side
    • tempted to change the gait clock/reward to chase a hardware symptom
    “倍周期(真机 1.23~1.32 Hz ≈ 时钟一半,v6/v7 连续两代;sim 从不出现):两条证据线互为对照——(a)… 真机重跑看频率是否回 2.5 Hz(假说:v7 没有任何奖励把步态钉在时钟上,真机 plant 有 armature/摩擦、高频贵,自由滑落到复摆自然频率 ~1.1 Hz);(b)真机吊挂录 fit_actuator.py … 看能否在仿真里复现 1.25 Hz。谁先给出阳性结果谁定 v9 的方向。”
    train/WALK_V8_SPEC.md § 9. 平行线 (倍周期)
  • Audit rewards by realized contribution (weight x achieved value) - a weight of 2.0 was really paying 0.04realized-contribution-audit
    Mechanism understoodwalkreward-shapingreward-shapingattribution

    Evaluate a reward table by each term's realized per-step contribution under the current policy, never by its weight column; if a term's realized value is ~0, escalating its weight is a no-op - change the term's structure instead.

    Symptom

    Foot dragging persisted through repeated weight escalation: feet_air_time had been raised 0.25 -> 1.0 -> 2.0 across versions with no behavioral change, and the training-side comment even recorded the fact ("几乎没有 单支撑相, 是拖着脚蹭") without the fix changing form.

    Context

    Computing realized per-term contributions in the trained state exposed the economy: feet_air_time contributed weight 2.0 x achieved 0.019 = 0.038 per step, against tracking's +1.20 - lifting the leg earned 3% of what tracking earned, so dragging was the rational optimum no matter the weight escalation. The same table acquitted the energy penalties (total -0.30 negative vs +1.74 positive) that a naive read of weights (-5.0 orientation!) would have blamed.

    Change

    Fix redirected from "raise the weight again" to "add a term whose realized contribution changes the optimum": a clearance penalty sized so its realized magnitude (~0.018/foot when dragging) is comparable to feet_air_time's, enough to flip the optimum without drowning tracking.

    Outcome

    With the term economy corrected (plus posture/range fixes), swing height reached 34 mm and tracking 87% by v6; weight escalation of the old term was abandoned.

    Mechanism

    A reward weight is only a multiplier on whatever the policy currently achieves on that term; when the achieved value is near zero (behavior absent), escalating the weight multiplies near-zero. Optimizer behavior is governed by realized per-step magnitudes, so audits must be conducted in that currency.

    Applies when

    • a behavior persists despite repeated weight increases
    • auditing whether penalties are "too strong" or rewards "too weak"
    • sizing a new reward term against existing ones
    “把 feet_air_time 权重从 0.25 → 1.0 → 2.0 一路加,但没有加高度项。量级算下来:feet_air_time 权重 2.0 × 实得 0.019 = 0.038,而跟踪奖励是 1.2。抬腿的边际收益只有跟踪的 3%,拖地当然是最优解。… 正项 +1.74,负项 −0.30。能量惩罚不是瓶颈,抬腿没收益才是。”
    train/WALK_DIAGNOSIS.md § 决定性证据(#1) / 各项奖励的实际量级
  • When the eval proxy has a known systematic bias, gate on within-proxy differences, not absolutesrelative-metrics-survive-proxy-bias
    Mechanism understoodwalkgate-batterygate-batterysim2simattribution

    Where the evaluator is known-biased, design gates as within-evaluator contrasts (left vs right, A vs B, pre vs post) that cancel common-mode error; reserve absolute thresholds for quantities whose proxy calibration has been checked.

    Symptom

    MuJoCo systematically overestimated yaw turn gain (1.58/2.45 where Isaac measured 1.04/1.02), and friction alignment recovered only part of the gap (mu 1.0 -> 0.6 pulled it to 1.43/2.27) - so any absolute turn-gain acceptance threshold in MuJoCo would be judging the proxy's bias, not the policy.

    Context

    The acceptance criterion was rewritten to use only the left/right difference of the turn gain (standard: <20%): both directions pass through the same biased proxy, so the bias largely cancels in the difference while the policy's chirality - the thing being gated - survives. The absolute-gain row was dropped: "转向只判左右差,不判绝对值". A parallel task was still opened to align MuJoCo contact parameters to Isaac's (mu, restitution), since drift proved highly friction-sensitive (straight-line drift -65 deg at mu 1.0 vs +1 deg at mu 0.4) - bias reduction and bias-robust metrics proceeded together.

    Change

    Gate metric changed from absolute turn gain to left-right gain difference; proxy-alignment work scheduled separately rather than blocking acceptance.

    Outcome

    Turn acceptance became meaningful across proxy versions (v5 70% -> v6 19% difference measured the real improvement) while the absolute bias question was pursued without holding the ladder hostage.

    Mechanism

    A systematic multiplicative or additive proxy bias applies to both arms of a mirrored measurement; differencing (or ratioing) mirrored conditions cancels the common-mode bias to first order, leaving the asymmetry signal. Metrics built this way remain valid while the proxy is imperfect - which it always is somewhere.

    Applies when

    • a sim proxy disagrees with the trainer or hardware in absolute terms
    • writing acceptance thresholds for direction-paired skills
    • proxy calibration work would otherwise block a ladder
    “转向只判左右差,不判绝对值 —— MuJoCo 的偏航增益系统性高估(实测 1.58/2.45 vs Isaac 1.04/1.02),摩擦只能解释一部分(μ 1.0→0.6 仅拉到 1.43/2.27)。差值是相对量。”
    train/WALK_V5_SPEC.md § 6. 验收
  • A field experiment is allowed when it is pre-scripted, single-variable, and self-reversing (RAM-only writes)reversible-single-variable-field-experiments
    Observed oncewalkreal-deployhardwarereal-acceptanceprocessattribution

    Permit hardware-side experiments only when scripted in advance with one variable, a log, and automatic reversion (volatile writes, git restore); keep every config mirror (yaml vs firmware) changed and restored as a unit.

    Symptom

    A hypothesis needed a hardware test - v5's "wild kicking" might trace to the RS00 torque cap being deployed at 11 N*m vs its trained 14 (-21%), on exactly the ankle-roll/hip-yaw joints doing lateral-yaw work - but changing limits in the field is the classic way to lose track of robot state.

    Context

    The experiment was written to be safe by construction: change exactly one number (robot.yaml RS00 tau_limit 11.0 -> 14.0), write to firmware RAM only (--write without --save, so a power cycle automatically rolls back to the saved 12/17/11), run the single logged trial, then restore both yaml (git checkout) and RAM immediately. The consistency requirement is explicit: the deploy tool's torque self-check compares against robot.yaml, so yaml and firmware must change and restore together; and the hypothesis scoping is itself single-variable - RS06's 67% cut was measured irrelevant (gait uses only 16% of rated) and hip torque was left alone for safety.

    Change

    Field experimentation policy refined: not "never touch hardware settings" but "only pre-scripted, one-variable, logged, auto-reverting changes with config/firmware kept consistent".

    Outcome

    The sub-experiment could answer the torque-cap hypothesis without any risk of the robot persisting in an undocumented state - forgetting to restore costs nothing but a git checkout.

    Mechanism

    The danger of field changes is state divergence (robot config drifting from the repo's record), not the change itself; volatile (RAM-only) writes bound the divergence lifetime to one power cycle, and single-variable scoping preserves attributability even in a field setting.

    Applies when

    • a hypothesis requires changing firmware limits or gains on the robot
    • field debugging tempts persistent config writes
    • designing safe escape hatches for deployment tooling
    “程序(--write 不带 --save = 只写 RAM,断电自动回滚)… deploy 的限扭自检是对 robot.yaml 比对的,所以 yaml 和固件必须同改同还原;忘了还原也没事,断电重启即回 12/17/11(上次 --save 的值),但 yaml 要 git checkout。”
    train/REAL_SWEEP_V5_V8.md § 4. 限扭子实验(可选二期,只对 v5,单变量)
  • Reward fixes come in causal chains - foot height, then landing impact, then foot spacingreward-chain-foot-height-landing-spacing
    Observed oncewalkreward-shapingreward-shaping

    Plan reward shaping as a chain, not a point fix: when you patch a degenerate gait behavior, pre-register which adjacent behavior the optimizer will exploit next and watch for it.

    Symptom

    Three problems appeared strictly in sequence: (1) swing feet lifted too low; (2) after fixing that, feet slammed down - "实际比视频里暴力得多" (far more violent in person than on video); (3) after fixing that, feet drifted too close together and collided.

    Context

    Each reward fix removed one degenerate optimum and exposed the next. The fix for foot spacing (COM lateral randomization +/-5 cm to force leg spread) itself caused base side-to-side sway, requiring a further foot-to-centerline distance penalty. Lucen had just solved its own foot-height problem (19mm -> 40mm swing height) and logged landing impact and foot spacing as the predicted next two problems.

    Change

    Chain of additions - (1) penalty when swing foot below 5 cm; (2) landing vertical-velocity penalty at touchdown; (3) COM lateral randomization +/-5 cm, then foot-centerline distance penalty to cancel the induced sway.

    Outcome

    Reference robot progressed through each stage; each individual fix worked and predictably surfaced the successor problem. For Lucen the chain served as a pre-registered roadmap of what breaks next.

    Mechanism

    Locomotion rewards are coupled through contact dynamics: raising swing height adds potential energy that must go somewhere at touchdown (impact); penalizing impact and forcing robustness to COM shifts changes lateral support strategy (spacing/sway). The optimizer always exploits the cheapest unpenalized channel, so fixing one channel routes the exploit to its neighbor.

    Applies when

    • adding a foot-height / clearance reward
    • feet slam or landing impact grows after a clearance fix
    • feet converge toward the centerline or self-collide
    • any single-reward fix to a coupled gait behavior
    “抬脚太低 → 加惩罚:摆动足低于 5 cm 就扣分 / 加完之后砸脚 → 抬起来了但落地极猛,"实际比视频里暴力得多" → 加落地速度惩罚 … / 两脚太近甚至互撞 → 先试质心横向随机化 ±5 cm … 有效但引发新问题——基座开始左右摇摆 → 再加足-中心线距离惩罚 … 这三条是串联的:每个修复都会暴露下一个问题。”
    Experience.md § 三个问题的解法链 (lines 72-77)
  • Training-log reward values and fixed-command eval values live on different distributions - comparing them once claimed a 44% improvement that was really 6-10%same-distribution-reward-comparison
    Mechanism understoodwalksim-evalmeasurementattributionprocess

    Quote reward-term values only with their distribution attached (command range, DR on/off, environment), and compare across runs only when those match; re-measure in a common environment before claiming any improvement percentage.

    Symptom

    A v6-era analysis concluded slip had dropped 44% by comparing the training log's Episode_Reward against values calibrated in a fixed-command play environment; a same-condition re-measurement showed the true improvement was 6-10%.

    Context

    The training log's reward is an expectation over the training command distribution (vx 0.15-0.5, yaw +/-0.6, with pushes and domain randomization); play-environment calibrations are taken at a single fixed command with DR off. Subtracting one from the other compares apples to oranges - the warning was written into the v7 pre-flight: "奖励数值只能在同一指令分布下比较 … 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次)".

    Change

    Rule adopted: any before/after reward-term comparison must hold the command distribution, DR state, and evaluation environment fixed; training-log values compare only against training-log values of runs with identical command/DR configs.

    Outcome

    The phantom 44% improvement was retracted; later term-level accounting (e.g. the C4 ignore-floor work) consistently specified its distribution before quoting numbers.

    Mechanism

    A reward term's expectation depends on the visited-state distribution as much as on the policy; changing the command distribution or DR moves every term's baseline. Cross-distribution differences therefore measure the distributions, not the policy change.

    Applies when

    • comparing reward telemetry across training runs or vs play evals
    • claiming improvement percentages from training logs
    • term-level reward accounting for diagnosis
    “奖励数值只能在同一指令分布下比较。训练日志的 Episode_Reward 是在训练指令分布上算的(vx 0.15~0.5 / 偏航 ±0.6 / 带推力与域随机化), 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次: 据此以为滑移降了 44%, 同条件对拍只有 6~10%)。”
    train/WALK_V7_SPEC.md § 3. 开训自查 ⚠️
  • Cross-simulator gate (Isaac Lab to MuJoCo) comes before any hardware attemptsim2sim-gate-before-sim2real
    Replicatedwalksim2sim-gatesim2simprocessgate-battery

    Gate every policy through a second simulator with an independently built plant before hardware; treat sim2sim failure as a contract or overfitting bug, and sim2real failure after a sim2sim pass as a plant/actuator gap.

    Symptom

    A policy that only ever ran in its training simulator carries untested dependencies on that simulator's solver, contact model, and defaults; the first place those dependencies surface should not be the real robot.

    Context

    Standing order of operations for the whole Lucen program, recorded as the opening line of the experience log: train in Isaac Lab, gate in MuJoCo, only then go to hardware. The MuJoCo side is the same plant used for evaluation batteries, so a sim2sim pass also validates the exported policy + contract (obs ordering, scales, defaults) outside the training stack.

    Change

    Pipeline rule adopted - every checkpoint must pass the MuJoCo evaluation battery (sim2sim) before it is considered for real deployment (sim2real).

    Outcome

    Used generation after generation as the cheap filter; hardware sessions only ever received policies that had already survived a second simulator.

    Mechanism

    Two simulators disagree exactly where a policy is overfit to simulator-specific artifacts (contact softness, integrator, default parameters, obs conventions); a cross-sim transfer catches contract bugs and solver overfitting at zero hardware risk, so real-robot failures that remain are attributable to genuine plant/actuator gaps.

    Applies when

    • planning the path from training to first hardware trial
    • exported policy behaves differently outside the training framework
    • triaging whether a real-robot failure is contract vs plant
    “先sim2sim - 从isaaclab 到mujoco / 再sim2real”
    Experience.md § opening lines (1-2)
  • Borrow reward values from other robots as ratios (to tracking weight, to leg length) - never as absolute numberstransfer-ratios-not-absolutes
    Mechanism understoodwalkreward-shapingreward-shapingprocess

    When importing any numeric from another robot's config or paper, identify its natural normalizer (tracking weight, leg length, sqrt(g*L), body mass) and transfer the dimensionless ratio; sanity-check against a same-scale robot when one exists.

    Symptom

    Published configs offered tempting absolute values (swing height target 0.06-0.08 m, weight -20) that would have been wrong for a robot with half the leg length and a different tracking weight.

    Context

    The cross-check normalized before transferring: G1's feet_swing_height weight -20 against tracking +1.0 is a 20x ratio, so with local tracking at 1.5 the equivalent is -30, not -20. G1's 0.06 m target on a ~0.70 m leg scales to ~28 mm on the local 0.325 m leg (Humanoid-Gym converts to ~23 mm), confirming the locally chosen 0.03 m and explicitly rejecting copying 0.05-0.08 absolutes. A same-class robot (Menlo's 16 kg) was used to sanity-check feet_air_time (+0.5 vs the local 2.0, flagged over-high). A nondimensional check of the same kind later validated the sidewalk speed target (v/sqrt(gL) = 0.081 vs hardware-verified 0.101 - 0.104 - inside the envelope, conservative).

    Change

    All borrowed values converted through ratios (weight/tracking-weight, height/leg-length, dimensionless speed) before entering the config.

    Outcome

    The scaled values worked (0.03 m target matched both the scaling law and measured 22-23 mm baseline); no cross-robot absolute was ever copied raw.

    Mechanism

    Reward economies are scale-relative (only ratios to the tracking term matter to the optimum) and kinematic quantities are morphology-relative (clearance scales with leg length, speed with sqrt(g*L)); absolutes encode the source robot's scale, ratios encode the design intent.

    Applies when

    • copying reward weights/targets from open-source configs or papers
    • setting clearance heights, speed targets, or impact thresholds
    • comparing your weights to published tables
    “G1 的 feet_swing_height 是 tracking 的 20 倍(−20 vs +1.0)。我们 tracking 是 1.5,按同比例应为 −30 … G1 目标 0.06 m / 腿长 ~0.70 m,换算到我们 0.325 m 腿长约 28 mm;Humanoid-Gym 换算约 23 mm。故 target 取 0.03 m 是对的 … 不必抄 0.05~0.08 的绝对值。”
    train/WALK_DIAGNOSIS.md § 修正 ②(权重放大) / 修正 ③(目标高度按腿长缩放)
  • The shipped checkpoint was chosen by scanning checkpoints on the full gate - neighbours 100 iterations apart failed 1 and 38 cells, late checkpoints degraded - never by taking the last one, and training stopped on signals, not on a schedulecheckpoint-choice-is-a-full-gate-scan
    Replicatedonelegsim-evalfork-selectiongate-battery

    Choose a release checkpoint by running the full acceptance battery over a band of checkpoints (including the transfer axis), stop training on measured signals rather than a fixed iteration count, and expect adjacent checkpoints to differ sharply.

    Symptom

    Gate results moved sharply and non-monotonically between checkpoints of the same run, and the last checkpoint was often not the best.

    Context

    One-leg V0r1: the 2,000 neighbourhood was best; from 2,500 on the nominal gates degraded (late overtraining); 2,000 itself had one real micro-hop (17.7 mm over 5 frames); 2,300 was all green and shipped. V0r2: failed cells per checkpoint 2,000:19, 2,100:38, 2,200:1, 2,300:3, 2,400:27, 2,500:12, 3,000:18 - 2,200 shipped. The recovery line learned the same from the other side: stopping v2_6 early at a scheduled point left a policy whose re-fall rate had spiked to 9-22% before consolidation healed it ("stop on signals, not on the schedule"), and a continuation's transfer decayed checkpoint by checkpoint while Isaac stayed perfect.

    Change

    The acceptance rule "scan checkpoints, do not look only at the last one" is written into the one-leg gates (called the S1 discipline); release candidates are chosen from the scan.

    Outcome

    Both one-leg releases were mid-run checkpoints (2,300 and 2,200) chosen by the full 40-cell battery.

    Mechanism

    PPO keeps changing the policy after the gates saturate; with no gradient toward the gate's conditions, later checkpoints wander, so gate quality is a noisy function of iteration.

    Applies when

    • picking which checkpoint of a run to export and stamp
    • a run is stopped at a fixed iteration budget
    • final-checkpoint results are worse than mid-run smoke tests
    “Isaac 侧 S1 纪律: 验收扫 checkpoint,不是只看最后一个。 … 扫描判决: 2000 邻域最优——2500+ 标称面退化(⑤③② 散挂, 晚期过训), 2000 有一例真微跳(L s100 μ1.2, 17.7mm/5帧), 2300 全绿。”
    git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §6 验收门 / §8 核查单 5 与 7
  • Adapting a lineage to one plant increment needs hundreds of iterations, not thousands - long runs only buy specializationcontinuation-budget-not-from-zero
    Mechanism understoodomnitraining-runcurriculumprocess

    Budget continuation rungs by increment class (hundreds of iterations for plant pins and smooth shifts, ~1000-1500 only for behavior-demanding changes like push), enforce a hard cap with frequent evaluation, and treat remaining budget as a reason to stop, not to continue.

    Symptom

    The default "6000 iterations per rung" (a from-zero-scale budget) was about to be applied to continuation rungs whose only change is one plant/DR increment - overspending compute and, worse, giving each rung thousands of iterations to specialize away retained skills.

    Context

    The 2026-08-07 budget table replaced the default with "最低适应窗口 + 每 100 iter 验收 + hard cap" scaled to the increment's difficulty: fixed-latency levels 300-500 (cap 500-800; the base has already seen in-band values, this only pins the plant); PD full-band 700 (cap 1000; kp+/-20%/kd+/-30% clearly widens the actuator family); COM +/-20 mm 500 (cap 800; a smooth dynamics shift); friction DR 700 (cap 1000; contact and actuator friction change the gait/contact solution together); push 1000 (cap 1500; a non-static plant change requiring recovery behavior - hardest). Rationale: "续训适应一个 plant 增量不需要从零量级的预算,跑长了只是给特化时间". The C ladder reused the scheme (per-rung caps 500-2000 by increment type), and the deep-training hazard got its own name when long runs sold quality ("深适应卖质量" - deep adaptation sells quality).

    Change

    Per-rung iteration budgets set by increment class with hard caps and 100-iter watch loops; checkpoint selection inside the window by the smoke curve, never "run to cap because budget remains".

    Outcome

    S2/C rungs completed in 300-1500 iterations each; the recurring late-run degradations (collapse valleys at 1500+, vx+0.30 decay) fell outside most rungs' caps instead of inside their runs.

    Mechanism

    A continuation rung's learning problem is local robustification around an existing optimum - low sample complexity; iterations past adaptation are spent sharpening onto the current distribution, which is exactly how retained skills and margins erode. Budgets sized to the increment bound both compute and the specialization damage window.

    Applies when

    • planning iteration budgets for a robustification or command ladder
    • a continuation run keeps improving its training metric late
    • retained skills decay in the back half of long continuation runs
    “「最低适应窗口 + 每 100 iter 验收(watch_ckpt --every 100)+ hard cap」—— 续训适应一个 plant 增量不需要从零量级的预算,跑长了只是给特化时间 … ⑥ push | 1000 | 1500 | 非静态 plant 变化,要学 recovery 行为,最难”
    train/OMNI_V0_SPEC.md § 4. 每级 iter 预算(2026-08-07 用户定)
  • A DR tail the robot never has is pure cost - stage deterministic plant levels instead of one wide uniformdr-tail-plant-continuation
    Mechanism understoodomnidr-tuningdomain-randomizationcurriculumactuator-modeling

    Set every DR range from the measured deployment distribution and cut tails that hardware cannot produce; when an axis changes the controller's character (delay, major gain regimes), prefer staged deterministic levels with gates over one wide uniform.

    Symptom

    Two consecutive lineages (s1e, s1f) trained under uniform latency DR (0, 0.06 s = 0-3 frames) both converged to drag-glide gaits - buying survival under heavy delay by giving up swing (3.6 mm) - even though the real pipeline never exceeds ~2 frames.

    Context

    The account: roughly 1/3 of training quality was spent on the >2 frame tail that hardware never presents ("uniform 尾部 ~1/3 训练质量 花在真机不出现的 >2 帧上"). The deeper reading came from the user: uniform 0-3 frames is not merely tail-heavy - it "把性质不同的控制系统 混进同一 PPO batch" (mixes qualitatively different control systems into one PPO batch); a 0-frame and a 3-frame plant demand different controllers, and one policy trained on the mixture serves neither. The S2 v2 ladder therefore redefined latency "从「随机化参数」重新定 义为 actuator/control plant 的一部分": deterministic FIFO levels, staged 1 frame then 2 frames (lo=hi so fractional interpolation degenerates to exact N frames, synonymous with the harness --delay N), each level gated by the fixed acceptance battery - a plant continuation, not a randomization.

    Change

    Latency DR replaced by staged deterministic levels covering the measured 1-2 tick reality with no tail; each stage a separate continuation rung with the standard gate and rollback.

    Outcome

    The s2_lag1 rung showed the clean-signal benefit immediately (survival 20/20, heading 6x recovery) with the swing cost booked honestly (21 -> 12 mm, half-pass, ladder paused for adjudication); the drag-glide attractor from uniform tails did not recur.

    Mechanism

    DR asks one policy to cover a plant family; when part of the family is fictitious, the policy pays real capability for fictitious robustness, and when family members demand structurally different controllers, gradient averaging produces a compromise controller optimal for none. A measured, discrete plant set matches the actual deployment support and keeps each rung's training signal coherent.

    Applies when

    • policies converge to degenerate gaits that buy worst-case survival
    • a DR range extends well past the measured hardware range
    • choosing between wide randomization and a staged ladder on an axis
    “两轮实证(s1e/s1f)宽尾延迟 DR 逼出拖地滑行 … uniform 0~3 帧不止尾重,而是把性质不同的 控制系统混进同一 PPO batch;1→2 帧确定性分级 = plant continuation,训练信号干净得多—— latency 从「随机化参数」重新定义为 actuator/control plant 的一部分。”
    train/OMNI_V0_SPEC.md § 4. v2 阶梯 (2026-08-07 用户定)
  • A proposal in the runbook - torque and action limits as versioned safety tiers (classroom / research / expert) written to motor RAM and read back, separate from the reward's effort penalty - recorded as a proposal, its implementation unrecordedsafety-limits-are-a-layer-not-a-reward
    Hypothesisinfrareal-deployhardwareprocessreal-acceptance

    Keep hardware limits as an explicit, versioned safety layer (tiers written and read back at start, the persisted default the safest one) and the effort penalty as a behaviour layer; when a skill needs more torque, change tier deliberately rather than trading one layer against the other.

    Symptom

    Running and jumping need more torque than the deployed limits allow, and the temptation is to trade the training-side effort penalty against the hardware limit, or to hand a new user a robot "tuned however the last person left it".

    Context

    A message pasted into the operator runbook (undated, citing Berkeley's practice of storing the full motor configuration as JSON with write and read-back scripts) proposes: configuration is a versioned artifact, not a verbal agreement; three safety tiers in robot.yaml beside the gain and policy profiles - classroom (RS06 limited to 10 N*m, lateral joints clamped: "however bad the policy, it only moves awkwardly"), research (14 N*m, clamps at twice the measured need, the default) and expert (the 36 N*m rating, joint limits only, requiring an explicit flag); deploy writes the tier to motor RAM at start and reads it back, while the stored copy stays classroom so a power cut returns to the safest state. It frames limits as the safety layer and the effort penalty as the behaviour layer - more torque for running means switching tier, not weakening the penalty.

    Change

    None recorded: the message ends by asking which to do first, a rollback or the tiers.

    Outcome

    The sources do not record the tiers being implemented; the deployed limits stayed at 12/17/11 N*m through the recovery and one-leg lines (the one-leg spec treats raising the RS06 limit as a separate, unapproved hardware decision). Related and recorded elsewhere: torque limits were written to RAM only in a scripted, self-reversing field experiment.

    Mechanism

    Hardware limits bound the damage any policy can do; reward terms shape what a policy prefers. Mixing them either weakens safety to buy behaviour or distorts behaviour to buy safety.

    Applies when

    • a new skill needs more torque than the deployed limits
    • robots are handed to students or new users
    • motor configuration lives in people's heads or in the firmware only
    “配置是版本化的产物,不是口头约定。 … deploy_policy 启动时按档写进电机 RAM 并读回校验(落盘的那份永远保持 classroom,断电自动回到最安全状态)。 … 限幅是安全层,dof_torques_l2 是行为塑造层,它们在不同的层,不冲突。跑步要更大力矩就换档,而不是去动训练里的省力惩罚。”
    RL系统/FOLLOW THIS copy 2.md § 面向 developer / 教育机构该怎么做 (pasted proposal, undated)
  • A 2-degree joint-zero calibration fix moved the whole runnable envelope - re-test old "cannot run" verdicts after recalibrationzero-offset-calibration-shifts-envelope
    Observed onceomnireal-deployreal-acceptancehardwareplant-calibrationattribution

    Date every hardware verdict with the calibration state; after any zero/mount recalibration, re-test previously condemned policy-power combinations and previously "unexplainable" posture offsets before attributing either to training or model.

    Symptom

    s1d was on record as "only runs at power 0.7" (kicked wildly at 0.8); after a calibration pass, the same policy ran 12 s at 0.8 with no kicking at all.

    Context

    The calibration had fixed a 2.08 deg zero offset on r_hip_roll - exactly the constant error source on the dominant joint of the kicking oscillation loop ("恰是乱踢振荡环主导关节的常值误差源"). The three-generation post-calibration hardware sweep also closed a second case: the robot's mysterious "backward lean" disappeared after calibration, and the sim-real posture difference collapsed from opposite-sign 5+ deg to same-sign ~2 deg ("后仰案实质了结") - the lean had been a sensing/zero artifact, not a mass-model error. Booked consequence: if the s1d recovery re-verifies, "真机可跑档整体 上移" - every policy's runnable power envelope shifts up, and downstream lineages' hardware expectations get revised.

    Change

    Joint-zero and mount calibration promoted from setup chore to a variable that dates hardware verdicts: verdicts about which power/scale levels a policy can run are conditioned on the calibration state they were measured under.

    Outcome

    One policy rehabilitated at a higher power level; one standing sim-real posture discrepancy closed without touching model or training; a pending re-verification booked rather than asserted.

    Mechanism

    A constant joint-zero error acts as a persistent disturbance injected at the feedback loop's most-loaded joint; near an oscillation threshold, removing a 2-degree bias is the difference between a stable and an unstable loop. Since the error is additive and machine-side, it shifts every policy's stability envelope simultaneously - which is why verdicts must carry their calibration date.

    Conflicts

    The s1d rehabilitation awaited one confirming re-run at the time of writing ("待复核一跑坐实") - the offset-as-cause reading is the head suspect, not a closed verdict.

    Applies when

    • a policy oscillates at a power level others tolerate
    • sim and real disagree on a constant posture offset
    • deciding whether to re-test old hardware verdicts after maintenance/calibration
    “发现①:s1d@0.8 能跑了(旧账「只有 0.7 能跑」)——12s 无乱踢。头号嫌疑 = 标定修正:r_hip_roll offset 修 2.08°,恰是乱踢振荡环主导关节的常值误差源。… 发现②:「后仰」标定后消失 … sim-real 姿态差从反号 5°+ 收敛到同号 2°,后仰案实质了结。”
    train/README.md § 真机 @0.8 三代横评(2026-08-07 标定后)
  • A binary reward band on the swing knee had zero gradient everywhere below it, so the one-leg policy parked in an unloaded "fake touchdown" that Isaac's 5 N threshold scored as success and MuJoCo showed as real pressing - a capped constant-gradient ramp, retrained from scratch, passed 40/40binary-band-reward-fake-touchdown
    Mechanism understoodonelegreward-shapingreward-shapingsim2sim

    Shape approach-to-target rewards as capped ramps with gradient from the starting posture, never as bands or indicators; and compare contact-based terms across simulators, because a policy riding just under a force threshold looks perfect in one and wrong in the other.

    Symptom

    At iteration 1,000 of the first one-leg run the swing foot never lifted: the policy stood with the "raised" foot resting lightly on the ground. In Isaac the contact-match term paid 96% of full marks; the same policy in MuJoCo pressed that foot on the ground for 450 frames.

    Context

    The swing-leg goal was "shank folded fully back" (knee 1.5-1.95 rad), rewarded as a binary band: +0.8 inside [1.5, 1.95], zero elsewhere. From knee 0.05 to 1.5 rad the term was flat. Contact is judged at a 5 N force threshold, so a foot carrying less than 5 N counts as lifted. The walk line had hit the same disease with a binary indicator (v4) and fixed it with a capped ramp (knee_swing_amplitude).

    Change

    swing_knee_fold changed from the binary band to a ramp clamp(|q|/1.5, 0, 1) - a constant gradient capped near 86 deg - and the policy was retrained from scratch (V0r1). After the first real-robot try showed the fold still too low, its weight went 0.8 -> 2.0 (V0.1).

    Outcome

    V0r1 model_2300 passed the full acceptance 40/40 (swing knee 1.72 rad, about 98.5 deg) and was stamped as oneleg_v0.onnx; the cross-simulator disagreement is recorded as the thing that caught the cheat.

    Mechanism

    A reward that is flat until the target is reached gives no gradient to approach it, so the policy settles for the nearest state other terms reward - here, a foot that satisfies the contact threshold without lifting; a second simulator with different contact force resolution exposes such threshold-riding.

    Applies when

    • rewarding a posture target with an in-band / out-of-band indicator
    • a contact threshold decides whether a foot counts as lifted
    • trainer-side contact terms are near full marks while the video looks wrong
    “初版二值带 [1.5,1.95] 在膝 0.05→1.5 全程零梯度,策略停在"卸力虚点地"(Isaac 5N 阈下 contact_match 96% 满分 / MuJoCo 同策略 450 帧实压——跨仿真器互证抓作弊);v4 二值指示同型病,按 knee_swing_amplitude 判例改常数梯度封顶 ramp,从零重训 … **oneleg_v0.onnx = V0r1 model_2300, 40/40 PASS**”
    git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §4 奖励表 swing_knee_fold 行 / §8 核查单 5
  • The run policy never left the ground and fell in the second simulator from the frontal plane - its DR (gains and latency only) covered the actuator axis, not the frontal-plane contact and inertia disturbances the doubled stride amplified; "is DR on" is the wrong questionthin-dr-judged-by-channel-coverage
    Replicatedrundr-tuningdomain-randomizationsim2simattribution

    Judge a DR recipe by whether its randomized terms cover the channel where the skill can lose stability, not by whether DR is enabled; when a new skill lengthens single support or enlarges motion in one plane, add disturbances in the plane it destabilizes before training.

    Symptom

    run R1 (6,000 iterations, 78 min): no flight phase ever appeared, and every one of 13 checkpoints failed the eight-gate MuJoCo smoke. In Isaac: zero terminations in 6,000 iterations, 4.2 deg tilt. In MuJoCo at delay 2: 1/6 survived, falls within 1.9-6.2 s at 50.8-58.7 deg, the most saturated joints all roll joints.

    Context

    The run contract doubled sagittal travel (knee action scale 0.9, knee swing peak 1.14 rad) with a 0.60 s period and 0.40 duty - long single support - while roll/yaw scales were deliberately left at 0.5. DR copied the s1e recipe: kp/kd (0.9, 1.1) and latency on; mass, COM, joint friction and push all off; ground friction pinned at (1.0, 1.0). Flight was read two independent ways: Isaac's per-foot contact reward stayed 0.845-0.857, never above 0.87 - the arithmetic ceiling of a gait with zero flight - and 30 of 36 MuJoCo seeds had flight fraction exactly 0 (the nonzero six were all tumbling falls). Foot lift itself worked (46-59 mm against a 50 mm design point): the walk-era "not enough travel" failure did not recur.

    Change

    Verdict FAIL, with the pre-registered first knob (exploration noise 1.0 -> 1.2) explicitly rejected as aimed at a different axis. The lesson was generalized and applied at the next line's design review: the one-leg spec made push, body mass, base COM and friction DR mandatory for its permanent single support and banned the thin recipe.

    Outcome

    The run line did not continue past R1 in the sources. The one-leg V0 with the wider DR passed its friction-variant gate (mu 0.4 and 1.2) inside a 40/40 acceptance.

    Mechanism

    Randomizing gains and latency covers the actuator's axis; a skill whose failure lives in frontal-plane contact and inertia needs randomization on that channel (push, mass, COM, friction), or the trainer's exact plant becomes the only one the policy can stand on - the omni_s1 transfer trap a second time, this time with DR switched on.

    Applies when

    • a policy is flawless in the trainer and falls immediately in a second simulator
    • reusing a DR recipe from a skill with a different support pattern
    • failures concentrate on one axis (roll, yaw) the DR does not touch
    “**机理**: 矢状面行程翻倍 (膝摆动峰 1.14 rad) + T 0.60 + duty 0.40 的长单支撑, 把额状面扰动放大了一个量级; 而 roll/yaw 通道按 §3 **刻意没有放大** (仍 0.5), DR 又是 s1e 复刻的薄配方 (mass/COM/关节摩擦/push **四关全关**, 地面摩擦钉死 (1.0, 1.0))。 … 说明**薄 DR 的判据不能只看"有没有开 DR"**, 要看**开的那几项 是否覆盖失稳所在的通道** —— kp/kd 与延迟是执行器轴向的, 对额状面接触/惯性 扰动零覆盖。 … 0.87 正是「零腾空的走路步态」的天花板算术”
    git:Lucen V2@origin/run-line:train/README.md § run R1 FAIL (2026-08-09, run 21-30-30_run_r1): 腾空零, 但病根在额状面不在探索