Skip to content

Training Coach

The coach reads a training run and returns a diagnosis, in-whitelist proposals and an experiment plan — and every claim cites one of the cards below. This is that corpus: 22 methodology rules distilled from a real sim2real programme, and 171 episode cards behind them, each with the sentence in the war history it came from.

Doctrine

A report may cite any of these as doctrine-N.

  1. doctrine-1Contract freeze and fingerprint discipline

    The policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.

    Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).

    Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.

    contract-freeze-and-checkerlegacy-profile-pinningderived-asset-staleness-checkgain-profile-belongs-in-the-stamprecovery-two-policies-and-a-state-machinewalk-recovery-fsm-handoff

  2. doctrine-2Attribution by resolved training params - never eval-override knobs

    Capability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.

    Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).

    Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.

    kd-bandwidth-mu-law-attributiondeploy-knob-attribution-before-retrainingcycle-time-override-is-oodresolved-config-is-source-of-truth

  3. doctrine-3PASS gates become constraints; FAIL gates become objectives

    Once a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.

    Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.

    Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".

    fixed-acceptance-matrix-per-rungpreregistered-stop-criteria-per-rung

  4. doctrine-4One variable per ladder rung - counted against what the checkpoint saw

    A rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.

    Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.

    Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.

    resume-state-dr-auditorthogonal-batch-with-ablation-orderfill-the-missing-factorial-cell

  5. doctrine-5Pre-register risks, readings, and stop criteria before the ladder

    Before a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.

    Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).

    Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.

    preregister-risks-and-fork-readingspreregistered-stop-criteria-per-rungpreregistered-real-expectationsfeasibility-accounts-lock-design-point

  6. doctrine-6Plant parameters are measured, never invented

    Every plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.

    Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).

    Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.

    friction-measured-not-guessedarmature-n2-rotor-inertiatorque-limit-shape-by-measured-peakslatency-lerp-reverse-extrapolationfeasibility-accounts-lock-design-pointsim-api-friction-columnsget-up-feasibility-accounts-before-trainingsingle-support-gain-authority-probe

  7. doctrine-7Sim2sim gate before sim2real - under deployment conditions

    Every checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).

    Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).

    Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.

    sim2sim-gate-before-sim2realeval-plant-honesty-contact-paramspipeline-latency-is-plant-not-drbody-frame-velocity-api-audittorque-penalty-bought-by-leg-bracingtorque-disagreement-between-simulators-unresolved

  8. doctrine-8Observation honesty - the actor's inputs are a hardware contract

    The actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.

    Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).

    Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.

    observation-honesty-critic-onlyhistory-obs-needs-plant-variationreward-observability-limitdeploy-heading-loop-and-align-training

  9. doctrine-9Reward economics are audited in realized currency

    Reward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.

    Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).

    Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.

    realized-contribution-auditreward-cost-of-ignoring-auditgate-new-reward-terms-by-commandignore-floor-diagnosiscalibrate-threshold-between-healthy-and-sickcalibration-threshold-with-withdrawal-clauseinert-reward-term-auditseated-basin-dead-exp-kerneltail-torque-needs-hinge-on-computed-demand

  10. doctrine-10The zero-cost option must be the desired behavior

    For every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.

    Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).

    Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.

    penalize-the-slip-not-the-jointsaturation-cheating-zero-rate-costmoving-gate-42x-stand-taxtermination-closes-degenerate-basinpenalty-gate-is-an-escape-hatchsoft-limit-penalty-charges-nominal-poseunpriced-foot-attitude-is-a-free-variableenumerate-cheapest-cheats-before-trainingbinary-band-reward-fake-touchdown

  11. doctrine-11Measurement discipline: independent referees, signs, distributions

    A disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.

    Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).

    Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.

    independent-referee-for-metric-disputesbody-frame-velocity-api-auditsame-sign-response-is-yaw-biasmedian-hides-bimodal-distributionratio-metrics-need-absolute-checkheading-integral-not-body-ratesame-distribution-reward-comparisonmultiseed-sign-test-for-driftsingle-impulse-recovery-is-chaotic

  12. doctrine-12The deployment pipeline is plant

    Irreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.

    Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).

    Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.

    pipeline-latency-is-plant-not-drpower-scale-hurts-nonforward-axesdeploy-scaling-not-training-equivalentteleop-command-band-per-axislatency-dr-covers-measured-pipelinedeploy-rate-limiter-windupslew-anchor-is-an-integratorbeta-anchored-action-targetpower-derating-cuts-full-range-contract

  13. doctrine-13DR budget is finite; its distribution is the measured support

    Robustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.

    Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).

    Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.

    push-dr-conditional-budget-conservationdr-tail-plant-continuationconstant-value-dr-overfits-margintask-shaping-before-plant-hardeningcom-randomization-forces-leg-spreadcom-dr-rollback-on-symptomthin-dr-judged-by-channel-coveragefriction-priority-re-measured-after-plant-change

  14. doctrine-14Gates measure what hardware feels: posture, margins, stripped assists

    Acceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.

    Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).

    Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.

    task-metrics-vs-posture-metricsstand-gate-posture-not-survivalconstant-value-dr-overfits-marginsuspension-probe-removes-free-stabilizerhip-roll-sum-predicts-lateral-driftchirality-scored-separatelylow-speed-commands-reward-draggingend-state-confusion-matrixfrozen-acceptance-distribution-and-pinned-seedvideo-as-acceptance-recordepisode-length-bounds-what-a-gate-seesremoved-wall-returns-on-hardware

  15. doctrine-15Fork and root selection: recoverability, maturity, frozen rewards

    Choose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.

    Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).

    Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.

    fork-root-recoverable-shortfallroot-maturity-vs-product-qualityfine-tune-reward-change-falsified

  16. doctrine-16Curricula: verified engagement, lineage counters, disease-phase gating

    Automatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.

    Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).

    Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.

    auto-curriculum-engagement-checkcurriculum-counter-lineage-stepsgate-penalties-to-the-disease-phaseaggregate-metrics-mask-subgroup-failurebucket-share-is-not-a-gradient-levercurriculum-criterion-conditioned-on-lagging-categoryper-step-income-drives-speed-time-gatetime-gate-vs-wide-stance-retire-the-fix

  17. doctrine-17Probe before training: feasibility first, hypotheses in tables

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.

    Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).

    Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).

    open-loop-probe-before-reward-tuninghypothesis-table-code-auditperiod-doubling-evidence-racesuspended-test-isolates-actuator-authorityconfiguration-probe-wall-not-slopeprone-dead-end-is-foot-placementdof-vel-penalty-is-not-a-pacing-knobamplitude-cut-falsified-yaw-fix

  18. doctrine-18External advice is recomputed locally; values transfer as ratios

    Every external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.

    Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).

    Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.

    external-advice-audit-against-own-arithmetictransfer-ratios-not-absoluteslatency-dr-covers-measured-pipelinereference-structure-fk-amplitude-divisioncycle-average-tracking-for-gait-quantitiesadvisor-paraphrase-vs-paper

  19. doctrine-19Hardware sessions are scripted experiments, not tuning sessions

    Real-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.

    Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).

    Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.

    risk-ordered-real-deploymentbattery-bracketed-real-abknow-zero-command-behaviorpush-test-chirality-protocolno-field-tuning-protocolreversible-single-variable-field-experimentssim-veto-needs-real-confirmationfirst-real-get-up-violent-stage-one-policystaged-hang-mat-floor-for-get-uppower-cycle-preflighttwo-machine-config-disciplinefall-guard-becomes-a-statehardware-log-is-the-attribution-input

  20. doctrine-20Close questions in writing; restart when the debt is structural

    Audited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.

    Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).

    Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.

    frozen-verdicts-semantic-boundariesstale-verdicts-under-old-stackzero-offset-calibration-shifts-envelopeplant-swap-invariants-vs-shiftsfreeze-lineage-fix-structure-restartminimal-reward-table-with-provenancewrite-hardware-verdicts-back

  21. doctrine-21Name the quantity in the space it lives in

    A goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.

    Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).

    Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.

    joint-space-proxy-for-task-space-quantitycontact-detector-single-signal-lieszero-partial-credit-is-not-an-iteration-problemheading-integral-not-body-rate

  22. doctrine-22Continuation needs a live gradient; a release is chosen by a scan

    Continue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.

    Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).

    Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.

    converged-continuation-is-poisoncheckpoint-choice-is-a-full-gate-scanstance-decided-by-get-up-pathstop-stacking-roll-back-and-auditcurriculum-history-is-part-of-the-productcontinuation-budget-not-from-zeroroot-maturity-vs-product-quality

Experience cards

171 cards matching “penalty-gate-is-an-escape-hatch”.

  • Mirror augmentation over an asymmetric default injects systematic error - symmetrize the default first and verify the transform bit-exactmirror-augmentation-needs-symmetric-default
    Mechanism understoodwalkobservation-designobservation-honestycurriculumprocess

    Before enabling any symmetry augmentation, make every constant inside the observation encoding exactly symmetric, and validate the mirror transform against forward kinematics to machine precision - an unverified augmentation is a new error source, not a regularizer.

    Symptom

    Mirror data augmentation was about to be added while both default poses (standing_pose, walk nominal_pose) were asymmetric - stale hand-tuned compensations from before a ground re-calibration, with hip_yaw differing 2.40 deg between sides and the foot soles actually tilted (pitch 2.88/1.35 deg, roll -2.47/+0.25 deg).

    Context

    The observation encodes joint_pos_rel = q - default. Under mirroring q_l -> -q_r, the relation (q-default)_l -> -(q-default)_r holds only if default_l = -default_r; with an asymmetric default, augmentation produces observation pairs that are NOT mirror images, i.e. "default 不对称时做镜像增强会引入系统性错误,比不做还糟" (worse than not doing it). The fix: adopt model geometric zero as standing default (MuJoCo FK verified: sole pitch/roll exactly 0, asymmetry 0.00 deg) and a symmetric crouch for walk (hip -0.25/knee -0.5/ankle -0.25 satisfying hip - knee + ankle = 0 to keep soles flat). The mirror transform itself was verified bit-exact before use: pseudovector vs polar-vector sign patterns (ang vel [-1,1,-1], gravity [1,-1,1], cmd [1,-1,-1]), joint swap-and-negate; FK check that left-foot pose under q equals the mirror of right-foot pose under mirror(q), measured error 0.00e+00.

    Change

    Defaults symmetrized first (with init heights recomputed by FK), stand policy retrained on the new default so both policies share one default; augmentation enabled only after the FK mirror test passed.

    Outcome

    stand_v1 achieved exact left/right pairing (l_knee -0.1013 / r_knee +0.1013), six-pair asymmetry 0.0 deg, height fluctuation 7 -> 1 mm, 33% less mean |action|.

    Mechanism

    Augmentation asserts an equivariance of the observation encoding; any asymmetric constant inside the encoding (the default) breaks the asserted symmetry, so the augmented data teaches a false invariance. Verifying the transform against FK geometry tests the assertion end to end, independent of the training stack.

    Applies when

    • adding mirror/symmetry augmentation to locomotion training
    • defaults or trims were hand-tuned per side at any point
    • observations are expressed relative to a default pose
    “观测里 joint_pos_rel = q − default。镜像下 q_l → −q_r,要让 (q−default)_l → −(q−default)_r 成立,必须 default_l = −default_r。default 不对称时做镜像增强会引入系统性错误,比不做还糟。… 位置误差与姿态矩阵误差实测均为 0.00e+00。”
    train/RETRAIN_v2.md § 2. 前提:default 姿态必须先对称化(不是可选项) / 3. 镜像变换
  • An auto-curriculum ratchet capped out at iter 248 and never engaged - stage difficulty manually or verify engagementauto-curriculum-engagement-check
    Observed oncewalkcurriculumcurriculumdomain-randomizationprocess

    Prefer manually staged difficulty with gated transitions; if you use an automatic curriculum, instrument its internal state and alarm when it stops engaging - a saturated curriculum is constant DR wearing a curriculum's name.

    Symptom

    A curriculum mechanism intended to grow difficulty adaptively (s1f's ratchet) hit its cap at iteration 248 and never bit again - for 96% of the run its effect was equivalent to constant DR, i.e. the curriculum existed in name only.

    Context

    When external advice suggested graded wz bands (start ±0.15, then ±0.30), the team agreed with grading but explicitly rejected automatic curriculum, citing the s1f episode. The same logic had already been paid for with push grading: ±0.6 failed twice, ±0.3 was feasible - grading matters, but the grade transitions were made by hand at verified checkpoints.

    Change

    Ladder policy: difficulty staged manually, one band per rung, each transition gated by the acceptance battery; automatic ratchets not used unless their engagement is monitored and demonstrated.

    Outcome

    Every C-ladder band change (wz ±0.12-0.25 first, wider later) was an explicit, attributable rung; no silent constant-DR-in-disguise runs recurred.

    Mechanism

    Adaptive curricula couple their own state machine to noisy training metrics; a ratchet that saturates early stops adapting but keeps its name, so the operator believes difficulty is progressing when it is frozen. Manual staging costs more decisions but each decision is observable and reversible.

    Applies when

    • choosing between auto-curriculum and staged bands for a new skill
    • a curriculum's difficulty parameter plateaus early in training
    • post-hoc attribution of what difficulty a lineage actually saw
    “C2 的 wz 分级(±0.15 → ±0.30,不要一上来 ±0.6)—— 与我们付过学费的 push 分级同形(±0.6 两轮 FAIL,±0.3 才可行)。但必须手动分级,不做自动课程 —— s1f 的课程化棘轮 iter 248 封顶未咬合,96% 时长等价常量 DR”
    train/C_LADDER_RUN.md § 1. 采纳 3 条 (C2 的 wz 分级)
  • With no term on foot attitude the robot stood on the edges of its feet (ankle roll -26 deg) from R0.5 to V2.5b; pricing it fixed that and turned stance width into the next free variable - the user saw both on video before any metric flagged themunpriced-foot-attitude-is-a-free-variable
    Replicatedrecoveryreward-shapingreward-shapingreal-acceptance

    List every posture quantity the hardware cares about (foot attitude, stance width, yaw, knee flexion) and make sure each is priced by a term with real gradient at the observed error; pricing one exposes the next, so re-inspect the feet-level video after every change.

    Symptom

    Watching the v2_5 video the user said the ankle roll after standing looked very strange - the feet were not flat. The feet view showed the right foot standing on its outer edge; median ankle roll at t = 8 s was -/+26 deg (limit +/-35), mirrored. This was the physical form of the ankle_roll saturation criterion that had failed since R0.5.

    Context

    No term priced foot attitude: feet_contact_upright counts an edge contact as contact, and stand_pose's wide exp kernel (sigma^2 = 9) gives almost no gradient at 0.45 rad. HoST carries an "ankle parallel" term (+20); this table never had one. Edge standing was made a blocking precondition for hardware (continuous ankle load, unstable contact, wear).

    Change

    V2.6 (single variable): flat_feet = (|q_l_ankle_roll| + |q_r_ankle_roll|) x upright gate x height gate, a linear hinge with a 5 deg margin, weight -2, continued from V2.5b.

    Outcome

    Ankle roll -/+26 -> 5.3/5.2 deg, ankle_roll saturation 6.6% (< 10%), feet flat on the feet-view video; success 99.6%, re-falls 0-1%. Then the user watched v2_6: hip yaw constantly tense and the legs very close together. The numbers: hip roll -/+4.9/4.8 deg against a nominal 25 - feet flat and hips open 25 deg cannot coexist without ankle compensation, the new term taxed that compensation, and nothing priced stance width. "Foot attitude as a free variable" was fixed and "stance width became the new free variable" - which the real robot then exposed as splits.

    Mechanism

    An optimizer spends every posture degree of freedom no term prices; closing one reallocates the slack to the next unpriced one.

    Applies when

    • a standing or landing posture looks wrong on video while gates pass
    • a saturation criterion keeps failing on one joint
    • a new posture term was just added
    “用户看 v2_5 视频:"起身之后 ankle_roll 非常奇怪,脚根本不是平着站立"。 … 机理:奖励表**无任何脚掌姿态项** —— feet_contact_upright 边缘接触也算触地, stand_pose 的 exp 核(σ²=9)对 26°=0.45 rad 梯度≈0。脚掌姿态是自由变量。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §41 预注册 V2.6:flat_feet —— ⑥ 老账的物理形态被用户目视锁定
  • A plateau in a training curve was a population mix, not a half-learned skill - 60% standing at 0.372 m and 40% sitting at 0.19 m - and a category at hard zero stayed at zero through 3,000 more iterationszero-partial-credit-is-not-an-iteration-problem
    Mechanism understoodrecoveryattributionmeasurementattributioncurriculum

    Before buying iterations for a plateau, split the metric by category and check whether it is bimodal; a category at hard zero with no partial credit is missing a capability or a reachable state, and more iterations under an unchanged config will only polish the categories that already work.

    Symptom

    After R0.1 the training-side base_height sat near 0.29 m and the curve was still climbing at the iteration cap, which read as "train it longer".

    Context

    Candidate A (user decision) was a child-run from R0.1's last checkpoint with zero config change - the logged env.yaml files differ only in log_dir - for 3,000 more iterations (R0.2). The per-category acceptance split was already available: prone had scored 0/156 with no partial credit.

    Change

    Continue training unchanged, then read the result by category rather than by the pooled curve.

    Outcome

    supine 91.5 -> 96.4%, side 82.2 -> 87.9%, get-up 0.96 -> 0.84 s, pose error 0.91 -> 0.54 - all improvements to categories that already stood. Prone stayed 0/153; mid 55.9 -> 41.2% was within noise (n=34). Height by category was binary - standing groups 0.372/0.373 m, seated groups 0.187/0.194 m, nothing between - so the pooled 0.29 was 0.61 x 0.372 + 0.39 x 0.19 = 0.30 (measured 0.307): six in ten standing, four in ten sitting. The rendered prone episode was still kneel-sitting at t = 8 s.

    Mechanism

    The pooled mean of a binary outcome only moves when the mix moves; PPO kept polishing the subpopulation that already succeeded while the failing one produced no advantage signal to follow.

    Conflicts

    R0.2 recorded the missing capability as "prone lacks rolling over"; R0.3's end-state confusion matrix retracted that - prone had righted its torso in 159/159 episodes and was failing to stand from the W-sit. The lesson that iterations could not fix it holds; the named cause was wrong.

    Applies when

    • a training curve plateaus while acceptance shows one category at zero
    • deciding between "train longer" and "change something"
    • pooled training metrics are read as the typical episode
    “**分类别 h 中位把"平台 = 人口混合"钉死了**:数值是**二值**的 —— 站立组 0.372/0.373,坐姿组 0.187/0.194,**中间没有过渡态**。 … **这也是本仓此后读该指标的通用告诫:全体混合的期望会把 双峰分布平均成一个不存在的中间值,必须分类别看。** … **结论:A 不能过门,原因确定为 prone 缺"翻身"这一技能,不是迭代不够。**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §13 R0.2(recovery_r0_2,child-run 续训):A 走完了 —— 推不动 prone
  • mj_objectVelocity returns inertial-principal-axis frame - one API assumption poisoned eval and observations for a whole linebody-frame-velocity-api-audit
    Mechanism understoodomnisim2sim-gatemeasurementsim2simobservation-honestyattribution

    Verify every frame-sensitive API against a hand-computed truth (rotate raw qvel yourself, or command a known world velocity and check where it lands) before trusting any evaluation or observation built on it - especially when a model's inertial frame is rotated from its body frame.

    Symptom

    Sidewalk vy read ~0 under every condition; separately, whole-policy performance was mysteriously mediocre in sim2sim while training-side numbers looked fine. Four training rungs were declared FAIL partly on these readings.

    Context

    base_link's URDF inertial frame is rotated 90 deg about x relative to the body frame (iquat = [0.7071, 0.7071, 0, 0]). mj_objectVelocity(flg_local=1) rotates into ximat - the inertial principal-axis frame - not the body frame, and returns center-of-mass point velocity, not body-origin velocity. Consequences measured: the "vy" column was actually vertical velocity vz (walking at cmd 0.25: old reading +0.0093 vs true -0.0424); the angular velocity fed to the policy in sim2sim was [wx, wz, -wy] - a different quantity than Isaac and the real IMU provide. RMS check over 8 s of walking: y/z axes swapped between v6[:3] and the qvel truth.

    Change

    Fixed sim2sim and both probes to compute ang_b = qvel[3:6] and lin_b = xmat.T @ qvel[0:3] (identical quantity to Isaac's root_ang_vel_b / root_lin_vel_b), with a standalone reproduction script (frame_bug_repro_0809.py).

    Outcome

    Re-scoring the "failed" C4 lineage under correct coordinates reversed the verdicts: c4r4 checkpoints showed vy 80-126% tracking (old reading: +/-2%) and vx+0.30 at 91-95% where the old metric said 0/5 - the bad frame both mis-measured vy and, via corrupted policy observations, systematically depressed all measured performance. Final product passed 260/260 cells.

    Mechanism

    A simulator API's frame convention is part of the observation contract; when the model's inertial frame is rotated relative to the body frame, frame-agnostic use of a "local" velocity silently permutes axes. Feeding a policy an axis-permuted angular velocity is an observation corruption that degrades behavior everywhere, not just on the axis being studied.

    Applies when

    • building or auditing a cross-simulator evaluation harness
    • one measured axis reads near-zero under all conditions
    • sim2sim scores are inexplicably worse than training-side metrics
    • URDF/MJCF inertial frames are rotated relative to body frames
    “base_link 的 iquat = [0.7071, 0.7071, 0, 0] … mj_objectVelocity 用的是这个 … 喂给策略的 base_ang_vel 是 [wx, wz, −wy] —— MuJoCo 侧观测与 Isaac / 真机 IMU 不是同一个量;vy_mean 报的是竖直速度 vz —— 前进 cmd 0.25 时旧读数 +0.0093,真值 −0.0424。”
    train/C_LADDER_RUN.md § 3l. ⚠️ mj_objectVelocity 读的是惯性主轴系 / 3m. 一 bug 坐实
  • An outer heading P-loop at deploy cut drift 10x because its output stays inside the trained command band - then training was aligned to itdeploy-heading-loop-and-align-training
    Mechanism understoodwalkreal-deployreal-acceptancecurriculumattribution

    Fix drift-class problems first with an outer loop whose output provably stays inside the trained command band; when adopting it permanently, align the training command generator to the deployment's actual command mixture (feedback-driven AND constant), matching law, gain, and clip exactly.

    Symptom

    Persistent heading drift on straight-line walking (v6 net yaw 60.3 deg over 15 s) that reward-side fixes had only partially tamed.

    Context

    The deploy stack added --heading: an external P loop wz = clip(0.5 * wrap_to_pi(theta0 - theta), +/-0.6), recomputed each frame and fed into the policy's ordinary wz command slot. Measured: net yaw walk_v6 60.3 -> 5.8 deg, walk_v5 17.4 -> 4.2 deg. A run-level audit later corrected the mechanism story: training had heading_command=False since v1 - the policy had NEVER seen heading-error feedback, so the loop works purely because its output lands inside the trained command distribution wz ~ U(+/-0.6): "收益真实,当时的机理解释写错了" (the benefit is real; the mechanism explanation had been wrong). v8 then closed the loop properly: training-side heading command enabled with rel_heading_envs=0.5 - half the envs get heading-error-driven wz, half get explicit constant wz, because deployment feeds wz BOTH ways (straight-line = heading feedback, turning = constant command) and rel=1.0 would have made constant-wz turning out-of-distribution. The law, gain, and clip were aligned item-by-item between trainer and deploy tool.

    Change

    Deploy-side outer loop first (no retrain needed); then v8-D enabled the matching training-side heading command at rel=0.5 with identical gain (0.5) and clip (+/-0.6), contract unchanged (wz slot carries the computed value).

    Outcome

    Drift handled at deploy (5.8 deg) generations before training caught up; the alignment removed the residual train/deploy distribution mismatch, with the accepted cost booked (open-loop straight walking becomes more OOD for heading-envs - irrelevant since acceptance and deployment always run the loop).

    Mechanism

    A learned velocity-tracking policy is a valid inner loop for any outer controller whose commands stay within the trained command distribution - the policy needs no knowledge of the outer objective. Full alignment then requires training on the same mixture of command sources the deployment actually uses, in the observed proportions.

    Applies when

    • heading/position drift on a velocity-tracking policy
    • designing outer loops over learned locomotion controllers
    • training command distribution differs from how deployment feeds commands
    “审计更正(2026-08-02,run 级 env.yaml):训练侧自 v1 复盘起就是 heading_command=False … 策略从未见过航向误差反馈。--heading 是评估/部署侧外加的航向 P 环(wz=clip(0.5·err,±0.6), 落在训练分布 wz~U(±0.6) 内)。实测净偏航 walk_v6 60.3° → 5.8° … 收益真实,当时的机理解释写错了”
    train/WALK_V7_SPEC.md § 0. 本轮之前已经改掉 (航向闭环, 含审计更正)
  • Diagnose a behavior failure by enumerating hypotheses and auditing each against the actual config, cheapest firsthypothesis-table-code-audit
    Replicatedwalkattributionattributionprocessreward-shaping

    Before changing anything, write the full hypothesis list for the symptom and audit each against the resolved config and measured magnitudes, cheapest check first; train only on the survivors.

    Symptom

    Real robot leaned forward "wanting to walk" but dragged its feet instead of lifting them - a symptom with many plausible causes and no obvious single fix.

    Context

    Seven hypotheses were listed and each checked against the actual training config files (velocity_env_cfg.py, isaac_values.py), ordered by check cost: missing foot clearance term (CONFIRMED, primary - feet_air_time existed but no swing-height term at all); energy penalties dominating (REJECTED - energy terms total -0.19 vs tracking +1.2, 16%); command range too narrow (CONFIRMED - (0.15,0.35)); nominal pose too crouched / action scale too small (HALF - knee 0.5 rad = 28.6 deg deep, scale fine); mixed PD across motor types (REJECTED - already grouped); missing base-height reward (REJECTED - present at -5.0); height-drop termination (REJECTED - none exists, which itself became finding #4 of the fix list).

    Change

    The audit produced a ranked fix list (add clearance penalty; widen speed range; reduce nominal crouch) with each rejected hypothesis documented so it would not be re-litigated.

    Outcome

    Three confirmed causes fixed over v5/v6: swing height went 22-23 mm -> 34 mm, tracking 81% -> 87%; the rejected hypotheses stayed rejected (no wasted rungs on energy weights or PD grouping).

    Mechanism

    Multi-cause symptoms invite guess-and-train loops; a written hypothesis table forces each candidate to be confirmed or rejected against actual values (not impressions), and cost-ordering the checks means most hypotheses die for the price of reading a config.

    Applies when

    • a real or sim behavior failure has multiple plausible causes
    • the team is about to "try a fix" without an audit
    • post-mortems keep re-proposing already-rejected causes
    “真机现象:躯干前倾像要走,脚抬不起来(拖着蹭)。按成本从低到高逐条核查 … | 1 | 缺 foot clearance | ✅ 成立,首要 | 有 feet_air_time,无任何摆动足高度项 | | 2 | 能量惩罚压过跟踪 | ❌ 不成立 | 能量类合计 −0.19,跟踪 +1.2,只占 16% |”
    train/WALK_DIAGNOSIS.md § walk 拖地问题 — 七条假设的代码核查结果
  • Under continuous 3-axis uniform sampling, pure straight-line walking is a zero-measure event the policy never trainedzero-measure-commands-need-mode-sampling
    Mechanism understoodomnicurriculumcurriculumobservation-honestygate-battery

    Enumerate the exact command points users will actually issue (straight, stop, in-place turn) and give each explicit probability mass via mode sampling with off-axes pinned to zero - never assume a continuous sampler covers its measure-zero subsets.

    Symptom

    "The robot drifts even in sim when told to walk straight" persisted across reward tunings - because with commands drawn as vx in [0.15,0.5] x vy ~ U(+/-0.2) x wz ~ U(+/-0.6), the event vy=0 AND wz=0 has probability zero: pure straight-line walking was never sampled even once.

    Context

    Restart evidence item #3: "纯直行是零测度点 … 'sim 里直行就漂'是分布的 必然,不是 reward 没调好" - the drift metric was legitimately drowned by commanded turning (v11's own comment self-documented this). The structural fix is discrete mode sampling: a custom ModeVelocityCommand that first draws a mode by share (stand/forward/back/turn/side/mixed), then draws values only on that mode's axes with all others pinned to exact zero - which is also what preserves single-variable discipline in the C ladder (native 3-axis uniform "采不出'离散模式桶' … 把 C1~C4 的单变量纪律直接毁掉"). The mixed mode later got an ellipsoid constraint rather than a cube for the same reason in reverse - corner combinations of a cube are unrepresentative extremes.

    Change

    Command generation moved from independent per-axis uniforms to mode-bucket sampling with pinned-zero off-axes (plus 20% rel_standing); acceptance likewise evaluates per mode.

    Outcome

    Straight-line behavior became a trained, testable mode instead of a measure-zero hope; the C ladder could add one mode per rung with provable isolation.

    Mechanism

    A policy optimizes expected reward under the command distribution; events of probability zero contribute nothing to the objective, so exact-zero-command behaviors (straight walk, stand, in-place turn) are only learned if the sampler gives them mass. Product-of-uniforms distributions concentrate mass on mixtures and give none to the pure behaviors users actually command.

    Applies when

    • a "simple" command (straight, stop) underperforms mixtures in sim
    • designing command distributions for velocity-tracking tasks
    • a ladder needs per-mode isolation for attribution
    “纯直行是零测度点:最终 command 为 vx∈[0.15,0.5] × vy∈U(±0.2) × wz∈U(±0.6) 连续均匀,vy=0∧wz=0 从未被专门采样 —— "sim 里直行就漂"是分布的必然,不是 reward 没调好 … Isaac 原生 UniformVelocityCommand 是三轴各自 uniform,采不出"离散模式桶"”
    train/OMNI_V0_SPEC.md § 0. 为什么从零 (3) / 三件前置 (1)
  • Prone get-up sat at 0/159 until two gated hinge terms moved the seated feet - first sideways (561 -> 360 mm), then fore-aft (-168 -> +56 mm) - and success went to 158/159 with nothing else changedprone-dead-end-is-foot-placement
    Mechanism understoodrecoveryreward-shapingreward-shapingattributioncurriculum

    When a stuck state and a successful state differ geometrically, penalize the discriminating quantity with a gated hinge that is exactly zero in the state the policy actually reaches (measure it - not the nominal), then re-probe: flattening one axis can move the discriminant to another.

    Symptom

    Prone falls always righted and then sat with the feet splayed wide or tucked behind the hips, from where the policy never stood (0% for four generations).

    Context

    Four lines of evidence pointed at foot position: the configuration probe (ankles 215 mm apart stood 52.3%, 561 mm apart 0.0%); FK showing the action contract's nominal (a = 0) is itself a 465 mm straddle, so the action_rate and still terms were pulling toward the splits; biomechanics (feet tucked under the body cut peak hip-extension torque 148.8 -> 32.7 N*m, -78%); and HoST's foot-displacement term, which this reward table lacked. The earlier "not a reward hole" reading was corrected to "a gradient hole, not a level hole": at the dead point the heaviest term (upright) was saturated with zero gradient, still paid for not moving, and the one live gradient (base_height) pointed at the thigh-horizontal torque barrier. A prone ROM scan had already ruled out pushing up from prone.

    Change

    R0.4: feet_spread_excess = clamp(ankle distance - 0.215, 0, inf) x upright gate, weight -2.0, plus a height-decay factor added after measuring that the policy's real standing stance was 406 mm, not the 215 mm nominal (the plain version would have taxed every successful stand 0.38/s). R0.5: the same shape on the fore-aft axis, feet_fore_seated = |fore-aft offset - 0.05| x upright gate x height decay, target +50 mm (the measured natural offset of standing postures). One variable per rung.

    Outcome

    R0.4: seated ankle distance 561 -> 360 mm, supine/side exactly unchanged, prone 0 -> 1.9%, mid 45.9 -> 62.2%; a probe then showed the discriminant had moved to the fore-aft axis (standing starts +42 to +51 mm, the prone seat -168 mm). R0.5: supine 99.4, prone 99.4, side 100, mid 100%, re-falls 0%; both geometry terms collapsed to ~0 near iteration 13,100 as base_height rose, and the prone fore-aft offset went -168 -> +56 mm - the term's own target, closing the causal chain. The cost, unmeasured at the time: action jitter rose 33% (sum |da|^2 6.82 -> 9.06).

    Mechanism

    An upright-gated hinge is inert while the robot rolls and exactly zero in the achieved stance, so it adds gradient only inside the stuck basin; a seated robot with its feet behind or outside its COM must make a kinematically unfavourable transition to stand, and moving the feet under the body removes it.

    Applies when

    • a get-up or transition skill fails from one start category only
    • successful and failed episodes differ in a measurable geometric quantity
    • a shaping term might tax the posture successful episodes already use
    “`recovery_r0_5`,唯一变量 = 追加 `feet_fore_seated`(与 R0.4 同形状,只换测量轴)。 … 对照 R0.4 的 prone(3.1%,360 mm,**−168 mm**):前后偏移从 −168 走到 +56, 正是这一项的目标量,**判别量被消掉后成功率随之到顶** —— 因果链完整。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §20 R0.5(前后向脚位置项):成功率门全过 —— prone 0/159 → 158/159
  • Narrowing the speed range to stop high-speed falls entrenched crouch-shuffling - judge gait quality at the speed that demands a gaitlow-speed-commands-reward-dragging
    Mechanism understoodwalkcurriculumcurriculumreward-shapinggate-battery

    Set command ranges to include speeds that physically demand the target behavior, and evaluate behavior-quality gates at those speeds; when a restriction is added to suppress a failure, check what new optimum it creates at the remaining commands.

    Symptom

    After the command range was narrowed to (0.15, 0.35) m/s (to treat walk_v1's 134% overspeed and 8.3 s fall at 0.5), the policy settled into crouched foot-dragging; tracking rose monotonically with speed (63% at cmd 0.2, 76% at 0.3, 87% at 0.45), showing low speeds were where the degenerate gait was optimal.

    Context

    The narrowing advice was the author's own and is retracted in the file: it treated the symptom (falls at speed) while reinforcing the root cause (at 0.15-0.35 m/s, shuffling in a crouch is globally optimal - the Froude number is so low that even humans would not lift their feet). A zero-cost experiment confirmed the flip side: at cmd 0.5 the same policy met BOTH tracking (81%) and clearance (23.0/23.2 mm) standards.

    Change

    Speed range widened back toward (0.15, 0.5) - upper bound deliberately slightly above the mechanically feasible ~0.44 m/s so the policy finds the boundary itself; acceptance re-pointed: gait-quality criteria (tracking, clearance) judged at 0.45-0.5 m/s, low speed kept only as a survival check.

    Outcome

    v4 -> v6 progression under the widened range delivered 87% tracking with 34 mm clearance; the "low command = drag" account was confirmed by the monotone tracking-vs-speed curve.

    Mechanism

    Command distribution is part of the reward: physics prices gaits per speed, and at very low speed the energetic optimum is no swing phase at all. Restricting training to that regime makes the degenerate gait the correct answer to the posed problem - and grading a gait at a speed that does not require stepping measures nothing.

    Applies when

    • a gait degenerates after a command-range restriction
    • quality metrics improve monotonically toward the range boundary
    • writing acceptance criteria for gait quality vs survival
    “现在看那个建议可能起了反作用:0.15~0.35 m/s 下蹲着蹭就是全局最优,抬腿反而亏。收窄治的是"高速摔倒"的症状,却强化了拖地的病根。… 验收标准里的 cmd 0.2 本身就是拖地速度(Froude 数极低,人在那个速度下也不抬脚)。accept_v2 应把速度跟踪与 clearance 的判定点改到 0.45~0.5 m/s”
    train/WALK_DIAGNOSIS.md § ② 放宽速度区间 / ① 零成本实验
  • Push-test protocol - positive side first, fragile side spotted, axes aligned in the log, and cross-machine push counts stay qualitativepush-test-chirality-protocol
    Mechanism understoodomnireal-acceptancereal-acceptancegate-batteryprocess

    Order disturbance tests from the robust side to the fragile side with protection scaled to sim-measured asymmetry, align and log frame conventions before testing, and treat cross-domain disturbance counts as qualitative evidence only.

    Symptom

    Hand-push testing on hardware risked falls on a side sim had already flagged as fragile, and push counts invited apples-to-oranges comparison with sim numbers.

    Context

    Sim chirality was explicit: descendants were far more fragile in -y (fric-3000@kd1.2: +6 N*s survived 15/20 vs -6 N*s only 3-9/20) while the s1e control was perfectly symmetric (40/40). The protocol therefore: push the positive direction first, keep a spotter for the negative side; before any push, record which real-robot side corresponds to sim's +y in the log ("上机前对一次坐标"); and - citing the chaos lesson ("混沌课文") - real push results are used only as qualitative corroboration, never compared numerically with sim survival counts across machines.

    Change

    Push testing became a scripted, chirality-aware protocol with frame alignment as a logged precondition and an explicit epistemic limit on cross-domain count comparison.

    Outcome

    The fragile side was tested with protection informed by sim's quantified asymmetry; logs stayed interpretable because the frame correspondence was recorded before the first push.

    Mechanism

    Disturbance-response chirality is a real, quantifiable lineage property, so test order should follow measured fragility; and perturbation outcomes are chaotic in the details (divergent trajectories from tiny differences), so counts do not transfer across domains even when qualitative rankings do.

    Applies when

    • planning push/disturbance tests on hardware
    • sim shows directional asymmetry in disturbance survival
    • someone proposes comparing real push counts to sim counts
    “先正向后负向, 负向留人扶 —— sim 手性明确: 后代在负 y 向显著更脆 (fric-3000 @kd1.2: +6 N·s 15/20 vs −6 N·s 3~9/20), 而 s1e@0.8 两向 40/40 完全对称。上机前对一次坐标 … 跨机不做二值结论 (混沌课文): 真机推力只作定性对照, 不与 sim 计数对比。”
    train/REAL_RUN_S2.md § 3. 抗推 (可选, 人手推; 做则按此协议)
  • Prove an armless get-up exists before training it - a connected static domain, 25% torque on the cheapest path, an 8 mm hand-over gap, a static roll-over - and write down what each scan cannot representget-up-feasibility-accounts-before-training
    Mechanism understoodrecoveryplant-calibrationplant-calibrationhardwareprocess

    Before training a get-up or any multi-contact skill, compute the quasi-static accounts - connectivity of the static domain, torque along the cheapest path, hand-over gaps, COM shift available for rolling - and state which configurations each scan cannot represent; when a policy gets stuck in one of those, extend the scan before blaming the reward.

    Symptom

    A torso-and-legs robot has no arms to push off the ground; whether it can get up from the floor at all was unknown when the line opened.

    Context

    recovery_feasibility.py ran three accounts before any training (the run line's "hard accounts first" discipline): a sagittal quasi-static scan (0.05 rad grid, 44,520 configurations, MuJoCo FK, flat-foot assumption). (1) The static standing domain (COM over the feet, torques in limit) has 29,586 cells, flood-fill connected with no islands, from a 0.097 m deepest squat to the 0.384 m stand. (2) The minimum-torque path peaks at 25% of the limits (knee 2.9/12, ankle 3.7/17 N*m) - a 4x margin. (3) All 508 ground-contact configurations have the contact behind the COM; the smallest gap to pure foot support is 8 mm. A roll-over account: swinging both straight legs to one side shifts the COM 96 mm against a 62 mm torso half-width - 1.6x, so rolling needs no momentum. Three conclusions were written down for later attribution: the legs are 80% of the mass (swinging them moves the COM), prone has no flat-foot hand-over face (merge into a supine/side sit first), and supine needs no sit-up (hip flexion is limited to 75 deg).

    Change

    The accounts gated opening the line and were cited in every later argument about what the robot can physically do.

    Outcome

    They held where they applied: in V1.0 every fall category was righted under a hard rate limit, which the spec records as the quasi-static roll-over account verified by training, and the 25% torque path was the basis for pursuing a slow get-up. They also misled once: account (3) is sagittal, and on 08-09 the spec corrected its scope - it cannot represent the splayed W-sit where the policy actually stalled. A follow-up prone hip-ROM scan (471,625 cells) found 3,912 two-foot-contact cells and none with both soles within 25 deg of level (best 40.2 deg): a flat-footed push-up from prone is infeasible on this robot, so the fix became where the feet go after sitting up.

    Mechanism

    A get-up needs a connected path through statically feasible configurations and enough torque along it; quasi-static accounts bound both cheaply, and momentum can only make the real problem easier. A reduced-dimensional scan, though, only speaks for the configurations it can express.

    Conflicts

    In R0.1-R0.2 the spec read account (3)'s "prone has no front hand-over" as "prone lacks the roll-over skill"; R0.3's confusion matrix showed prone had righted its torso 159/159, and the spec then restricted account (3) to the sagittal configurations it models.

    Applies when

    • opening a get-up, recovery or climbing skill on a new robot
    • a robot lacks arms or other obvious contact options
    • a policy stalls in a configuration a feasibility scan never modelled
    “本机 **torso + legs、无手臂可撑地**,开训前先证明存在不依赖手臂的物理解 … 双腿同侧直腿摆最大横移 **96 mm = 1.6×** … **静态摆腿即可翻身,无需动量**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §1 机械可行性判决(2026-08-09,recovery_feasibility.py 三笔账)
  • After seven patch-generations, freeze the lineage as a regression baseline, fix the structural debts, and retrain from zerofreeze-lineage-fix-structure-restart
    Observed oncewalkprocessprocesscontract-freezecurriculum

    When successive rungs keep trading one symptom for another, ask whether the remaining problems are structural (contracts, latency, sampling, reward-table architecture); if so, freeze the lineage as regression baselines, pay the structural debts, and restart minimal - carrying forward laws and instruments, not weights and weights' patches.

    Symptom

    The v5-v11 walk lineage had accumulated interacting patches (reward terms, gates, clamps, per-joint scales) faster than it converged on the user's goal; v12's spec itself was superseded before training by an external review's verdict that the remaining problems were structural, not parametric.

    Context

    The 2026-08-05 status banner records the pivot: the walk profile was rolled back wholesale to v10b parameters, the v5-v11 lineage frozen "只作回归对照" (kept only as regression baselines), and four structural debts were named as prerequisites for a from-zero straight-walk baseline: the action-latency FIFO (fixed with its own test), the ONNX manifest contract, discrete command sampling, and a minimal reward table. The v12 spec - fully designed, partially implemented - was suspended: "本规格挂起,不再按此开训".

    Change

    Strategy switched from "one more patch generation" to freeze-fix-restart: lineage checkpoints retained as comparison anchors, infrastructure hardened first, then a clean retrain with a minimal reward table (this restart produced the s* generation that later became the real-robot SOTA line).

    Outcome

    A designed-and-ready training generation was deliberately not run - the review's structural findings outranked sunk design cost; the restart line inherited seven generations of laws (calibrations, gate batteries, falsified fixes) without inheriting their entangled reward table.

    Mechanism

    Patch lineages accumulate coupled terms whose interactions eventually cost more to reason about than a restart costs to train; the knowledge worth keeping is the laws and instruments (measured plant values, calibrated gates, falsified directions), not the entangled weights. A restart on hardened structure converts the lineage's lessons into a clean initial design instead of another delta.

    Applies when

    • repeated rungs shuffle symptoms without net progress
    • an external review flags infrastructure/contract debts
    • deciding between another patch generation and a clean retrain
    “同日外部评审定调换路线:冻结 v5~v11 血统(只作回归对照),修结构性问题(latency FIFO 已修 tests/test_action_latency.py、ONNX manifest 契约、离散命令采样、最小奖励表)后从零训直行基线。本规格挂起,不再按此开训。”
    train/WALK_V12_SPEC.md § ⚠️ 状态 (2026-08-05)
  • Raising a command bucket's share does not strengthen its per-state gradient - it only starves the other modesbucket-share-is-not-a-gradient-lever
    Mechanism understoodomnicurriculumcurriculumreward-shapingdomain-randomization

    When a skill is not learning, first prove its per-state signal is nonzero (ignore-floor and probe checks); only rebalance sampling shares to fix genuine sample starvation, and account the regression risk to the diluted modes before doing it.

    Symptom

    Sidewalk was not learning, and the reflex proposal was to give the side bucket a larger share of sampled commands.

    Context

    The C4-redo3 rung explicitly kept the 20/40/20/20 bucket (stand/forward/turn/side) with the reasoning written out: PPO computes advantages per state, so bucket proportion does not change the per-state gradient of side states; at 4096 envs x 20% x 24 steps the rollout already contained ~19.7k sidewalk states - sample count was not the bottleneck. And the cost side was already measured: cutting forward from 60% to 40% had made vx+0.30 die at +400 in an earlier run - more cuts would only collapse it sooner.

    Change

    Bucket proportions held constant across the entire C4 redo series; the actual bottlenecks (metric frame bug, reward variance penalty, exploration form) were pursued instead.

    Outcome

    Sidewalk was eventually fixed with zero bucket changes (feed-forward delivery, +100 iters); forward/turn skills never suffered starvation-induced regressions during the redo series.

    Mechanism

    Policy-gradient credit is assigned per visited state; oversampling a mode multiplies its states in the batch but not the informativeness of each, so if the per-state gradient is ~0 (behavior unreachable or reward indifferent), N times zero is still zero - while the displaced modes genuinely lose data and regress.

    Applies when

    • proposing to oversample a failing task/command mode
    • a majority mode regresses after share rebalancing
    • budgeting env count vs mode share for a multi-skill policy
    “比例不动:PPO 逐状态算优势,桶占比不改变单状态梯度;4096 env × 20% × 24 = 每 rollout 已有 1.97 万个侧走状态,样本数不是瓶颈;而 forward 60%→40% 已实测让 f30 在 +400 处死掉,再加码只会更早塌。”
    train/C_LADDER_RUN.md § 3i. 桶 20/40/20/20 不动(比例不动)
  • Fine-tuning through a reward-table change was falsified (scatter, half-recover, collapse) - continuation training is legal only with the reward frozenfine-tune-reward-change-falsified
    Replicatedomnicurriculumcurriculumfork-selectionreward-shapingprocess

    Never fine-tune through a reward-table change - retrain from zero; reserve checkpoint continuation for frozen-reward plant/DR widening, reset noise_std when branching, and watch for the scatter/half-recover/collapse signature as the abort trigger.

    Symptom

    The s1c A/B experiment: arm B fine-tuned from an existing checkpoint under the revised reward (same contract, same network, changed reward + small DR) and failed with a characteristic signature - scatter, half-recover, fall back ("打散→半恢复→摔回"); arm A trained from zero under the same config won decisively (full shaping lifted swing to 21.6 mm within 500 iters; shipped at 5500).

    Context

    Verdict recorded: "从零 + 强塑形是本机唯一验证过的发育路径" (from-zero plus strong shaping is this machine's only validated development path). The signature became a standing stop criterion in every later rung that touched a reward ("s1c B 臂签名,出现即停"). Crucially the boundary of the law was drawn explicitly when S2 continuation training was proposed: "当年证伪的是「奖励表中途改版的 fine-tune」… S2 续训奖励表全程冻结、只逐级加宽 plant/DR,属另一类" - continuing a checkpoint with the reward FROZEN while widening plant/DR one rung at a time is a different class and was allowed (and then worked, powering the whole S2/C lineage) - with the honest fallback that if frozen-reward continuation ever collapses, that rung retrains from zero and the doctrine gets re-examined with data. Fine-tune arms also need mechanical care: reset the checkpoint's collapsed noise_std (terminal 0.033 "会杀死探索") and account for iteration counters re-zeroing (curriculum gates fire immediately).

    Change

    Reward changes and lineage continuation permanently separated: reward revisions -> from-zero retrain; plant/DR widening -> frozen-reward continuation with per-rung gates; the B-arm signature promoted to a universal tripwire.

    Outcome

    No later reward revision was attempted by fine-tune; frozen-reward continuation carried S2 (PD/COM/friction rungs) and the C command ladder successfully from the s1e root.

    Mechanism

    A trained policy sits in an optimum of its reward's geometry; changing the reward moves the optimum but leaves the policy's exploration noise near-zero and its value function calibrated to the old returns - it disassembles the old solution faster than it can assemble the new one. Widening DR under a frozen reward instead keeps the optimum's identity and asks only for local robustification.

    Applies when

    • proposing to fine-tune an existing policy under a revised reward
    • planning a robustification ladder from a validated checkpoint
    • a continued run scatters then partially recovers then collapses
    “B 臂 fine-tune 证伪(打散→半恢复→摔回——从零 + 强塑形是本机唯一验证过的发育路径)。… 当年证伪的是「奖励表中途改版的 fine-tune」(B 臂,塑形突变致终盘摔回);S2 续训奖励表全程冻结、只逐级加宽 plant/DR,属另一类;若 s2_lag1 续训本身塌方,回退方案 = 该级从零重训,续训教义再议(拿数据说话)。”
    train/OMNI_V0_SPEC.md § 3. S1.3 / 4. 与 s1c fine-tune 证伪的关系
  • The restart reward table lists a reason for every term AND a lesson for every exclusion - absent terms are removed, not zero-weightedminimal-reward-table-with-provenance
    Mechanism understoodomnireward-shapingreward-shapingprocesscontract-freeze

    Maintain the reward table as an evidence ledger: every term cites the episode that justifies it, every excluded term cites the episode that convicted it (including the development stage it is valid at), and retired terms are deleted from the config, never left at weight zero.

    Symptom

    Seven walk generations had accumulated an entangled reward table where nobody could say which term earned its place; the restart needed a table that could be audited line by line.

    Context

    The minimal table v2 was built under three written principles: "一项管 一件事、结构性反抬脚的项一个不留、塑形只留一套相位逻辑;不在表里的一律 不加" (one term per job; zero structurally-anti-lift terms; exactly one phase-shaping logic; nothing outside the table gets added). Every row carries its provenance (e.g. base_height target = standing height cites the v4 crouch lesson; split x/y tracking cites the merged-exp gradient hole; world-frame yaw cites the v1/v3 body-frame lesson). Every EXCLUSION carries its same-type precedent: feet_landing_vel out because it is poison while the gait is unformed (drag pays 0, lifting pays - the v4-clearance / v8a-B "reverse threshold" family) though it was fine in v7 when the gait already existed - term validity depends on developmental stage; feet_air_time out with its measured non-lever evidence (v6 had it at 2.0 and still lifted 4 mm); and replaced terms are REMOVED from the config ("置 None,不是权重 0 挂着") so audits see truth, not dormant weight.

    Change

    Reward table rebuilt as ~20 rows each with weight + provenance column; exclusion list maintained alongside with the falsifying episode for each; dormant terms deleted rather than zeroed.

    Outcome

    Later revisions (S1.2/S1.3) modified the table by citing and updating specific rows' evidence rather than re-arguing the whole design; the table doubled as the lineage's reward-lesson index.

    Mechanism

    A reward table is a set of standing hypotheses; attaching each row's evidence makes revisions targeted and reversible, and recording why a term is absent prevents the cycle of re-adding known poisons. Deleting vs zero-weighting matters because config audits and DR interactions see the term either way - a zero-weight term is dormant complexity waiting to be flipped on wrongly.

    Applies when

    • designing a reward table for a restart or new task
    • someone proposes re-adding a previously removed term
    • auditing which reward rows still earn their place
    “原则:一项管一件事、结构性反抬脚的项一个不留、塑形只留一套相位逻辑;不在表里的一律不加 … feet_landing_vel(评审 #4):拖地时代价恒 0、抬脚才收费——与 v4-clearance/v8a-B 同属「反抬脚门槛」家族,步态未成形时是毒;v7④ 加它时步态已存在。… 已从 cfg 移除(置 None),不是权重 0 挂着。”
    train/OMNI_V0_SPEC.md § 3. 最小奖励表 v2 / 明确不带
  • A real-robot verdict is (policy x deployment stack) - when the stack changes materially, old verdicts expirestale-verdicts-under-old-stack
    Mechanism understoodwalkreal-acceptancereal-acceptanceattributionprocess

    Date every hardware verdict with the deployment-stack version it was measured under; after any material stack change, re-test before trusting old condemnations or old praises - with the interpretation of each possible result written down first.

    Symptom

    walk_v5 stood condemned as "kicks wildly" and walk_v6 as "cannot walk unassisted" - but those verdicts were issued under an earlier deployment stack (--heading did not exist yet, several fixes had just landed); only v7 had ever run under the current unified stack.

    Context

    New sim evidence sharpened the doubt: a same-harness four-version sweep showed v6 was the HEALTHIEST archive at the real operating point (cmd 0.15: dominant frequency locked at 2.50, steps 35:35 perfectly symmetric, foot distance 203/190 mm best of four, mean tilt 4.2 deg, saturation 15%). A full re-test under the unified stack was scheduled with per-version questions and a pre-filled interpretation table ("判读表(预填假设,回来对号)"): e.g. v6 walks + frequency ~2.5 -> old verdict was the stack's fault, v6 becomes the comparison champion; v6 walks but at ~1.25 -> period-doubling on hardware = confirmed plant gap, actuator fitting promoted to mainline; v5 no longer kicks -> the kicking was an old-stack artifact.

    Change

    All four versions re-queued on hardware under one stack (same torque limits, slew profile, heading loop, logging), with the version-specific legacy profile pinned; verdicts held provisional until re-issued.

    Outcome

    The re-test design separated policy properties from stack artifacts before any policy was permanently written off - and turned each outcome into a specific conclusion via the pre-filled table.

    Mechanism

    A deployed behavior is produced by the policy plus everything between it and the motors (heading loop, slew limits, torque caps, clock); verdicts implicitly condition on that whole stack. Fixing the stack invalidates the conditioning, so old failures may be stack artifacts and old successes may not survive either.

    Applies when

    • deployment tooling (limits, filters, loops) changed since a policy was last judged
    • deciding which historical policy is the rightful baseline
    • a sim sweep contradicts an old hardware verdict
    “只有 v7 在完整的今日部署栈下上过真机 … v5"左右乱踢"、v6"未能自主"的判决全部来自更早的栈(--heading 尚不存在, 部分修复刚落地)——判决已过期。且 2026-08-02 四代同机仿真横测翻出了新证据:v6 在真机工况(cmd 0.15)下是四代里最健康的仿真档案”
    train/REAL_SWEEP_V5_V8.md § 0. 为什么重测
  • A binary reward band on the swing knee had zero gradient everywhere below it, so the one-leg policy parked in an unloaded "fake touchdown" that Isaac's 5 N threshold scored as success and MuJoCo showed as real pressing - a capped constant-gradient ramp, retrained from scratch, passed 40/40binary-band-reward-fake-touchdown
    Mechanism understoodonelegreward-shapingreward-shapingsim2sim

    Shape approach-to-target rewards as capped ramps with gradient from the starting posture, never as bands or indicators; and compare contact-based terms across simulators, because a policy riding just under a force threshold looks perfect in one and wrong in the other.

    Symptom

    At iteration 1,000 of the first one-leg run the swing foot never lifted: the policy stood with the "raised" foot resting lightly on the ground. In Isaac the contact-match term paid 96% of full marks; the same policy in MuJoCo pressed that foot on the ground for 450 frames.

    Context

    The swing-leg goal was "shank folded fully back" (knee 1.5-1.95 rad), rewarded as a binary band: +0.8 inside [1.5, 1.95], zero elsewhere. From knee 0.05 to 1.5 rad the term was flat. Contact is judged at a 5 N force threshold, so a foot carrying less than 5 N counts as lifted. The walk line had hit the same disease with a binary indicator (v4) and fixed it with a capped ramp (knee_swing_amplitude).

    Change

    swing_knee_fold changed from the binary band to a ramp clamp(|q|/1.5, 0, 1) - a constant gradient capped near 86 deg - and the policy was retrained from scratch (V0r1). After the first real-robot try showed the fold still too low, its weight went 0.8 -> 2.0 (V0.1).

    Outcome

    V0r1 model_2300 passed the full acceptance 40/40 (swing knee 1.72 rad, about 98.5 deg) and was stamped as oneleg_v0.onnx; the cross-simulator disagreement is recorded as the thing that caught the cheat.

    Mechanism

    A reward that is flat until the target is reached gives no gradient to approach it, so the policy settles for the nearest state other terms reward - here, a foot that satisfies the contact threshold without lifting; a second simulator with different contact force resolution exposes such threshold-riding.

    Applies when

    • rewarding a posture target with an in-band / out-of-band indicator
    • a contact threshold decides whether a foot counts as lifted
    • trainer-side contact terms are near full marks while the video looks wrong
    “初版二值带 [1.5,1.95] 在膝 0.05→1.5 全程零梯度,策略停在"卸力虚点地"(Isaac 5N 阈下 contact_match 96% 满分 / MuJoCo 同策略 450 帧实压——跨仿真器互证抓作弊);v4 二值指示同型病,按 knee_swing_amplitude 判例改常数梯度封顶 ramp,从零重训 … **oneleg_v0.onnx = V0r1 model_2300, 40/40 PASS**”
    git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §4 奖励表 swing_knee_fold 行 / §8 核查单 5
  • Never referee a suspect metric with another metric from the same code - they can share the diseaseindependent-referee-for-metric-disputes
    Mechanism understoodomniattributionmeasurementattributionsim2simprocess

    To adjudicate a disputed measurement, compute the quantity by an independent method from raw state; never accept a sibling column from the same pipeline as the tiebreaker.

    Symptom

    A triple reversal on one question: sidewalk sign diagnosis (correct) was retracted using a second metric from the same script, then the retraction itself had to be retracted when that second metric turned out to be the buggy one - two opposite-direction errors on the same problem in one day, both written into the execution sheet.

    Context

    The probe's net-displacement metric suggested the sidewalk reference sign was inverted. Worried about yaw-drift pollution of net displacement, the author checked the same table's body-frame vy_mean column (~0.003 everywhere, 20-50x smaller) and retracted the sign diagnosis. But vy_mean came from mj_objectVelocity, which was silently reporting vertical velocity due to a frame bug - the "referee" was the diseased measurement. Re-measured with a truly independent computation (xmat.T @ qvel, world trajectory), the original diagnosis was confirmed: saw -0.5 gave vy +0.058/-0.130 (76%/106%), consistent with the net-displacement values all along (yaw pollution was real but only 10-21 deg, nowhere near reversal-sized).

    Change

    Lesson written twice, verbatim, as a hard rule: when questioning a measurement, the referee must be an independent algorithm (different code path, different physical derivation), e.g. rotate qvel by the body matrix directly, or inspect the raw world trajectory.

    Outcome

    With the independent referee in place the frame bug was confirmed, fixed, and the whole C4 line re-scored - revealing sidewalk had been working (see body-frame-velocity-api-audit).

    Mechanism

    Metrics sharing a code path (or an upstream API) share failure modes; agreement between them is evidence about the code, not the world. Only a measurement with an independent derivation can break the tie, because its errors are uncorrelated with the suspect's.

    Applies when

    • two metrics of the same quantity disagree
    • about to retract a conclusion based on a second readout
    • auditing evaluation code after a surprising result
    “我用一个坏指标去质疑一个好指标,并把撤回写进了执行单。教训(写死):质疑一个测量时,不能用同一份代码里的另一个测量当裁判 —— 它们可能同源同病。裁判必须是独立算法(这次的裁判应该一开始就是 xmat.T · qvel[:3],或直接看世界轨迹)。”
    train/C_LADDER_RUN.md § 3m. 二 我今天犯了两个方向相反的错 / 3n. 五 元教训
  • A single run's drift direction may be a limit cycle, not a policy bias - check the sign distribution across seedsmultiseed-sign-test-for-drift
    Mechanism understoodwalksim-evalmeasurementattributiongate-battery

    Distinguish "bias" from "broken symmetry limit cycle" by the sign distribution over many seeds; report drift as (mean, sign split), and never compare single-run drift magnitudes across versions.

    Symptom

    Net yaw over 15 s appeared to worsen from -41 deg (v2) to -84 deg (v4), inviting the conclusion that the new version drifted more.

    Context

    The Isaac-side view across 32 environments told a different story: per-env yaw was mixed-sign (20 negative / 12 positive) with mean ~0 - the drift is a limit cycle whose direction depends on initial conditions, not a systematic policy bias. The single MuJoCo run had sampled one draw from that distribution, so its magnitude could not be compared across versions as if it were a property.

    Change

    Evaluation rule: before classifying drift as systematic, run multiple seeds and examine the sign distribution; single-trajectory drift magnitudes are samples, not properties.

    Outcome

    The v2-vs-v4 drift "regression" was reclassified as not-established; later drift work (hip_roll l+r bias) used cross-policy, cross-seed evidence instead.

    Mechanism

    Symmetric dynamical systems can settle into either of two mirrored limit cycles; the selected cycle is decided by noise and initial state. A statistic whose sign is initial-condition-dependent has no meaning as a single sample - only its distribution does.

    Applies when

    • comparing heading drift or lateral drift across policy versions
    • a symmetric-looking behavior shows a consistent direction in one run
    • deciding whether to fix "drift" in reward or calibration
    “偏航反而变差(−41° → −84°):注意 Isaac 侧 32 env 的逐 env 偏航是正负混合(20/12)、均值 ≈0,说明这是极限环性质(方向随初值)而非策略偏置 —— MuJoCo 单次跑测到的是分布里的一个样本,不能当作系统性偏差。要判断需多种子统计。”
    train/WALK_DIAGNOSIS.md § walk_v4 独立验收 读法 (偏航)
  • Guessed joint friction was 2.5x low and damping 5x high - measure, then DR around nominalfriction-measured-not-guessed
    Mechanism understoodwalkplant-calibrationplant-calibrationdomain-randomization

    Measure frictionloss and damping separately (they need different rigs), put the measured value at DR center, and express DR as an additive band around that nominal - a DR range around a guessed value can exclude the real robot entirely.

    Symptom

    Old MJCF friction values were invented, not measured; when finally measured, every guessed value was wrong by a large factor in some direction.

    Context

    Joint friction split into Coulomb (frictionloss, tau_c) and viscous (damping b). Measured with the robot hung from a crane (吊机测) while armature was measured no-load; the two measurements are deliberately separated. Old MJCF: frictionloss 0.05, damping 0.1, DR joint_friction range [0, 0.1] "凭空拍的" (made up out of thin air).

    Change

    Replace guessed values with measured ones - frictionloss: RS06 0.15 / RS02 0.12 / RS00 0.13 N*m (old 0.05, i.e. 2.5x too low); damping: 0.02 N*m*s/rad on all three motor types (old 0.1, i.e. 5x too high). DR reshaped from an absolute made-up range [0, 0.1] to an additive band around measured nominal: joint_friction_add [-0.05, +0.10].

    Outcome

    "摩擦定稿(与 armature 一起, plant 参数第一次全部来自实测)" - friction frozen as part of the first fully-measured plant; DR now brackets a measured truth instead of spanning an invented interval.

    Mechanism

    Coulomb friction and viscous damping have opposite behavioral signatures (constant-torque threshold vs velocity-proportional drag); guessing both wrong in opposite directions gives a plant that is simultaneously too easy to start moving and too hard to move fast. DR centered on a wrong nominal makes the policy robust to a family of plants that does not contain the real one.

    Applies when

    • plant friction/damping values have no measurement provenance
    • DR ranges are absolute intervals rather than bands around a nominal
    • policy is over- or under-damped on hardware relative to sim
    “测关节摩擦。吊机测,而armature应该空机测试。… frictionloss τ_c (N·m) │ 0.15 │ 0.12 │ 0.13 │ 0.05(低 2.5×) … damping b (N·m·s/rad) │ 0.02 │ 0.02 │ 0.02 │ 0.1(高 5×) … DR │ joint_friction_add: [−0.05, +0.10] 叠标称 │ 旧 [0, 0.1] 凭空拍的”
    Experience.md § 摩擦定稿表 (lines 12-25)
  • The median of a bimodal metric lands in the empty gap - check the distribution, and never judge swing on 3 seedsmedian-hides-bimodal-distribution
    Replicatedomnisim-evalmeasurementgate-batteryprocess

    Before quoting a median or mean, look at the distribution; report suspected-bimodal metrics as mode share plus per-mode ranges, use small-seed smoke runs only to screen trends, and size the seed count for decisions by the share resolution you need (here: 20).

    Symptom

    Years of "high swing variance" and undecidable 3-seed swing readings turned out to be one fact: the metric was bimodal all along - "历代 swing 高方差与 3-seed 不可判由此定性:一直在测双稳态系统" - and every median reported from it (e.g. 12.1 mm) described a value no seed ever produced.

    Context

    Concrete instances: s2e_pd-1400's 20 seeds split 2.6-4.9 mm vs 19.3-24.0 mm with zero seeds between; s1e-500 read 3.9 mm on seeds 0-2 but 23.1 mm median over 20 seeds; a 3-seed reading of 19.3 was logged as "double-peak optimism, lesson recurrence #4". The selection re-audit codified the sampling rule: "3-seed 的 swing 读数不可判点, 只能筛带,选点必须 20-seed" - 3 seeds may screen a band, only 20 seeds may pick a point.

    Change

    Swing (and any suspected multi-modal metric) reported as mode shares plus per-mode ranges instead of a bare median; 3-seed smoke numbers demoted to band-screening; all shipping/selection decisions moved to 20-seed batteries.

    Outcome

    The "swing debt" bookkeeping was reinterpreted as basin probability (see swing-bistability-damping-switch), and checkpoint selection stopped being whipsawed by which basin the first three seeds happened to fall into.

    Mechanism

    Central-tendency statistics presuppose unimodality; on a bimodal distribution the median tracks the mode SHARE, not any achievable behavior, and small samples alias the share entirely. Mode-aware reporting (share + per-mode stats) is the only faithful summary, and the needed sample size is set by the share resolution required.

    Applies when

    • a quality metric shows chronic high variance across seeds
    • 3-seed smoke readings contradict 20-seed batteries
    • reporting swing height, clearance, or any basin-prone metric
    “中位数落在空档里,「swing 债 −11mm」实为「50% 概率掉进拖地吸引子」。历代 swing 高方差与 3-seed 不可判由此定性:一直在测双稳态系统。… swing 跨 seed 双峰 (500 在 seed0~2 只读 3.9mm, 20-seed 中位 23.1) —— 3-seed 的 swing 读数不可判点, 只能筛带, 选点必须 20-seed。”
    train/README.md § swing 双稳态定性 / s1e 选点重审
  • Every rung gets written stop criteria - hit any one, stop; tightened when priors say results should come fastpreregistered-stop-criteria-per-rung
    Replicatedomnitraining-runprocessgate-battery

    Freeze per-rung stop criteria (old-skill floors, new-skill deadline, oscillation signature, known pathology signatures) before training, stop on first hit - and shorten the deadline in proportion to how fast your mechanism says results should appear.

    Symptom

    Continued training past the point of degradation had previously destroyed capabilities (C1 trained away the root's backward skill); without hard stop rules, sunk-cost reasoning keeps runs alive too long.

    Context

    Standard rung protocol - eval_c_matrix every 100 iters at 5 seeds with a fixed criteria list, frozen before training ("开训前写死,事后不许改"): (1) stand or vx+0.15 at <=3/5 for two consecutive points; (2) vx+0.15 tracking <70% for two consecutive points (vs recorded root baseline); (3) oscillation signature - survival repeatedly crossing zero between adjacent checkpoints, stop on first occurrence, do not wait for confirmation; (4) new-skill-no-progress deadline (e.g. wz tracking <30% or still same-signed after iter 400); (5) a known lineage pathology signature (s1c arm-B: scatter -> half-recover -> collapse).

    Change

    Stop budgets scale with prior knowledge: when C4-redo4's feed-forward was already proven open-loop, the no-progress deadline was tightened from +500 to +100 iters ("前馈是开环就给 71~96% 的 … 若 +100 还没有 … 不值得再烧").

    Outcome

    Multiple rungs were stopped exactly on criterion (C4-redo criterion 4 at iter 1200; redo2/redo3 at absolute 900), converting each into a clean hypothesis test instead of a drifting run; no rung burned its full cap on a dead hypothesis.

    Mechanism

    Degradation of retained skills is a lineage-level injury that later training often cannot undo, so detection latency is capital loss; and a stop rule written before the run cannot be bent by hope. Tightening deadlines when priors predict immediate results converts "no result yet" into evidence against the mechanism.

    Applies when

    • starting any resumed/curriculum training rung
    • deciding whether to keep training a run that shows early regression
    • a mechanism-backed change should produce results immediately
    “每 100 iter eval_c_matrix.py --seeds 5,命中任一即停 … 存活在相邻 checkpoint 间反复跨越 0(出现即停)… 收紧版:绝对 iter 900(= +200)时 vy±0.10 跟踪仍 <30% 即停”
    train/C_LADDER_RUN.md § 3h. 预注册停梯判据(开训前写死,事后不许改)
  • The walk phase machine structurally cannot express flight - rebuild the representation for duty < 0.5, teach flight with a mask tax, never a cliff bountyphase-machine-structural-limits
    Mechanism understoodrunreward-shapingreward-shapingcurriculumgate-battery

    When a new gait changes the contact pattern's structure, audit whether the phase/mask representation can express it and rebuild the representation if not; teach the new contact pattern with graded mask-mismatch pressure and count it in acceptance with artifact-proof definitions (minimum segment length), never with cliff bounties.

    Symptom

    Running requires both feet airborne, but the walk-era phase machine switches legs by the sign of sin(phase) - with duty < 0.5 the two swing windows must OVERLAP during flight, which a sign-switching representation cannot express at all.

    Context

    The run phase machine was re-architected rather than patched: per-leg phases (left = phi, right = phi+0.5 mod 1) with leg_phase < duty defining stance, aligned to the walk sin convention at duty=0.5 so the machines agree where their domains overlap. Flight is taught by the SAME mechanism that once cured foot-dragging, direction reversed: in the two planned flight windows the contact mask is (0,0) and feet_contact_number_duty charges -0.3 per foot still on the ground - a mild ~0.16/step tax, deliberately NOT a cliff: "悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文". The reference shape (half-sine bump over swing progress) is zero at window boundaries by construction, eliminating the clearing-window that the C4 probe measured to cost 13-19% on non-sinusoidal references. The shape self-check ("零代价选项是什么") was run on three behaviors: standing pays ref everywhere (known cmd=0 stepping risk, booked), walking pays only the flight-window tax, proper running collects full marks.

    Change

    New rewards.py run section (leg_phase_duty / stance_mask_duty / ref_run / clearance_run / contact_duty) with the walk versions untouched byte-for-byte; flight acceptance metric defined with a segment-length floor (>=40 ms to count) so numeric contact flicker cannot fake flight.

    Outcome

    Flight became expressible and taught by a calibrated mild pressure; the walk lineage's phase code stayed frozen as its own contract.

    Mechanism

    A phase representation defines which contact patterns exist in the reward's vocabulary; duty cycling below 0.5 introduces states (double-flight) outside a half-period sign convention's language, so no weight tuning can teach them. And rare desirable events taught by cliff-shaped bounties invite hacks (jumping in place); the graded mask tax prices the planned pattern without creating a jackpot.

    Applies when

    • extending a walking stack to running/jumping (duty < 0.5)
    • a desired contact pattern never appears despite reward increases
    • defining flight/contact acceptance metrics
    “duty<0.5 时摆动窗 (1−duty)T > T/2,两腿摆动窗在腾空段重叠 —— sin 符号切腿的机制结构上表达不了"双脚同时在空中"。… feet_contact_number_duty 对"还踩着地"持续 −0.3/脚 —— 与 walk 治拖地同一机制,方向相反。不腾空的税 ~0.16/步 … 梯度温和不构成悬崖(悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文)。”
    train/RUN_V0_SPEC.md § 5. 相位机设计
  • Set torque limits per joint from measured gait peaks - a uniform percentage is the wrong shape, and training must use the deployed numberstorque-limit-shape-by-measured-peaks
    Mechanism understoodwalkactuator-modelingactuator-modelinghardwareplant-calibration

    Measure per-joint torque peaks in the actual gait and set each limit as measured-peak x margin capped at rating; then propagate the same numbers into training and add an automated deploy-time consistency check - never derate by a uniform percentage, never let training assume torque deployment will not grant.

    Symptom

    A uniform 50% torque derating (18/8.5/7) had piled safety margin on the joints that never use it while cutting the busiest joint below half its measured demand.

    Context

    Per-joint gait peaks were measured (walk_v5 at cmd 0.3/0.6): RS06 (hip_pitch/knee) uses 5.5-5.9 N*m = 15-16% of its 36 N*m rating - cutting it to 12 is a free safety win; RS02's ankle_pitch runs at 16.2 N*m = 95% of its 17 N*m rating - "它是速度的硬件瓶颈", no room to cut; RS00 measured 36-44%, capped at 11. The resulting shape 12/17/11 replaced the uniform percentage. Sweeps across several limit sets (rated / 50% / 14-17-11 / 12-17-11) produced identical speed, lift, and landing force - within this range the limits do not shape the gait; what matters is consistency: "关键是训练和硬件必须是同一个数", because the exporter fills effort_limit from tau_limit, and a policy trained at rated 36/17/14 "会假设有三倍力矩可用" while deployed at 12/17/11 (exactly the v5 cross-generation inconsistency later suspected in its wild kicking).

    Change

    robot.yaml tau_limit set to the measured-shape 12/17/11, firmware written to match, and train/isaac_values.py regenerated so training sees the same limits; the deploy tool self-checks limits against robot.yaml on every run.

    Outcome

    Free safety margin captured where demand is low, the real bottleneck joint left at rating, and the train/deploy torque worlds unified with an automated consistency check.

    Mechanism

    Torque demand is grossly unequal across joints in a gait (15% vs 95% of rating here); a uniform percentage misallocates the safety budget by construction. And since the trainer treats effort_limit as a plant truth, any train/deploy mismatch is an invisible plant gap of exactly the mismatch ratio.

    Applies when

    • choosing safety torque limits for a legged platform
    • training-vs-deployment actuator limit audit
    • one joint runs near rating while others idle
    “曾用统一 50%(18/8.5/7)是错的形状: 把余量堆在用不到的 RS06 上, 却把 ankle_pitch 砍到需求的 52%。… RS02 在 0.6 m/s 已用到额定 95%, 它是速度的硬件瓶颈 … 实测多组限幅 … 完全一致 —— 限幅在这个范围对步态零影响, 关键是训练和硬件必须是同一个数。… 若训练仍按额定 36/17/14, 学出的策略会假设有三倍力矩可用。”
    train/WALK_V6_MINIMAL.md § 3. 训练侧必须同步的一件事
  • Late training leaned on sampling noise as a stability crutch - deterministic play collapsed while training metrics stayed greennoise-crutch-deterministic-collapse
    Replicatedomnitraining-runsim2simprocess

    Evaluate the deterministic policy in an external harness on a fixed cadence during training (not just at the end), select checkpoints on that curve, and treat a collapsing noise_std with rising training reward as a warning that noise is load-bearing.

    Symptom

    omni_s1's final checkpoint (model_5999) fell at 4 s even in Isaac's OWN deterministic play, while checkpoints from iter 1700-4000 were fine - and no training metric flagged anything. Policy noise_std had collapsed to 0.045 by iter ~990 (final 0.033).

    Context

    Diagnosis: the policy had learned to use its exploration noise as a dither/stabilizer - "策略把采样噪声当稳定拐杆,训练指标看不见" (the training metrics cannot see it, because training always runs with noise on). Countermeasures: entropy_coef 0.005 -> 0.01 to slow the std collapse, and - the structural fix - an in-training smoke loop (watch_ckpt.py): every 500 iters, export ONNX directly, run 3-seed MuJoCo evaluation, log CSV/TensorBoard curves plus three-view videos. The doctrine line was written in bold: "训练指标全绿不再是发育健康的 证据,冒烟曲线才是" - green training metrics are no longer evidence of healthy development; the smoke curve is. The follow-up run s1b showed the drift metric follow a U-shape (73 -> 8.6 at iter 3500 -> 76), making checkpoint selection BY the smoke curve (early stop at 3500) the shipping mechanism, with terminal re-degradation booked as known and unresolved.

    Change

    entropy floor raised; watch_ckpt smoke loop instituted as standing infrastructure; checkpoint selection moved from "last iteration" to "best point on the deterministic smoke curve".

    Outcome

    s1b shipped from iter 3500 (the U-bottom) instead of a degraded terminus; every later lineage (s1c/s1e, the C ladder's --every 100 loops) inherited the watcher as the standard guardrail.

    Mechanism

    PPO evaluates and improves the stochastic policy; if noise itself stabilizes the gait (dither smoothing a marginal limit cycle), the deterministic mean policy is a different, worse controller that training never measures. External deterministic evaluation on an independent simulator is the only readout of what will actually be deployed.

    Applies when

    • final checkpoints underperform mid-training ones
    • noise_std collapses early while training reward climbs
    • deciding which checkpoint to export and ship
    “训练后期确定性脆化——noise_std iter~990 收到 0.045(终 0.033),model_5999 连 Isaac 确定性 play 都 4 s 摔(1700~4000 正常):策略把采样噪声当稳定拐杖,训练指标看不见。对策:entropy_coef 0.005→0.01 + train/watch_ckpt.py 训练中冒烟曲线 … 训练指标全绿不再是发育健康的证据,冒烟曲线才是。”
    train/OMNI_V0_SPEC.md § 3. S1.1 修订记录 ②
  • Audit rewards by realized contribution (weight x achieved value) - a weight of 2.0 was really paying 0.04realized-contribution-audit
    Mechanism understoodwalkreward-shapingreward-shapingattribution

    Evaluate a reward table by each term's realized per-step contribution under the current policy, never by its weight column; if a term's realized value is ~0, escalating its weight is a no-op - change the term's structure instead.

    Symptom

    Foot dragging persisted through repeated weight escalation: feet_air_time had been raised 0.25 -> 1.0 -> 2.0 across versions with no behavioral change, and the training-side comment even recorded the fact ("几乎没有 单支撑相, 是拖着脚蹭") without the fix changing form.

    Context

    Computing realized per-term contributions in the trained state exposed the economy: feet_air_time contributed weight 2.0 x achieved 0.019 = 0.038 per step, against tracking's +1.20 - lifting the leg earned 3% of what tracking earned, so dragging was the rational optimum no matter the weight escalation. The same table acquitted the energy penalties (total -0.30 negative vs +1.74 positive) that a naive read of weights (-5.0 orientation!) would have blamed.

    Change

    Fix redirected from "raise the weight again" to "add a term whose realized contribution changes the optimum": a clearance penalty sized so its realized magnitude (~0.018/foot when dragging) is comparable to feet_air_time's, enough to flip the optimum without drowning tracking.

    Outcome

    With the term economy corrected (plus posture/range fixes), swing height reached 34 mm and tracking 87% by v6; weight escalation of the old term was abandoned.

    Mechanism

    A reward weight is only a multiplier on whatever the policy currently achieves on that term; when the achieved value is near zero (behavior absent), escalating the weight multiplies near-zero. Optimizer behavior is governed by realized per-step magnitudes, so audits must be conducted in that currency.

    Applies when

    • a behavior persists despite repeated weight increases
    • auditing whether penalties are "too strong" or rewards "too weak"
    • sizing a new reward term against existing ones
    “把 feet_air_time 权重从 0.25 → 1.0 → 2.0 一路加,但没有加高度项。量级算下来:feet_air_time 权重 2.0 × 实得 0.019 = 0.038,而跟踪奖励是 1.2。抬腿的边际收益只有跟踪的 3%,拖地当然是最优解。… 正项 +1.74,负项 −0.30。能量惩罚不是瓶颈,抬腿没收益才是。”
    train/WALK_DIAGNOSIS.md § 决定性证据(#1) / 各项奖励的实际量级
  • Fix a too-deep nominal pose before adding any penalties - the default stance defines the basin training starts innominal-posture-before-penalties
    Mechanism understoodwalkreward-shapingreward-shapingplant-calibration

    Before tuning penalties on a degenerate gait, audit the nominal pose and height targets against morphology and published ratios; if the default stance encodes the degenerate behavior, fix it first - and recompute dependent quantities (init height) by FK, not by hand.

    Symptom

    Policy lived in a crouched shuffle; nominal knee angle was 0.5 rad (28.6 deg) - deeper than published configs (Unitree G1 0.3 rad / 17.2 deg, Booster T1 0.4 rad) - so the policy's starting point and its action-space center both sat inside the crouch basin.

    Context

    Initially ranked "secondary" in the local diagnosis, this was promoted to co-first priority by the cross-check against published reward tables, which states that with nominal knee flexion above ~0.4 rad, fixing the posture must precede adding any penalties ("改这个之前别加任何 惩罚都是白费"). Companion base-height items: walk profile had weakened base_height_l2 to -5.0 (base class -10, field standard -10 to -20, "the second most common cause of death"), and the height target must be the STANDING height (0.384), not the crouch height.

    Change

    Nominal knee 0.5 -> 0.3 rad with init_base_height recomputed by MuJoCo FK (0.3739 -> 0.3802); base_height_l2 restored to -10 with standing height target; both bundled as first-priority alongside the clearance term.

    Outcome

    Part of the v5/v6 package that lifted swing height to 34 mm and tracking to 87%; the crouch basin stopped being the default answer.

    Mechanism

    The nominal pose is the fixed point every regularizer pulls toward and the point where action=0 lands; if that point is itself the degenerate posture, every penalty fights the geometry. Correcting the attractor is prior to shaping the gradient field around it.

    Applies when

    • policy converges to a crouched or collapsed posture
    • nominal joint angles were chosen for stability rather than gait
    • base-height reward targets or weights were locally weakened
    “研究明确说"nominal 膝屈超过 ~0.4 rad 必须先改,改这个之前别加任何惩罚"。我们是 0.50,超标。… base_height_l2 在 walk profile 里被减到 −5.0(基类是 −10)。研究说这是"第二常见死因"且应 −10 ~ −20。改回 −10。目标高度用站立高 0.384 是对的(研究要求 target 必须是*站立*高度而非蹲姿)。”
    train/WALK_DIAGNOSIS.md § 修正 ①(升级优先级) / 修正 ④
  • Brief the operator on the lineage's measured zero-command and untrained-axis behavior before handing over the joystickknow-zero-command-behavior
    Replicatedomnireal-deployreal-acceptanceprocess

    Before any teleop/demo, measure and write down the policy's zero-command behavior and per-axis competence, label untrained axes explicitly as not-bugs, and set the floor/procedure to accommodate the known drift.

    Symptom

    A teleop session was about to start on a policy that does not stand still at zero command and has never been trained on lateral commands - behaviors an unbriefed operator would report as bugs or emergencies.

    Context

    Three measured facts were written into the teleop instructions ("都有 实测依据, 不是猜"): (1) A/D (lateral) keys will get essentially no response - probe-measured sidewalk tracking ~3%, an untrained axis: "这正是 C4 要解决的事, 不是 bug"; (2) no keypress = cmd 0, and this lineage does not stand still at zero command - a three-generation lineage property: paces in place, drifts right ~5 cm/s, net rotation -30 deg/20 s; sim survival is 20/20 (it will not fall) but it walks away slowly, so leave floor margin especially on the right; (3) S (backward) WILL respond - probe-measured 20/20 survival, 67% tracking untrained, which is also why this root was chosen for the C ladder. Plus a keybinding dry-run while suspended before touching down.

    Change

    Operator briefing became part of the deployment artifact: expected response per key, expected idle behavior with magnitudes and directions, and the distinction between untrained (expected, not a bug) and abnormal.

    Outcome

    The session proceeded with correct interpretations available in advance; the known zero-command wander was handled by floor margin and start-with-command procedure rather than misdiagnosed on the spot.

    Mechanism

    A learned policy's off-nominal behaviors (idle drift, untrained axes) are lineage properties, stable and measurable in sim beforehand; operator surprise converts known properties into false incident reports and unsafe reactions. A briefing transfers the measured behavior model to the person holding the controller.

    Applies when

    • handing a learned policy to an operator or demo audience
    • the policy idles in a non-stationary way at zero command
    • some command axes are untrained in the current lineage
    “A/D 基本不会有反应 —— s1e 从未训过非零 vy, 选根探针实测侧走跟踪率 ~3% … 这正是 C4 要解决的事, 不是 bug。… 不按键 = cmd 0, 而 s1e 在零指令下不站定 —— 血统属性, 三代实录: 原地踏步 + 右漂 ~5 cm/s + 净旋 −30°/20s。”
    train/REAL_RUN_S2.md § 附: WSAD 遥控 上机前必须知道的三条
  • Freeze the deployment contract, stamp every export, and let an automated checker catch wiring bugscontract-freeze-and-checker
    Replicatedomniprocesscontract-freezeprocesssim2sim

    Freeze and fingerprint the policy I/O contract; ship contract changes as new versioned profiles that leave old artifacts bit-identical; and extend the automated contract checker with every pipeline change, forcing the new path to execute in the check.

    Symptom

    Contract-level changes (observation layout, action pipeline) are where silent sim/real divergence is born; two real wiring bugs appeared the one time the action pipeline was extended.

    Context

    The 215-dim observation contract was frozen ("纪元 3,三机 digest" - an era number plus a digest agreed across three machines); proposals that would break it (e.g. a GRU memory) were rejected on contract grounds. Every exported ONNX is stamped and verified with a manifest (onnx_manifest --stamp / --verify), and deployment refuses mismatched combinations. When C4 added the lateral feed-forward, it went in as a NEW profile (omni_ff) leaving the existing omni profile's behavior bit-identical; the checker (check_contract) was extended to force the feed-forward path to actually execute (cmd_vy=0.13) and promptly caught two genuine bugs: (1) re-clamping with soft_joint_pos_limits after the feed-forward (0.23 rad deviation) instead of reusing the parent's clip; (2) indexing processed actions by asset.joint_names instead of the action term's own contract-ordered _joint_names, which landed the feed-forward on the wrong joints (l_hip_yaw / r_ankle_pitch).

    Change

    Contract discipline as implemented: frozen dims + digest; manifest stamping and refusal; contract changes only via new versioned profiles; checker updated in the same commit as any pipeline change, with inputs chosen so new code paths are exercised.

    Outcome

    Both wiring bugs caught before any training or deployment ("两个都是 check_contract 当场抓出来的 —— 这次它值回票价"); old deployments provably unaffected by the new profile.

    Mechanism

    The contract is the only interface the policy and robot share; freezing plus fingerprinting makes divergence detectable, and an executable checker turns "the contract holds" from a belief into a test - but only if its inputs actually drive the new code path.

    Applies when

    • modifying the action or observation pipeline of a deployed policy
    • exporting policies for hardware
    • proposals that would change observation dims or history structure
    “契约校验抓到的两个真错误(记账,别再犯):1. 前馈后误用 soft_joint_pos_limits(URDF 限位 ×0.9)重钳 → 0.23 rad 偏差 … 2. 用 asset.joint_names 索引 _processed_actions → 前馈落到 l_hip_yaw/r_ankle_pitch 上 … 两个都是 check_contract 当场抓出来的 —— 这次它值回票价。”
    train/C_LADDER_RUN.md § 3j. 契约级改动 / 契约校验抓到的两个真错误
  • A ratio metric flipped the verdict - spectral share rose while absolute high-frequency energy fell 16%ratio-metrics-need-absolute-check
    Mechanism understoodwalksim-evalmeasurementattributionprocess

    Never compare share/percentage/centroid metrics across conditions whose totals differ; pair every ratio with its absolute numerator before issuing a verdict, and log retracted judgments so they are not re-derived.

    Symptom

    walk_v6 was provisionally judged "more jittery" than v5 because the joint-velocity spectral centroid rose 2.91 -> 3.62 Hz and the >4 Hz energy share rose 12.6% -> 17.8%.

    Context

    Absolute measures said the opposite: first differences of actions fell 1.54 -> 1.31, second differences 2.55 -> 2.19, and absolute high-frequency energy fell 16% (0.490 -> 0.413). The shares and centroid rose only because low-frequency content fell even more - the denominator shrank. The interim judgment was retracted in writing so it would not be reused.

    Change

    Metric discipline noted: "占比类指标在总量变化时不能直接比较" - share/ratio metrics are not comparable across conditions when the total changes; verdicts about smoothness must cite absolute energies or difference norms.

    Outcome

    v6 correctly classified as smoother, not jitterier; the retracted judgment logged under "被推翻的一个中间判断(记下来免得复用)".

    Mechanism

    A ratio confounds numerator and denominator; any intervention that removes low-frequency content raises every high-frequency share without adding a single joule of jitter. Only absolute quantities support cross-condition comparison when totals move.

    Applies when

    • comparing smoothness/jitter/spectral metrics across versions
    • any percentage-based metric moves after an intervention
    • writing an eval report that includes normalized quantities
    “我一度说"v6 动作更抖" … 错了: 动作一阶差 1.54→1.31、二阶差 2.55→2.19 都在降 … 谱质心升高只是因为低频成分掉得更多, 绝对高频能量实际下降 16%(0.490→0.413)。占比类指标在总量变化时不能直接比较。”
    train/WALK_DIAGNOSIS.md § 过程中被推翻的一个中间判断(记下来免得复用)
  • Add a termination that makes the degenerate strategy fatal - no height cut-off meant crouch-shuffling could live forevertermination-closes-degenerate-basin
    Observed oncewalkreward-shapingterminationreward-shaping

    For each known degenerate strategy, check whether the termination set makes it fatal; if the robot can live indefinitely inside the degenerate posture, add a termination just past the intended operating envelope rather than escalating penalties.

    Symptom

    Crouched foot-dragging survived indefinitely because the termination set contained only bad_orientation (40 deg) and base contact - there was no height termination at all, so a deep squat was a viable long-term strategy.

    Context

    The hypothesis audit found the missing termination (hypothesis 7); the cross-check against published configs found the field practice: Booster terminates at 0.45 m (38% of body height) and the research warning is that the termination height must not be so low that crouching survives it. The proposed value: 0.32 m, just below the walk crouch base height 0.3739 - a deep squat terminates immediately, "断掉蹲着蹭的活路" (cutting off the crouch-shuffle's livelihood).

    Change

    Add height termination at 0.32 m as a second-priority item of the walk fix package, alongside restoring base_height_l2 to -10.

    Outcome

    Entered the v5/v6 fix package under which the crouch-shuffle optimum disappeared (34 mm clearance, 87% tracking by v6).

    Mechanism

    Termination conditions define which strategies exist at all: a reward penalty prices a behavior, but a termination deletes its future returns entirely. Degenerate basins that are merely penalized can remain optimal under enough tracking pressure; a termination placed between the degenerate posture and the intended one makes the basin unreachable as a steady state.

    Applies when

    • a degenerate but stable behavior persists across reward tunings
    • auditing termination conditions for a locomotion task
    • a policy exploits the gap between penalized and terminated states
    “加终止高度:研究第 6 条"终止高度不能低到让蹲着也能活"。我们完全没有高度终止。建议 0.32 m(略低于 walk 蹲姿基座高 0.3739,深蹲即终止)。… 加终止高度 0.32 m(深蹲即终止,断掉蹲着蹭的活路)”
    train/WALK_DIAGNOSIS.md § 修正 ④ / 最终改动清单 第二优先
  • Verify changes in the run's resolved config (and checkpoint md5), never in the source you editedresolved-config-is-source-of-truth
    Replicatedomniprocessattributionprocesscontract-freeze

    Attribution and single-variable claims must be made on the resolved per-run config (and checkpoint hashes), not on source diffs; verify every intended variable landed before burning compute, and verify every rollback byte-level against the historical resolved config.

    Symptom

    An intended arm-B config change never reached the training run - the run was grid-identical (117/117 cells) to its C2 predecessor - and the burn was only understood afterwards.

    Context

    The repo's discipline hardened around the logged resolved config (logs/<run>/params/env.yaml) as the only source of truth: (1) the C2 root-cause analysis was performed against the checkpoint's logged env.yaml, not the code ("以真相源 23-19-25/params/env.yaml 核实"); (2) C4 added a pre-flight: grep the landed env.yaml for the new keys, and compare the first checkpoints of the two arms - identical md5 means the variable did not land, stop immediately; (3) the C4 full rollback was accepted only after starting a 1-iter run and byte-comparing its resolved env.yaml against the historical 700-era file (identical except 4 dormant schema fields, each verified to be at its no-op default).

    Change

    Standing pre-flight and post-change verification: dump/diff the resolved config that the run actually consumed; use checkpoint hash equality as a cheap "variable landed" detector between arms.

    Outcome

    Caught the not-landed variable class of failure; made the rollback provably equivalent to the historical training state rather than believed-equivalent.

    Mechanism

    Between edited source and the running experiment sit layered overrides, env-var switches, and registration logic; only the resolved, serialized config reflects their composition. Diffing at that level tests the actual experiment; diffing source tests intent.

    Applies when

    • launching an A/B pair or any single-variable rung
    • rolling back to a historical training state
    • a run behaves as if a change was never applied
    “开训前先验落盘 cfg(上一轮臂B 的改动没进 run,与 C2 逐格 117/117 相同):grep -E "base_com|joint_friction|push_robot|track_lin_vel_y_exp" logs/<run>/params/env.yaml 另:两臂第一个 checkpoint 的 md5 若相同 = 变量没进去,立刻停。”
    train/C_LADDER_RUN.md § 3d. ⚠️ 开训前先验落盘 cfg / 3l. 回退清单(验证)
  • The action-delay was implemented as lerp - beyond one step it extrapolated BACKWARD, so a whole lineage trained on a fictitious actuatorlatency-lerp-reverse-extrapolation
    Mechanism understoodinfraactuator-modelingactuator-modelingplant-calibrationsim2sim

    Unit-test plant-model code (delays, filters, randomizers) against hand-computed truth across its FULL configured range, not just the nominal case - a delay must be a queue, and any interpolation used outside [0,1] is a silent plant corruption that training will faithfully absorb.

    Symptom

    Every walk/stand model up to v11 had been trained on a silently wrong plant: the action-latency implementation lerp(cur, prev, lag) is only an interpolation for lag <= 1 - at lag 3 it computes 3*prev - 2*cur, a REVERSE extrapolation. With the configured (0, 0.06) s at 50 Hz (lag in [0,3]), about 2/3 of environments were adapting to actuator dynamics that do not exist.

    Context

    Listed as evidence item #1 for the full restart: "全部旧模型训在错误 plant 上" and the head suspect for the real robot's wild kicking. The fix replaced it with a true FIFO delay line (commit 7f13793) plus its own regression test (tests/test_action_latency.py) - but every exported ONNX predated the fix, which is part of why the lineage was frozen rather than patched.

    Change

    Delay implementation rewritten as an honest FIFO with unit tests; the restart baseline trained on the corrected plant from day one.

    Outcome

    A generation-scale training investment was revealed to have a corrupt plant underneath; the class of bug (plausible-looking math that silently changes meaning outside its valid range) got a permanent test.

    Mechanism

    lerp(a, b, w) leaves the segment for w > 1; used as a delay it fabricates high-gain inverted dynamics precisely in the largest-delay draws, so the policy learns compensation for an actuator that cannot exist - and DR then trains robustness to the artifact rather than to reality. No training metric can catch this: the sim is self-consistent, just wrong.

    Applies when

    • implementing or auditing action delay / filtering in a trainer
    • a lineage behaves as if compensating dynamics nobody modeled
    • deciding whether old checkpoints are salvageable after a plant bug
    “动作延迟旧实现 lerp(cur, prev, lag) 在 lag>1 时是反向外插(w=3 → 3·prev−2·cur),配置 [0,0.06]s@50Hz 即 lag∈[0,3],约 2/3 env 在适应不存在的执行器动态。7f13793 已换真 FIFO,但所有 ONNX 均训于修复之前 —— 真机"乱踢"的头号嫌疑。”
    train/OMNI_V0_SPEC.md § 0. 为什么从零 (1)
  • Export every CAD part in the whole-machine frame so URDF rotations are zero and inertia is exacturdf-shared-origin-export
    Observed onceinfraplant-calibrationplant-calibrationhardwareprocess

    Generate the model so that correctness is structural: shared-origin STL export, zero rotations, subtraction-only origins, and an explicit 1e-9 g*mm^2 -> kg*m^2 conversion - never hand-rotate inertia tensors.

    Symptom

    Hand-assembled URDFs accumulate per-link rotation/origin errors and unit-conversion mistakes in inertia tensors - silent plant corruption that no later calibration can cleanly fix.

    Context

    Documented CAD -> URDF -> USD procedure from a successful Isaac Lab deployment, kept as the recipe if Lucen regenerates its model.

    Change

    (1) In CAD, align the whole robot to Z-up, X-forward (Isaac Lab convention) and ground the assembly; (2) export each STL with other parts hidden but the machine's shared origin kept, so all parts share one origin, every URDF rotation is 0, and inertia matrices equal CAD values directly; (3) units: Fusion 360 gives g*mm^2, URDF wants kg*m^2 - multiply by 1e-9; (4) link origin = negative of the joint position; link COM = CAD COM minus joint position; joint origin = difference of the two joint positions; (5) after URDF -> USD import, open the USD separately and set it instanceable before saving.

    Outcome

    A URDF whose rotations are all zero and whose inertia tensors are CAD-exact, eliminating an entire class of hand-transcription plant errors.

    Mechanism

    Keeping one shared origin turns every frame transform into a pure translation computable by subtraction, and leaves inertia tensors in the frame CAD already computed them in - no rotation of inertia tensors, the most error-prone manual step, is ever needed.

    Applies when

    • building or regenerating URDF/MJCF from CAD
    • inertia or frame bugs suspected in the plant model
    • importing URDF into Isaac Lab / USD
    “导出 STL 时隐藏其他零件但导出整机——这样所有零件共享同一原点,URDF 里所有 rotation 全是 0,惯量矩阵直接等于 CAD 值 / 单位:Fusion 360 给 g·mm²,URDF 要 kg·m²,乘 1e-9 / link origin = 该关节坐标取负 … URDF → USD 导入后必须单独打开 USD 设成 instanceable 再存”
    Experience.md § URDF 制作流程 (lines 87-92)
  • FK-verify a borrowed reference's structure, then size its amplitude by the reference's job - it pins phase, the policy adds liftreference-structure-fk-amplitude-division
    Mechanism understoodwalkreward-shapingreward-shapingcurriculumplant-calibration

    When borrowing a reference trajectory: verify its structural claim against your own kinematics (an invariant like flat-foot), assign it the phase-pinning job, and size amplitude low enough that the policy contributes the lift - moving toward a proven foreign value in halves, not jumps.

    Symptom

    walk_v4 had big knee swing (40-46 deg) but only 18-24 mm foot lift - amplitude without hip/knee/ankle phase coordination; later, walk_v5's real-robot swing ballooned to 73.6 deg (sim 55.7) with violent footfalls - amplitude over-driven by the reference.

    Context

    Structure first: Humanoid-Gym's 1:2:1 hip:knee:ankle reference was verified on the local model before adoption - the ratio exactly satisfies the locally derived flat-foot constraint hip - knee + ankle = 0, FK-tested at multiple amplitudes with sole pitch 0.00 deg throughout. Amplitude second, and here the first reasoning failed honestly: FK said shorter legs need LARGER reference scale (0.30 for 30 mm lift), and the FK was correct - but the premise was wrong ("FK 没错, 但前提错了"): it assumed foot lift must come from the reference. HighTorque Pi, same scale, uses 0.08 with a 0.02 m foot-height target - proof that lift is added by the policy ON TOP of the reference, whose actual job is pinning the phase relationship. Scale 0.30 made the reference the entire gait: over-constrained and over-driven. The correction went to 0.15, deliberately not Pi's 0.08: "一次只走一半, 留退路" (walk half the distance, keep a retreat).

    Change

    target_joint_pos_scale 0.30 -> 0.15 as one of v6-minimal's three changes, treating both the footfall force and the lateral kicking (yaw momentum scales with leg swing amplitude).

    Outcome

    v6 improved landing force 1.72x -> 1.55x, suspended tilt 45.9 -> 23.0 deg, turn-gain asymmetry 70% -> 19%; the later v6-halved-shaping experiment (35 mm -> 4 mm collapse) confirmed the reference still carries the gait's existence on this machine - the division of labor is real but machine-specific.

    Mechanism

    A joint-space reference plays two separable roles: encoding structure (phase relations that keep the foot flat) and injecting amplitude (energy). Structure transfers across robots and is checkable by FK against an invariant; amplitude is a negotiation with the policy, and over-assigning it to the reference removes the policy's freedom to modulate lift with state.

    Applies when

    • importing a reference gait / imitation target from another codebase
    • reference amplitude reasoning based on leg length alone
    • real swing amplitude far exceeds sim's under a strong reference
    “FK 没错, 但前提错了。我默认抬脚必须由参考轨迹产生。HighTorque Pi 同尺度机器人 … 用 0.08, 而它 target_feet_height = 0.02 m —— 说明抬脚是策略在参考之上加出来的, 参考只负责钉住髋/膝/踝的相位配合。我们取 0.30 等于让参考本身就是整个步态, 过约束 + 过驱动”
    train/WALK_V6_MINIMAL.md § ① target_joint_pos_scale 0.30 → 0.15
  • When training fails repeatedly, inject the target behavior open-loop - stop tuning rewards for an unverified behavioropen-loop-probe-before-reward-tuning
    Mechanism understoodomniattributionattributionprocesscurriculum

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop on the real plant/sim first, and only resume training once you hold a measured, safe, sign-verified target trajectory.

    Symptom

    Three sidewalk training rounds failed; hypotheses multiplied (exploration failure / wrong reference waveform / insufficient authority) with no way to pick between them by running more training.

    Context

    Instead of a fourth reward guess, the team wrote probe_side_ref.py: the candidate reference is injected open-loop on top of a frozen policy's output (bypassing PPO entirely), directly measuring "what happens if the robot literally does this waveform" - separating all three hypotheses in one experiment (5 seeds x 8 s per condition, several waveform families and gains). The probe immediately eliminated the authority hypothesis (full-amplitude execution, 5/5 survival) and localized the problem to the waveform/measurement side. The closing principle was written down after the saga: without a verified target behavior, tuning rewards is "在黑暗里试钥匙" (trying keys in the dark).

    Change

    Standing method: before opening another training rung on a failing skill, build an open-loop (or task-space) generator of the intended behavior, measure whether the physical system can express it and what it looks like - then train toward a verified, quantified target.

    Outcome

    The probe chain produced the verified waveform (reversed-sign triangle, half gain), quantified safe amplitude (tilt 8.2 deg at band top, foot distance clear of the wall), exposed the metric bug when probe and training disagreed, and ultimately supplied the feed-forward that made C4 pass in +100 iters.

    Mechanism

    Training couples exploration, reward design, and feasibility into one opaque outcome; open-loop injection cuts the loop and tests feasibility and waveform alone. A behavior demonstrated open-loop converts the remaining failure into a pure credit-assignment/reward question - and its measured trajectory becomes the reference itself.

    Applies when

    • repeated training failures on one skill with multiple live hypotheses
    • uncertainty whether the platform can physically express the behavior
    • a reference trajectory's shape/sign/amplitude is guessed, not measured
    “三轮 FAIL 之后不再猜,写 train/probe_side_ref.py 把参考开环注入到策略输出之上(绕过 PPO),直接量「照这个波形做会怎样」,一次分开三个假说:甲 探索 / 乙 波形 / 丙 权限。”
    train/C_LADDER_RUN.md § 3f. C4 真因定谳(开环探针) / 3k. 建议的下一步
  • A suspended (no-load) test acquits or convicts the actuator before you blame authoritysuspended-test-isolates-actuator-authority
    Mechanism understoodomnireal-acceptancehardwarereal-acceptanceattribution

    Before attributing a failure to actuator authority, measure no-load tracking error and steady-state torque fraction; blame authority only if the task fails while the error grows with demanded force - and then fix gains or targets, not training.

    Symptom

    hip_roll sagged 0.21 rad on the ground and saturation questions loomed over the sidewalk plan - was the roll axis physically too weak (authority), or was something else limiting it?

    Context

    Before C4, the roll-authority question was settled by measurement triage: suspended test (--suspend, feet off ground) showed hip_roll tracking error 0.0008 rad - actuator acquitted; the entire 0.21 rad ground sag is load-induced. Steady-state torque was 25% of limit - 75% margin remains. Since sidewalk needs lateral force, not exact angles, authority was ruled "not a hard limit", with a pre-registered criterion for when it WOULD become one: sidewalk fails to track AND roll error keeps growing - then the fix is raising hip_roll kp or lowering the vy target, not more training.

    Change

    Hypothesis "roll authority insufficient" demoted from blocker to a monitored branch with an explicit trigger condition; C4 proceeded.

    Outcome

    Later open-loop probes confirmed the actuator could produce the behavior (sidewalk feed-forward ran at full amplitude, 5/5 survival), and the eventual C4 failure causes were measurement and reward, never authority.

    Mechanism

    Suspended vs loaded comparison separates the actuator's closed-loop competence from the load path: tiny no-load tracking error means the motor/controller is fine and any loaded deviation is statics (gravity / stiffness budget, kp trading error for force). Torque-fraction measurement then bounds how much force headroom actually remains.

    Applies when

    • suspecting an axis is "too weak" for a new skill
    • large position sag on a loaded joint
    • deciding between hardware fix, gain change, and more training
    “吊挂(--suspend)实测 hip_roll 跟踪误差 0.0008 rad → 执行器无罪,地面下垂 0.21 rad 全是负载所致;稳态占限扭 25% → 仍有 75% 扭矩余量。… 判据:若 C4 出现「侧走跟不动且 roll 误差继续变大」,那才是权限账 … 解法是提 hip_roll 的 kp 或降 vy 目标,不是硬训。”
    train/C_LADDER_RUN.md § 3d. roll 权限:已部分澄清,不是硬上限
  • IMU observation age cut 52-68 ms to ~4 ms by moving AHRS onto the MCU - as a single variableimu-age-move-fusion-downstream
    Observed oncewalkreal-deployhardwarereal-acceptanceprocess

    Audit observation age end-to-end and move time-critical fusion as close to the sensor as possible - and when you fix a latency, change only that one variable so the gain is attributable.

    Symptom

    IMU-derived observations reaching the policy were 52-68 ms old because attitude fusion ran in Python on the loaded host computer - stale attitude is a direct feedback-loop delay the policy was not trained with.

    Context

    The fix was scoped deliberately narrowly: move the AHRS computation from Python to the STM32 H7 (MC02). CAN topology explicitly unchanged, so the change is a clean single variable.

    Change

    AHRS fusion relocated Python -> H7. Before/after - IMU age: 52-68 ms -> ~4 ms; CAN timing: unchanged; Python load: high -> ~0.

    Outcome

    IMU age reduced by an order of magnitude with no confound; host CPU headroom recovered ("把计算单元搬在stm32上, 这样imu有剩余").

    Mechanism

    Sensor age is pipeline latency, not sensor quality: fusing on the MCU next to the sensor removes host scheduling jitter and interpreter overhead from the critical path. Keeping the bus topology fixed makes the improvement attributable to the relocation alone.

    Applies when

    • measured sensor-to-policy age far exceeds sensor sample period
    • attitude fusion or filtering runs on a loaded host CPU in an interpreted runtime
    • planning infrastructure changes during a sim2real campaign
    “AHRS 搬到 H7——这个不改 CAN 拓扑,只是把一段计算从 Python 挪到 MC02,单变量:IMU age 52–68 ms → ~4 ms / CAN 时序 不变 / Python 负载 高 → ≈0”
    Experience.md § AHRS 搬到 H7 (lines 28-35)
  • Cutting swing amplitude 40% raised yaw-momentum demand 53% - the falsified fix is recorded so nobody walks that road againamplitude-cut-falsified-yaw-fix
    Mechanism understoodwalksim-evalmeasurementattributionreward-shaping

    Test gait fixes against the quantity the ground must actually supply (torque/force rates vs friction ceilings), not against kinematic proxies; record falsified fixes with their mechanism so the search space shrinks permanently.

    Symptom

    Support-foot yaw slip stayed at ~223 deg (vs 225 deg) after walk_v6 cut joint swing amplitudes by ~40% (hip_pitch 20.6->10.4 deg, knee 31.3->18.8 deg) - the change built on the theory "smaller swing = less yaw momentum to dump into the ground".

    Context

    Direct measurement inverted the theory: yaw-momentum amplitude ROSE 16% (+/-0.1000 -> +/-0.1159) and its rate of change rose 53% (1.86 -> 2.84 N*m demanded from the ground), pinned exactly at the foot's supply ceiling (2.8-3.3 N*m at mu 0.6-0.7) - so slip could not drop. The extra demand lives in higher harmonics: v6's crisper foot placement (better clearance 34 mm, lower landing force) shortens the momentum- exchange window, concentrating the same exchange into less time. The verdict was written as a closed road: "下一轮不要再走这个方向". An honest residue was also booked: WHY amplitude down but momentum up 16% remained unresolved, with the named next step (per-rigid-body decomposition of H_z, since per-joint RMS sensitivity ignores phase correlations).

    Change

    The "reduce amplitude to reduce yaw momentum" lever was removed from the planning space; future yaw-slip work redirected toward the supply side (friction) and momentum-rate mechanics.

    Outcome

    Slip unchanged (225 -> 223 deg); the falsification and its mechanism became a permanent constraint on the fix search space.

    Mechanism

    Ground yaw torque demand scales with the rate of change of angular momentum, not its amplitude; kinematic amplitude cuts that also sharpen contact timing can raise dH/dt while lowering range. When demand exceeds the friction-limited supply ceiling, slip is set by the ceiling, so demand-side changes below the ceiling do nothing visible.

    Applies when

    • attacking foot slip or yaw drift via gait shape changes
    • a fix targets an amplitude while the constraint is a rate
    • documenting a failed intervention after a version comparison
    “walk_v6 | ±0.1159 (+16%) | 2.50 Hz | 2.84 N·m (+53%) … 脚的供给上限 2.8~3.3 N·m(μ 0.6~0.7)—— v6 正好顶在天花板上, 所以滑移一点没降。… 结论: "减小摆动幅度以降低偏航动量"这条被 v6 证伪 —— 砍 39% 幅度, 需求反升 53%。下一轮不要再走这个方向。”
    train/WALK_DIAGNOSIS.md § ② 未生效的机理: 减小摆动幅度反而让偏航需求上升

Next page