Training Coach
Doctrine
A report may cite any of these as doctrine-N.
doctrine-1Contract freeze and fingerprint disciplineThe policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.
Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).
Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.
doctrine-2Attribution by resolved training params - never eval-override knobsCapability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.
Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).
Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.
doctrine-3PASS gates become constraints; FAIL gates become objectivesOnce a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.
Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.
Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".
doctrine-4One variable per ladder rung - counted against what the checkpoint sawA rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.
Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.
Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.
doctrine-5Pre-register risks, readings, and stop criteria before the ladderBefore a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.
Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).
Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.
doctrine-6Plant parameters are measured, never inventedEvery plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.
Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).
Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.
doctrine-7Sim2sim gate before sim2real - under deployment conditionsEvery checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).
Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).
Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.
doctrine-8Observation honesty - the actor's inputs are a hardware contractThe actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.
Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).
Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.
doctrine-9Reward economics are audited in realized currencyReward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.
Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).
Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.
doctrine-10The zero-cost option must be the desired behaviorFor every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.
Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).
Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.
doctrine-11Measurement discipline: independent referees, signs, distributionsA disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.
Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).
Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.
doctrine-12The deployment pipeline is plantIrreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.
Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).
Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.
doctrine-13DR budget is finite; its distribution is the measured supportRobustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.
Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).
Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.
doctrine-14Gates measure what hardware feels: posture, margins, stripped assistsAcceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.
Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).
Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.
doctrine-15Fork and root selection: recoverability, maturity, frozen rewardsChoose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.
Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).
Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.
doctrine-16Curricula: verified engagement, lineage counters, disease-phase gatingAutomatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.
Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).
Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.
doctrine-17Probe before training: feasibility first, hypotheses in tablesAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.
Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).
Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).
doctrine-18External advice is recomputed locally; values transfer as ratiosEvery external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.
Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).
Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.
doctrine-19Hardware sessions are scripted experiments, not tuning sessionsReal-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.
Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).
Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.
doctrine-20Close questions in writing; restart when the debt is structuralAudited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.
Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).
Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.
doctrine-21Name the quantity in the space it lives inA goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.
Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).
Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.
doctrine-22Continuation needs a live gradient; a release is chosen by a scanContinue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.
Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).
Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.
Experience cards
162 cards matching “tail-torque-needs-hinge-on-computed-demand”.
Late training leaned on sampling noise as a stability crutch - deterministic play collapsed while training metrics stayed green
noise-crutch-deterministic-collapseEvaluate the deterministic policy in an external harness on a fixed cadence during training (not just at the end), select checkpoints on that curve, and treat a collapsing noise_std with rising training reward as a warning that noise is load-bearing.
Symptom
omni_s1's final checkpoint (model_5999) fell at 4 s even in Isaac's OWN deterministic play, while checkpoints from iter 1700-4000 were fine - and no training metric flagged anything. Policy noise_std had collapsed to 0.045 by iter ~990 (final 0.033).
Context
Diagnosis: the policy had learned to use its exploration noise as a dither/stabilizer - "策略把采样噪声当稳定拐杆,训练指标看不见" (the training metrics cannot see it, because training always runs with noise on). Countermeasures: entropy_coef 0.005 -> 0.01 to slow the std collapse, and - the structural fix - an in-training smoke loop (watch_ckpt.py): every 500 iters, export ONNX directly, run 3-seed MuJoCo evaluation, log CSV/TensorBoard curves plus three-view videos. The doctrine line was written in bold: "训练指标全绿不再是发育健康的 证据,冒烟曲线才是" - green training metrics are no longer evidence of healthy development; the smoke curve is. The follow-up run s1b showed the drift metric follow a U-shape (73 -> 8.6 at iter 3500 -> 76), making checkpoint selection BY the smoke curve (early stop at 3500) the shipping mechanism, with terminal re-degradation booked as known and unresolved.
Change
entropy floor raised; watch_ckpt smoke loop instituted as standing infrastructure; checkpoint selection moved from "last iteration" to "best point on the deterministic smoke curve".
Outcome
s1b shipped from iter 3500 (the U-bottom) instead of a degraded terminus; every later lineage (s1c/s1e, the C ladder's --every 100 loops) inherited the watcher as the standard guardrail.
Mechanism
PPO evaluates and improves the stochastic policy; if noise itself stabilizes the gait (dither smoothing a marginal limit cycle), the deterministic mean policy is a different, worse controller that training never measures. External deterministic evaluation on an independent simulator is the only readout of what will actually be deployed.
Applies when
- final checkpoints underperform mid-training ones
- noise_std collapses early while training reward climbs
- deciding which checkpoint to export and ship
“训练后期确定性脆化——noise_std iter~990 收到 0.045(终 0.033),model_5999 连 Isaac 确定性 play 都 4 s 摔(1700~4000 正常):策略把采样噪声当稳定拐杖,训练指标看不见。对策:entropy_coef 0.005→0.01 + train/watch_ckpt.py 训练中冒烟曲线 … 训练指标全绿不再是发育健康的证据,冒烟曲线才是。”
train/OMNI_V0_SPEC.md § 3. S1.1 修订记录 ② Fix the task first, harden the plant second - DR budget spent on a dying task is wasted
task-shaping-before-plant-hardeningFreeze the task/command distribution before spending DR budget on plant robustness; if the task will still change, schedule plant hardening as a final pass and book the interim robustness gap explicitly.
Symptom
Tempting default ordering was to keep the plant-hardened (S2) lineage and teach it new commands; but the S2 plant adaptation had been earned on the straight-walk task, and the new omni tasks (sidewalk, in-place turn) use completely different contact patterns.
Context
The team had direct evidence that DR robustness is a budget that gets reallocated when the data distribution changes ("push/μ 两轮已实证 DR 预算有限且会被重分配") - robustness trained under one task/command distribution does not persist when training continues under another.
Change
Ladder order set to: first C (task shaping - add command modes until the task family is final), then a second S2 pass (plant hardening) on the C product. The plant-robustness gap this creates mid-ladder is accepted and booked explicitly ("此处不欠账" - the debt is assigned to the second S2 pass, not denied).
Outcome
The first S2 pass was not wasted: its laws (kd bandwidth <-> low mu, push need not be trained, ground mu need not be trained, bistability) let the second pass drop from five rungs to three. The C ladder itself ran on the softer plant band without incident.
Mechanism
DR robustness is carried by the policy's visited-state distribution; changing the task changes that distribution, so robustness bought under the old task partially dissolves. Hardening before the task is final means paying for robustness on states that will no longer be visited - "给一个即将不存在的任务花预算" (spending budget on a soon-to-not-exist task).
Applies when
- deciding ordering between skill/command expansion and DR hardening
- a hardened lineage is proposed as the root for a task change
- robustness regressions appear after adding new command modes
“S2 的 plant 适应是为直行步态调的,C4 侧走/C3 原地转是完全不同的接触模式,先硬化再改任务 = 给一个即将不存在的任务花预算(push/μ 两轮已实证 DR 预算有限且会被重分配)。故顺序改为 先 C(任务定型)→ 再 S2(plant 硬化)。”
train/C_LADDER_RUN.md § 0. 决策逻辑 = 短板可不可恢复 (末段) A 10 s acceptance episode left 6-7 s of standing to observe - a narrow stance held for that window and split on hardware; a gate cannot see instability slower than its own horizon
episode-length-bounds-what-a-gate-seesSize the standing phase of an acceptance episode, and its disturbances and floor friction, to what deployment will impose; a pass on a short static window certifies only that window.
Symptom
v2_6 passed every simulation gate (success 99.6%, re-falls 0-1%) and then, on the real robot, stood up and slid into the splits several times; prone starts stood and then fell backwards.
Context
Acceptance ran 10 s episodes; a ~1-2 s get-up left roughly 6-7 s of static standing on a nominal floor with no disturbance. The narrow stance's lateral margin and the straight-knee stance's lack of any flex buffer are both failure modes that need time, disturbance or lower friction to show.
Change
The gap was booked as a known blind spot of the gate ("long-duration standing stability") alongside the task-space stance criterion; later rungs added MuJoCo friction sweeps at mu 0.4 to every checkpoint scan.
Outcome
The spec through §50 records the blind spot but no longer standing window or disturbance row in the recovery acceptance itself.
Mechanism
An acceptance episode observes only the dynamics that unfold within its horizon under its conditions; slow drifts and disturbance-triggered failures are outside it by construction.
Applies when
- a policy passes sim gates and fails on hardware after a delay
- acceptance episodes are short relative to deployment use
- stability is judged without pushes or friction variation
“**sim 门为什么没逮住**:10 s episode 起身后只站 ~6-7 s,静态窗口内窄站距 撑得住;真机站立时长/扰动谱在门口径之外 —— 长时站立稳定性记为口径缺口。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §43 判读:sim 门为什么没逮住 Train with self-collisions ON (filtering nested-link ghost pairs) - the reward wall prevents, the physics makes cheating impossible
self-collision-physics-plus-reward-wallNever train a contact-risk behavior with self-collisions disabled; enable them with an audited filter list for nested/overlapping pairs (zero contacts across a pose sweep), record the fps cost, and keep a calibrated distance penalty as the preventive layer on top.
Symptom
walk_v8 logged 107 frames of leg-on-leg contact while still earning 0.751 tracking score - because training-side self-collisions were OFF, leg clipping was literally imperceptible to the policy ("碰腿在训练里 根本感知不到").
Context
Enabling self-collisions naively is its own trap: an Isaac audit had shown PhysX auto-filters adjacent bodies (base-hip clean for free) but nested links generate ghost forces - calf and ankle_roll overlap 65 mm at the zero pose, producing 12x body-weight phantom forces. The v10 recipe: enable self-collisions, explicitly filter only the two nested pairs (l/r calf-ankle_roll), then run a zero-contact audit at three poses (nominal stand, walk crouch, swing-extreme) requiring contact count = 0, adding any residual pair to the filter and re-auditing; a 500-iter sanity run for NaN and an fps-cost record (measured -8.8%). Redundancy with the reward-side foot-distance wall was argued, not assumed: "N2 离得远(奖励侧预防),SC 碰了疼(物理侧兜底)" - the reward keeps distance at range, the physics makes contact hurt - so the v8-style "clip legs and still score" outcome becomes physically impossible.
Change
enabled_self_collisions=True + 2-pair filter + three-pose zero-contact audit (re-verified at 0.00 N after the later mass update) + fps budget recorded.
Outcome
Leg contact entered the training signal; the audit protocol caught the nested-pair ghost-force hazard before it corrupted training; combined with the calibrated distance wall, later versions held contact = 0 on hardware and in sim.
Mechanism
A hazard absent from the training physics cannot be learned about, no matter the reward; but collision meshes that interpenetrate at rest inject large fictitious forces if enabled blindly. Filtered enabling plus a pose-swept zero-contact audit gives true contact physics with no phantom energy - and layering prevention (reward) with consequence (physics) covers both learning and enforcement.
Applies when
- real robot self-contacts while training scored it healthy
- enabling self-collisions on a model with nested collision meshes
- deciding between reward-side and physics-side fixes for clipping
“PhysX 自动过滤相邻体(base↔hip_pitch 免费干净),幽灵力只在 calf↔ankle_roll(零位嵌套 65mm,12 倍体重)。… 与 N2 互补不冗余:N2 离得远(奖励侧预防),SC 碰了疼(物理侧兜底)—— v8 那种 107 帧互碰拿 0.751 跟踪分的事从此物理上不可能。”
train/WALK_V10_SPEC.md § 4. SC —— 训练侧自碰撞(范围已探明,比想象便宜) A quadratic clearance reward stalled for 3500 iters near target - switch to an indicator on accumulated height
indicator-reward-avoids-gradient-decayWhen a shaped term plateaus near its target, check the gradient profile: replace vanishing-gradient forms with threshold/indicator forms for the final approach, and prefer delta-accumulation over absolute positions to immunize against frame offsets.
Symptom
The quadratic-error clearance term froze at -0.002 from iteration 2000 to 5500 - thousands of iterations with no progress on foot lift.
Context
Diagnosis: a quadratic penalty's gradient vanishes as the error approaches target, so exactly where the last millimeters must be earned the incentive fades to nothing. The replacement (Humanoid-Gym form): accumulate the swing-phase height climb per foot, reward a BINARY indicator |accumulated - target| < 0.01 masked to the planned swing window, reset on contact, weight +1.6 as a positive reward. Two properties: the indicator's incentive is constant until the threshold is crossed (no decay zone), and accumulating height DELTAS makes any constant sole-frame offset cancel automatically - which structurally sidesteps the earlier 0.0585 m zero-point bug ("顺带绕开我先前那个'忘了减 0.0585 导致惩罚恒为 0'的坑").
Change
Clearance reformulated from quadratic penalty on instantaneous height to indicator on per-swing accumulated climb (target 0.03 m by leg-length scaling, weight +1.6).
Outcome
Part of the v5 package under which lift finally moved (v5 29 mm, v6 34 mm vs the stalled 18-24 mm era); the offset-cancellation property removed one whole bug class from the term.
Mechanism
Policy-gradient learning follows the reward's local slope; quadratic shaping concentrates slope far from target and starves it near target, so convergence stalls precisely at the finish line. An indicator pays a constant bounty until the goal is met; formulating on deltas rather than absolutes removes sensitivity to reference- frame constants.
Applies when
- a reward term's value freezes short of target for thousands of iters
- designing clearance/height/precision terms
- reward code depends on absolute link positions
“现行二次型在接近 target 时梯度趋零 —— 这正是 clearance 从 iter 2000 到 5500 卡在 −0.002 不动的原因。… 二值指示在跨过阈值前梯度恒定,没有衰减区 … 累积 delta 让 SOLE_OFFSET 自动抵消”
train/WALK_V5_SPEC.md § 3. clearance 改峰值型(去掉二次型的梯度衰减) Real robot walked at half the sim clock for two generations - resolved by racing a reward-side and a plant-side evidence line, not by guessing
period-doubling-evidence-raceFor a hardware-only pathology, refuse to guess: pre-register one probe per side of the sim2real boundary (can the reward mechanism change it on hardware? can fitted plant parameters reproduce it in sim?) and let the first positive result direct the next version.
Symptom
The number-one sim2real gap: on hardware v6/v7 stepped at 1.23-1.32 Hz - almost exactly half the 2.50 Hz gait clock they were trained and simulated at; sim never reproduced it, two generations running.
Context
Instead of committing training budget to a guess, v8 pre-registered two mutually controlled evidence lines and kept the clock OUT of the training variables: (a) reward-side - if the v8 saturation fix revives joint_pos_ref (the term that pins the gait to the clock), re-run hardware and see whether frequency returns to 2.5 Hz (hypothesis: v7's frozen actions meant NO reward was pinning the gait to the clock, and the real plant - with armature and friction making high frequencies expensive - slid down to the leg's pendulum natural frequency ~1.1 Hz); (b) plant-side - record suspended joint data (fit_actuator), fit armature/friction, load the fitted values into sim2sim and see whether the 1.25 Hz reproduces IN SIM. Decision rule fixed in advance: "谁先给出阳性结果谁定 v9 的方向 (奖励侧 vs plant 侧)" - whichever line goes positive first sets the next version's direction.
Change
Period-doubling excluded from the v8 change set; both diagnostic lines scheduled in parallel as non-blocking work; frequency reported factually in acceptance with no pass/fail attached ("倍周期是否消失 不设判定,它是 §9 的关键证据").
Outcome
The gap was routed into a decisive-experiment structure rather than a speculative retrain; the plant-side line pointed at exactly the unmodeled armature/friction that were later measured and installed as the plant baseline. Resolution (era-2c full-plant retest): the family had TWO causes - v8's low-speed period-doubling vanished once measured armature+friction were installed (1.30 -> 2.50 Hz, bifurcation-edge machine sensitivity), while v7's stood untouched at 1.20 Hz (saturation-freeze-driven policy property) - both evidence lines paid off, one per case.
Mechanism
A behavior appearing only on hardware has candidate causes on both sides of the sim2real boundary; changing training to fix it tests only one side per expensive cycle. Two cheap parallel probes - one intervening on the reward mechanism, one making sim reproduce the real behavior - localize the cause to a side before any training money is spent, and sim-reproduction of a real pathology is itself the strongest form of plant validation.
Applies when
- a gait pathology appears on hardware but never in any simulator
- deciding whether a sim2real gap is reward-side or plant-side
- tempted to change the gait clock/reward to chase a hardware symptom
“倍周期(真机 1.23~1.32 Hz ≈ 时钟一半,v6/v7 连续两代;sim 从不出现):两条证据线互为对照——(a)… 真机重跑看频率是否回 2.5 Hz(假说:v7 没有任何奖励把步态钉在时钟上,真机 plant 有 armature/摩擦、高频贵,自由滑落到复摆自然频率 ~1.1 Hz);(b)真机吊挂录 fit_actuator.py … 看能否在仿真里复现 1.25 Hz。谁先给出阳性结果谁定 v9 的方向。”
train/WALK_V8_SPEC.md § 9. 平行线 (倍周期) Order hardware runs by sim risk, gate each stage on the last, and put the fragile cell last with a spotter
risk-ordered-real-deploymentScript hardware sessions as a risk ladder: baseline first, sim-riskiest last with a spotter, suspended smoke before ground, each stage gated on the previous, environment (floor mu) recorded as a selection input - and stop at the stage that misbehaves.
Symptom
Five policy-x-gain combinations had to go on hardware in one session, with sim survival ranging from 20/20 down to 17/20 (and zero-command survival down to 2/20) - an unordered session risks breaking the robot on an avoidable run.
Context
The execution sheet fixed the order as sim-risk low to high, control baseline first (current SOTA establishes the floor reference), the fragile cell (fric-2400@kd1.0) last with a person spotting throughout. Stage gating: suspended smoke (feet off ground, 10 s each, all five pass before anything touches down) -> suspended with IMU and forward command (gait forms in the air) -> grounded runs -> speed raise only for combos that survived the previous stage -> zero-command tests only with a spotter, ordered by sim zero-cmd survival, with the 2/20 cell skipped by default. Preconditions include recording the floor material and estimating mu (if mu <~0.6, sim says pick the kd1.2 gain as main), port/CAN self-check, calibration frozen. Any stage failing stops the session at that stage: "任一段出问题就停在那一段, 不要跳到下一段".
Change
Session structured as a risk ladder with per-stage gates instead of a flat checklist; per-combo sim survival numbers written into the run table as the ordering key.
Outcome
The session design localized any failure to the cheapest stage that could reveal it, kept the robot safe for the informative fragile run, and made the control baseline available before any comparison run.
Mechanism
Hardware sessions consume a shared budget (robot integrity, battery, floor time); ordering by predicted risk means information is bought cheapest-first, and stage gates convert an expensive failure into a cheap earlier one. Baselines run first because every later reading is relative to them.
Applies when
- taking multiple policies/configs to hardware in one session
- a candidate is known-fragile in sim but must be measured
- writing a deployment runbook for a new robot
“跑序 = sim 风险从低到高, 最险的放最后 (依据 = 存活门/零指令存活) … ⑤ 是 sim 里最脆的一格 … 放最后跑, 全程留人扶, 起步即给 cmd, 零指令不做。… 任一段出问题就停在那一段, 不要跳到下一段。”
train/REAL_RUN_S2.md § 上机名单 / 全部命令 Privileged signals (true velocity, foot force, foot height) go to the critic only
observation-honesty-critic-onlyTreat the actor observation vector as a hardware contract: every element must exist on the real robot with realistic noise; privileged simulator truths belong in the critic only.
Symptom
Policies trained on ground-truth base linear velocity work in sim and fail on hardware, where only a drifting IMU and encoders exist - the policy has learned to depend on a signal that does not survive deployment.
Context
Many open-source locomotion stacks feed simulator ground-truth linear velocity to the actor. The reference team refused: the real robot has no ground-truth velocity. Asymmetric actor-critic keeps the training benefit of privileged information without deploying the dependency.
Change
Route ground-truth velocity, foot contact forces, and foot heights to the critic only; the actor observes exclusively signals that exist on hardware (IMU-derived quantities, encoders, commands, previous actions).
Outcome
Recorded as adopted doctrine in Lucen's experience log; the trained actor's input contract matches what the real robot can actually produce.
Mechanism
The critic is discarded at deployment, so it may consume any privileged state to reduce value-estimation variance; the actor's observation set is a deployment contract - anything in it that hardware cannot supply (or supplies with different noise/drift) becomes a train/deploy distribution shift the policy was never trained to handle.
Applies when
- designing actor/critic observation spaces
- reviewing a config where the actor sees base_lin_vel or contact forces
- sim policy is strong but real robot drifts, oscillates, or falls without obvious actuator cause
“很多开源代码库把真值线速度喂给策略,Asimov 团队没有,因为真机上没有真值速度,只有会漂的 IMU 和编码器;用完美速度训练出来的策略会依赖它,然后在硬件上失效。真值速度、足底力、足高统统只给 critic”
Experience.md § 观测空间的诚实性 (line 7) The fallen-state reset was designed, not sampled from SO(3) - fixed category shares with jitter, a low drop that settles physically, equal left/right shares for mirror augmentation, and a numeric check before training
fallen-pose-reset-distributionBuild a fallen-start distribution from named, physically plausible categories with jitter and a settle phase, keep mirrored categories at equal probability, and check the realized shares and penetration numerically before spending a training run on it.
Symptom
A get-up policy can only learn from the fallen states its resets produce; uniformly random orientations produce ground-penetrating and limit-jammed states the robot can never be in.
Context
R0 reset_root_fallen: supine 30%, prone 30%, side_l 15%, side_r 15%, mid (random axis 50-125 deg) 10%, +/-15 deg jitter, full yaw, dropped from 0.28-0.40 m and left to settle under physics, joints uniform inside the soft limits with a 5% margin plus small random velocities. Random SO(3) was rejected (the advisor agreed). side_l and side_r must have equal probability because mirror augmentation turns a left fall into a right fall. The advisor had proposed supine and prone only for R0; the spec included side and mid because the feasibility accounts showed physical solutions for all of them, and wrote "narrow back to supine+prone" down as the first fallback. With no display on the training box the reset was checked numerically instead of by eye.
Change
Category mix as above; realized shares, settle height and penetration measured over 512 envs before the first run. A fallen-state bank (real falls, settled and stored) was pre-registered for R2.
Outcome
Realized shares 29.3/31.6/16.4/17.8% against the config, settle +0.262 m, final penetration 0/512 (a 0.10 m peak at the write instant, ankle links only, pushed out within 80 ms because the 0.28 m drop floor is shorter than a fully extended leg). R0's failure was a reward basin, not a reset artifact. The fallen-state bank stayed unbuilt through V3.1 (checklist item open); R0.3 later re-sliced the prone share into roll_l/roll_r bands, which is what forced the acceptance distribution to be frozen separately.
Mechanism
A category-structured, physically settled start distribution keeps training on states the robot can actually occupy, and equal mirrored shares keep mirror augmentation a pure doubling of data rather than a bias.
Applies when
- designing reset distributions for get-up, recovery or multi-contact skills
- mirror/symmetry augmentation is on and the task has chiral start states
- no viewport is available to inspect resets on the training machine
“角度 jitter ±15°、yaw 全域、0.28~0.40 m 低空放下由物理沉降,关节软限位内 均匀(留 5% 余量)+ 小随机速度。**不用 random SO(3)**(会采出穿地/极限卡死 等现实不可能状态,顾问同判) … side_l/side_r **概率必须相等**(镜像增强的样本同分布前提)”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §3 R0 任务定义 / §9 核查单 The trainer read a stale USD after the URDF mass update - regenerate derived assets and gate on an automated equality instrument
derived-asset-staleness-checkFor every derived plant artifact (USD from URDF, generated value files), pair the generation step with an automated source-vs-derived equality instrument, prove the instrument can fail, gate training on its PASS, and re-run physical audits after every regeneration.
Symptom
Measured link masses had been committed to URDF/MJCF (total 9.58 -> 9.792 kg, weighed values), but Isaac reads the derived USD asset - which still carried the old masses: a silent 2.2% mass fork between the training plant and the evaluation plant.
Context
The v12 checklist made USD regeneration a hard precondition ("硬性 前置") and, crucially, backed it with an instrument: check_usd_mass.py compares USD vs URDF per-link mass AND inertia trace, validated by showing it FAILs the stale asset naming 9 offending links, then PASSes after re-conversion (13/13 links consistent, total 9.7920). The self-collision filter audit was re-run after the regeneration (three poses, 0.00 N) because a regenerated asset invalidates physical audits done on the old one. This milestone was also where the three plants first aligned: "三边 plant 首次对齐(armature+摩擦+ 实称质量)就在这一代".
Change
convert_urdf re-run on the training machine, regenerated asset committed, check_usd_mass.py PASS required as an acceptance gate for the generation; dependent audits repeated post-regeneration.
Outcome
The 2.2% plant fork was closed before it could distort a generation's acceptance numbers; the staleness class of bug now has a permanent detector instead of a memory.
Mechanism
Source-of-truth edits do not propagate to derived binary assets by themselves; any consumer reading the derivative silently trains or evaluates on the old plant. An automated equality check between source and derivative - proven able to fail - turns an invisible staleness into a red gate, and regeneration invalidates every audit performed on the old artifact.
Applies when
- editing masses/inertia/geometry in URDF or MJCF sources
- a trainer or evaluator consumes converted/derived assets
- plant numbers differ between simulators for no visible reason
“25ba997 把连杆质量更新为实称值(总重 9.58→9.792 kg,URDF/MJCF 已改),但 Isaac 读的是 train/assets/laika_v2.usd —— 仍是旧质量。… 否则 Isaac(9.58)与 MuJoCo(9.79)质量分叉 2.2%,v12 验收数字失真。验收门:python tools/check_usd_mass.py 必须 PASS … 对旧资产实测 FAIL/9 连杆点名,仪器已验证”
train/WALK_V12_SPEC.md § 7. 核查单 (⚠️ 先重转 USD) Sort external training advice into adopt / already-have / modify / would-trap by recomputing it on your own config
external-advice-audit-against-own-arithmeticNever apply external tuning advice directly: recompute each claim on your own reward table and probe data, classify it adopt / have / modify / trap, and record why - and verify external citations actually exist.
Symptom
External AI/literature advice for the omni ladder arrived plausible-sounding but was written without knowledge of this robot's actual reward table, contract, and history; following it blindly would have broken single-variable discipline and, in one case, made sidewalk unlearnable.
Context
Before the C ladder, every external suggestion was audited: 3 adopted (ellipsoid command sampling; staged wz bands; command-switch acceptance), 3 already present (unified reward; frame history - the frozen 168-dim 5-frame window; per-100-iter acceptance), 2 modified (stand share kept at 20% to avoid a second variable; back share NOT raised because the probe showed backward works untrained 20/20@67%, so oversampling would only crowd out forward), and 1 flagged as a trap: "start vy very small (0.06-0.15)" - on THIS reward table vy was only an L2 tax, so ignoring a vy=0.06 command costs 0.4% of the vx tracking scale, 28-180x cheaper than ignoring forward, with quadratic shrinkage making small commands weaker still. A separate retrieval-reliability note: two search agents returned fabricated verbatim quotes from arXiv PDFs (2 papers, verified fake and discarded); only HTML/abstract/source-verifiable material was used.
Change
Advice classified only after recomputing each claim with local numbers; the "start small" trap was replaced by adding a gated lateral tracking term (the ladder's only true reward surgery) instead of shrinking the command.
Outcome
The adopted items (ellipsoid modes, staged wz, transition acceptance) entered the ladder; the trap was avoided; one external factual error (calling s1g the mainline start - it was falsified 0/20) was caught. Later, one initially-dismissed item (sigma=0.15 too narrow) turned out right for a different reason than claimed - see cycle-average-tracking-for-gait-quantities.
Mechanism
External advice encodes the advisor's reward table and robot, not yours; the transfer-validity test is whether the claim survives recomputation under your own arithmetic (reward margins, probe baselines, contract freeze). Items that survive become experiments; items that don't become documented traps.
Applies when
- incorporating LLM or literature advice into a training plan
- advice conflicts with locally measured baselines
- an external claim depends on reward-table details the advisor cannot know
“其建议 C4「先很小,vy = ±0.06~0.15」—— 在我们这张奖励表下会让侧走学不起来 … 忽略侧走比忽略前进便宜 28~180 倍,且指令越小激励越弱(平方缩放)—— "先很小"在稳定性上对、在梯度上正好把信号缩没了。”
train/C_LADDER_RUN.md § 1. 外部 AI 训练建议的评估(采纳 / 已有 / 要改 / 会踩坑) After installing the measured plant, re-run all generations paired on old and new plant - identity metrics carry verdicts, physics metrics re-baseline
plant-swap-invariants-vs-shiftsTreat every plant upgrade as an era boundary: re-run the retained policy set paired (same seeds/flags) on both plants, carry forward only verdicts whose metrics proved plant-invariant, re-baseline the rest - and mine the systematic shifts as measurements of the old plant's biases.
Symptom
With the plant finally fully measured (weighed masses 9.792 kg, bench-identified armature, in-situ friction), no historical sim number was comparable to new runs - "历史 sim 数字跨纪元不可比" - and it was unknown which historical verdicts still held.
Context
The era-2c cross-test ran all seven walk generations on the complete plant under one harness (14/14 survived), then re-ran the same 14 configurations on the OLD plant retrieved from git, same flags and seeds, with a self-check (one historical record reproduced digit-for-digit). The split was clean. Policy-identity metrics moved essentially zero across the plant swap - dominant frequency (v7's period-doubling 1.20 -> 1.20), knee amplitude (9.9 -> 10.0), foot distance (+/-2 mm), saturation (100 -> 100, 0 -> 0) - so all seven cross-generation verdicts (freeze signature, saturation-line closure, knee-collapse location, slip-penalty accounting, N2 lineage, v6's balance, v11's triad) were re-confirmed on the honest plant. Plant-physics metrics shifted systematically with ordering preserved: slip down 10-25% (measured friction makes ground-twisting costlier), landing vertical velocity down 15-50% ("旧 plant 高估落地 凶度"第二次独立证实), landing force mixed (mass up 2.2% vs friction braking the swing - two effects fighting). Exactly ONE behavior-level change: v8's low-speed period-doubling vanished (1.30 -> 2.50, lift normalizing) - confirming it had been machine-dependent bifurcation-edge behavior that armature+friction push off the knife edge, while v7's period-doubling stood untouched: saturation-freeze-driven, a policy property, not a numerical accident.
Change
Era re-baselining protocol: after any plant upgrade, one paired same-seed sweep of all retained generations on old and new plant; verdicts keyed to identity metrics carry over, thresholds re-read against the new-plant table, and the differences themselves become plant-physics findings.
Outcome
Seven verdicts survived with evidence rather than assumption; two causes of the period-doubling family were separated with plant-side proof; and the sim's landing-violence overestimate was independently confirmed a second time.
Mechanism
A policy's structural properties (frequencies, amplitudes, frozen joints) are functions of its weights and survive plant changes; contact-mediated quantities are joint properties of policy and plant and shift when the plant becomes honest. Pairing seeds across plants isolates the plant's contribution exactly, so the sweep both validates history and measures what the old plant had been lying about.
Applies when
- installing measured masses/armature/friction into the sim
- historical thresholds are cited across a plant change
- a hardware-only behavior might be bifurcation-edge sensitivity
“策略身份指标逐位不动:主频(v7 1.20→1.20)、膝摆 … 这些是策略属性,plant 换代携带无损,历史定论因此全部成立。… 唯一行为级变化:v8@0.15 的倍周期消失(主频 1.30→2.50…)——印证当时"分岔边缘、机器相关"的判定:armature+摩擦把 v8 推离刀锋;v7 的倍周期纹丝不动(1.20→1.20),它是饱和冻结驱动的深层属性,不是数值巧合。”
train/README.md § 纪元 2c 全代同机横测 (2026-08-04): 完全体 plant 上历史结论全部存活 A nonzero response with the same sign for + and - commands is bias, not ability
same-sign-response-is-yaw-biasBefore crediting any directional skill, test both command signs: response must flip sign with the command; a same-signed pair is a bias to subtract, not an ability to report.
Symptom
Root-selection probe showed nonzero wz "tracking percentages" on turn commands, tempting the read that candidates could partially turn.
Context
During C-ladder root selection, s1e-500's measured yaw rate was +0.084 rad/s for cmd +0.3 and +0.093 rad/s for cmd -0.3 - same sign both ways. The same check on the C2 baseline gave wz+0.20 -> -0.13 and wz-0.20 -> +0.12 (again same sign), while the alternative root s2e_pd-1400 gave +0.16 / -0.16 - opposite signs, i.e. a genuine 16% command response.
Change
Reading corrected and written into the execution sheet: percentages on directional commands are meaningless unless the +cmd and -cmd responses have opposite signs; all three candidates were re-classified as "cannot turn, cannot sidewalk - C2/C3/C4 learn from zero". Acceptance criteria thereafter required "tracking >=50% AND left/right opposite-signed".
Outcome
Prevented crediting turn/sidewalk ability that did not exist; the antisymmetry clause became a standing part of every turn and sidewalk PASS condition (C2, C4, C4-redo levels all carry "且左右反号").
Mechanism
A constant yaw (or lateral) bias projects onto any command's sign convention and shows up as fake fractional tracking; only sign-antisymmetry under command reversal distinguishes a feedback response to the command from an open-loop offset.
Applies when
- evaluating turn/sidewalk/any signed-command tracking percentages
- a candidate shows partial tracking on an axis it was never trained on
- writing PASS criteria for a new directional skill
“C2/C3 那些非零的 wz 百分比不是转向能力 —— 转向+ 与 转向− 的实测同号(s1e:cmd +0.3 → +0.084,cmd −0.3 → +0.093 rad/s),那是恒定偏航偏置。… 三个候选都不会转、都不会侧走。”
train/C_LADDER_RUN.md § 0. 读数纠正(重要,别引错) Brief the operator on the lineage's measured zero-command and untrained-axis behavior before handing over the joystick
know-zero-command-behaviorBefore any teleop/demo, measure and write down the policy's zero-command behavior and per-axis competence, label untrained axes explicitly as not-bugs, and set the floor/procedure to accommodate the known drift.
Symptom
A teleop session was about to start on a policy that does not stand still at zero command and has never been trained on lateral commands - behaviors an unbriefed operator would report as bugs or emergencies.
Context
Three measured facts were written into the teleop instructions ("都有 实测依据, 不是猜"): (1) A/D (lateral) keys will get essentially no response - probe-measured sidewalk tracking ~3%, an untrained axis: "这正是 C4 要解决的事, 不是 bug"; (2) no keypress = cmd 0, and this lineage does not stand still at zero command - a three-generation lineage property: paces in place, drifts right ~5 cm/s, net rotation -30 deg/20 s; sim survival is 20/20 (it will not fall) but it walks away slowly, so leave floor margin especially on the right; (3) S (backward) WILL respond - probe-measured 20/20 survival, 67% tracking untrained, which is also why this root was chosen for the C ladder. Plus a keybinding dry-run while suspended before touching down.
Change
Operator briefing became part of the deployment artifact: expected response per key, expected idle behavior with magnitudes and directions, and the distinction between untrained (expected, not a bug) and abnormal.
Outcome
The session proceeded with correct interpretations available in advance; the known zero-command wander was handled by floor margin and start-with-command procedure rather than misdiagnosed on the spot.
Mechanism
A learned policy's off-nominal behaviors (idle drift, untrained axes) are lineage properties, stable and measurable in sim beforehand; operator surprise converts known properties into false incident reports and unsafe reactions. A briefing transfers the measured behavior model to the person holding the controller.
Applies when
- handing a learned policy to an operator or demo audience
- the policy idles in a non-stationary way at zero command
- some command axes are untrained in the current lineage
“A/D 基本不会有反应 —— s1e 从未训过非零 vy, 选根探针实测侧走跟踪率 ~3% … 这正是 C4 要解决的事, 不是 bug。… 不按键 = cmd 0, 而 s1e 在零指令下不站定 —— 血统属性, 三代实录: 原地踏步 + 右漂 ~5 cm/s + 净旋 −30°/20s。”
train/REAL_RUN_S2.md § 附: WSAD 遥控 上机前必须知道的三条 An exponential kernel on instantaneous velocity punishes gait oscillation - track the cycle average
cycle-average-tracking-for-gait-quantitiesReward velocity tracking on gait-cycle averages (or filtered values), not instantaneous samples, whenever the desired behavior oscillates at stride frequency; widening the kernel does not fix a variance penalty.
Symptom
Even while the robot genuinely sidewalked (verified after the metric fix), the Isaac-side tracking reward sat on the ignore-floor: true sidewalk scored 0.178 vs 0.189 for ignoring the command - the reward was mildly punishing the desired behavior.
Context
Sidewalking is inherently oscillatory: per-frame vy std was 0.177 while the tracking kernel width was sigma = 0.15, applied to the instantaneous value. A kernel-width scan showed widening sigma 0.15 -> 0.50 still loses (-0.14 -> -0.07): "指数核惩罚的是方差,而侧步天生带方差" (the exponential kernel penalizes variance, and side-stepping inherently carries variance). Modeling with measured parameters: replacing instantaneous vy with the mean over one gait cycle (0.5 s) flips the margin decisively (true sidewalk 1.888 vs ignore 1.281, +0.607), half-cycle is neutral (+0.006), two cycles adds nothing more. Explicitly flagged as extrapolation pending Isaac-side implementation. This also vindicated a previously dismissed external note (sigma too small) - right conclusion, different mechanism than claimed (variance, not gradient).
Change
Proposed fix recorded: change the tracked quantity from instantaneous vy to a one-gait-cycle running average; widening sigma alone rejected by the scan.
Outcome
Diagnosis complete and quantified; the C4 product shipped via feed-forward before the reward change was implemented, so the cycle-average fix remained a verified-by-model, not-yet-trained change.
Mechanism
E[exp(-(v-c)^2/sigma^2)] decreases with Var(v) even when E[v] = c exactly; a gait's phase-locked oscillation guarantees variance at the stride frequency, so instantaneous tracking rewards structurally prefer standing still at the command mean. Averaging over exactly one cycle removes stride-frequency variance while preserving command-following error.
Conflicts
The cycle-average fix itself is model-extrapolated ("⚠️ 这一条是外推,须在 Isaac 侧实装并复量后才能当结论") - the diagnosis is measured, the remedy untested in training at the time of writing.
Applies when
- tracking rewards for lateral/turn/any oscillation-carrying velocity
- a verified behavior scores below the ignore-floor
- choosing sigma for exp-kernel tracking terms
“侧走时 vy 的逐帧摆幅 std = 0.177,而 track_lin_vel_y_exp 核宽 σ = 0.15,且作用在瞬时值上 … 真侧走(均值 66%,振荡 ±0.18)0.178 | 完全无视指令 0.189 … 真侧走的得分比无视指令还低。… σ 从 0.15 放到 0.50,侧走仍然吃亏 … 把跟踪目标从瞬时 vy 换成一个步态周期(0.5 s)的平均 vy:… 1.888 vs 1.281”
train/C_LADDER_RUN.md § 3n. 二/三 Isaac 训练奖励为何一直坐在「无视底分」/ 修法不是放宽 σ Training-log reward values and fixed-command eval values live on different distributions - comparing them once claimed a 44% improvement that was really 6-10%
same-distribution-reward-comparisonQuote reward-term values only with their distribution attached (command range, DR on/off, environment), and compare across runs only when those match; re-measure in a common environment before claiming any improvement percentage.
Symptom
A v6-era analysis concluded slip had dropped 44% by comparing the training log's Episode_Reward against values calibrated in a fixed-command play environment; a same-condition re-measurement showed the true improvement was 6-10%.
Context
The training log's reward is an expectation over the training command distribution (vx 0.15-0.5, yaw +/-0.6, with pushes and domain randomization); play-environment calibrations are taken at a single fixed command with DR off. Subtracting one from the other compares apples to oranges - the warning was written into the v7 pre-flight: "奖励数值只能在同一指令分布下比较 … 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次)".
Change
Rule adopted: any before/after reward-term comparison must hold the command distribution, DR state, and evaluation environment fixed; training-log values compare only against training-log values of runs with identical command/DR configs.
Outcome
The phantom 44% improvement was retracted; later term-level accounting (e.g. the C4 ignore-floor work) consistently specified its distribution before quoting numbers.
Mechanism
A reward term's expectation depends on the visited-state distribution as much as on the policy; changing the command distribution or DR moves every term's baseline. Cross-distribution differences therefore measure the distributions, not the policy change.
Applies when
- comparing reward telemetry across training runs or vs play evals
- claiming improvement percentages from training logs
- term-level reward accounting for diagnosis
“奖励数值只能在同一指令分布下比较。训练日志的 Episode_Reward 是在训练指令分布上算的(vx 0.15~0.5 / 偏航 ±0.6 / 带推力与域随机化), 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次: 据此以为滑移降了 44%, 同条件对拍只有 6~10%)。”
train/WALK_V7_SPEC.md § 3. 开训自查 ⚠️ Changing the gait clock silently flipped a hardwired threshold's meaning - write derived constants as expressions
derived-constants-must-track-their-baseBefore changing any base parameter (clock, control rate, scale), enumerate every constant derived from it and every constant that must NOT change; convert derived literals into expressions of the base so the next change cannot silently flip a term's meaning.
Symptom
Slowing the clock 0.40 -> 0.50 s would have silently inverted the feet_air_time threshold's semantics: the 0.25 s threshold was hardwired, so at ct 0.40 the swing window (~0.20 s) sat below it (constant pressure to lengthen strides), while at ct 0.50 the window (~0.25 s) equals it - the term's meaning flips from "push longer" to "neutral" with no code error anywhere.
Context
The clock change audit walked every dependent quantity: most followed automatically (joint_pos_ref / clearance / contact_number cycle_time params, gait_phase observation, deploy/sim2sim/policy_io, export) - wiring confirmed, zero hand edits; the air_time threshold was the one hardwired constant, fixed by preserving the RATIO: 0.25 -> 0.3125 = 0.625 x ct, with the recommendation to commit it as the expression 0.625*ct "一劳永逸" (solved once and forever). The same audit also listed what must NOT follow the clock (50 Hz control rate, physics dt/decimation, 47-dim contract, action_latency absolute seconds, PD/torque limits) - the change's blast radius stated in both directions.
Change
feet_air_time threshold re-expressed as a fraction of cycle_time; auto-following vs must-not-change lists written into the spec for the clock migration.
Outcome
The clock migration (v10, repeated in v11) carried no silent semantic flips; the expression form removed the trap for every future clock change.
Mechanism
Constants derived from a base parameter encode a ratio at their birth; storing the evaluated number severs the dependency, so changing the base leaves stale semantics with no failing test. Expressions preserve the intent; and an explicit both-directions dependency list (follows / must-not-follow) is what makes a base-parameter change reviewable.
Applies when
- changing gait clock, control frequency, or units
- a reward threshold interacts with a phase/window duration
- config audit finds literals that encode ratios
“feet_air_time 阈值 0.25 是写死的,不跟 ct 走——0.40 时摆动窗 ~0.20s<0.25(恒拉长压力),0.50 时摆动窗 ~0.25s≈阈值(语义翻转)。按比例保原压力:0.25 → 0.3125(=0.625×ct;建议直接写成 0.625 * ct 表达式,一劳永逸)。”
train/WALK_V10_SPEC.md § 3. T —— 慢时钟 (训练侧必做一件) Push-test protocol - positive side first, fragile side spotted, axes aligned in the log, and cross-machine push counts stay qualitative
push-test-chirality-protocolOrder disturbance tests from the robust side to the fragile side with protection scaled to sim-measured asymmetry, align and log frame conventions before testing, and treat cross-domain disturbance counts as qualitative evidence only.
Symptom
Hand-push testing on hardware risked falls on a side sim had already flagged as fragile, and push counts invited apples-to-oranges comparison with sim numbers.
Context
Sim chirality was explicit: descendants were far more fragile in -y (fric-3000@kd1.2: +6 N*s survived 15/20 vs -6 N*s only 3-9/20) while the s1e control was perfectly symmetric (40/40). The protocol therefore: push the positive direction first, keep a spotter for the negative side; before any push, record which real-robot side corresponds to sim's +y in the log ("上机前对一次坐标"); and - citing the chaos lesson ("混沌课文") - real push results are used only as qualitative corroboration, never compared numerically with sim survival counts across machines.
Change
Push testing became a scripted, chirality-aware protocol with frame alignment as a logged precondition and an explicit epistemic limit on cross-domain count comparison.
Outcome
The fragile side was tested with protection informed by sim's quantified asymmetry; logs stayed interpretable because the frame correspondence was recorded before the first push.
Mechanism
Disturbance-response chirality is a real, quantifiable lineage property, so test order should follow measured fragility; and perturbation outcomes are chaotic in the details (divergent trajectories from tiny differences), so counts do not transfer across domains even when qualitative rankings do.
Applies when
- planning push/disturbance tests on hardware
- sim shows directional asymmetry in disturbance survival
- someone proposes comparing real push counts to sim counts
“先正向后负向, 负向留人扶 —— sim 手性明确: 后代在负 y 向显著更脆 (fric-3000 @kd1.2: +6 N·s 15/20 vs −6 N·s 3~9/20), 而 s1e@0.8 两向 40/40 完全对称。上机前对一次坐标 … 跨机不做二值结论 (混沌课文): 真机推力只作定性对照, 不与 sim 计数对比。”
train/REAL_RUN_S2.md § 3. 抗推 (可选, 人手推; 做则按此协议) A soft joint-limit penalty charged the standing pose itself - the geometric-zero knee sat on its hard limit, so stand_v1 bent its knees to dodge 0.419 per step and leaned 4.1 deg forward; excluding the knee gave 0.24 deg
soft-limit-penalty-charges-nominal-poseBefore training, evaluate every penalty at the nominal pose; if a joint's soft limit sits inside the pose the task requires (a straight knee on its hard stop), exclude that joint from the soft-limit penalty and let the action clip enforce the hard limit.
Symptom
stand_v1 (retrained after the default pose moved to the CAD geometric zero and mirror augmentation was added) fixed left/right asymmetry (6.8 -> 0.0 deg) but settled at a 4.1 deg forward lean, where pure PD at the same default settled at 0.1 deg - the policy was actively pushing itself forward, which is exactly the real robot's failure direction.
Context
soft_joint_pos_limit_factor = 0.9 shrank the knee's soft limit to -/+0.1047 rad, while the geometric-zero default has the knee at q = 0, exactly on the hard limit. Standing in the nominal pose therefore paid 0.2094 x 2.0 = 0.419 per step in dof_pos_limits (alive earned only 0.5). The policy's way out was to bend the knees to -/+0.1013 rad, and the cost was the forward lean.
Change
stand_v1b: the knees excluded from dof_pos_limits (a straight knee IS the standing pose; the hard limit is still enforced by the action clip). No other change.
Outcome
stand_v1b: max tilt 0.3 deg, steady tilt 0.24 deg, asymmetry 0.1 deg, height 0.384 m - exactly nominal - with knees at -0.0007 / +0.0005 rad. It became the standing release used on the real robot, and later the standing side of the recovery switch. The same exclusion was carried into the recovery contract (knee at the clip in the standing pose) and the one-leg reward table.
Mechanism
A limit penalty whose soft boundary lies inside the nominal pose turns the nominal into a taxed state, and the policy buys its way out with whatever posture change is cheapest - here a knee bend paid for with lean.
Applies when
- a standing or default pose has a joint at or near its hard limit
- a policy settles in a small steady tilt that pure PD does not show
- soft-limit factors shrink limits uniformly across joints
“`soft_joint_pos_limit_factor=0.9` 把膝软限位内缩到 ∓0.1047,而几何零位 default **膝盖 q=0 正好压在硬限位上** ⇒ 站在标称姿态每步白扣 `0.2094 × 2.0 = 0.419` (alive 才 +0.5)。策略只能屈膝到 ∓0.1013 躲罚,代价是躯干前倾 —— 恰好是真机的 失效方向。 … 修掉"软限位罚标称姿态"后重训(`dof_pos_limits` 排除膝盖)。**前倾问题彻底消失** … 不再屈膝躲惩罚,高度正好落回标称 0.3840。”
train/README.md § 三期: 镜像对称增强 + 站立 v1 (2026-07-28) / stand_v1b (2026-07-28): 站立定版 Score left and right separately - averages hide chirality breaking that mirror augmentation does not prevent
chirality-scored-separatelyReport every mirrored skill as two numbers with an explicit gap budget; never accept an average, and never assume augmentation guarantees symmetry - measure it per lineage and treat breakage as hard to reverse.
Symptom
Policies developed quantified left/right asymmetry (e.g. C2-700 turned right at 82% but left at 67% - a 15 pp gap; push tolerance 40/40 symmetric on the root vs 17/40 on a deep-trained descendant), and averaged metrics would have reported healthy midpoints.
Context
The repo had policy-level symmetry-breaking evidence strong enough to make separate scoring a battery rule: "左右必须分开打分 … 平均 vy 跟踪会把它掩盖". Notably, chirality broke and never recovered even though mirror augmentation (command-level mirror_prob 0.5) was on the whole time - augmentation reduced but did not prevent asymmetry, and once broken it stayed broken through subsequent rungs. PASS conditions therefore carried explicit symmetry budgets (left/right tracking gap <=10 pp), and sim's predicted asymmetry (700: right faster than left) was flagged for direct real-robot timing confirmation.
Change
Battery rule: every directional skill reports left and right (CW/CCW) as separate rows with a max-gap budget; mirror augmentation treated as mitigation, not proof of symmetry.
Outcome
The 700-vs-A800 asymmetry gap (15 pp vs 7 pp) became a first-class selection criterion; C4 product shipped with a measured 5 pp gap.
Mechanism
Averaging over mirrored conditions cancels antisymmetric error exactly where it matters; and symmetry lost during training is a lineage injury (like plasticity loss) that later rungs do not spontaneously heal, so it must be gated, not assumed.
Applies when
- evaluating turn/sidewalk/push-recovery or any mirrored skill
- relying on mirror/symmetry augmentation
- selecting between checkpoints with similar average scores
“左右必须分开打分(left/right lateral、CW/CCW turn 各自一行)—— 本仓已有 policy-level symmetry breaking 的量化证据,平均 vy 跟踪会把它掩盖。”
train/C_LADDER_RUN.md § 5. 固定验收矩阵 (左右分开打分) A binary reward band on the swing knee had zero gradient everywhere below it, so the one-leg policy parked in an unloaded "fake touchdown" that Isaac's 5 N threshold scored as success and MuJoCo showed as real pressing - a capped constant-gradient ramp, retrained from scratch, passed 40/40
binary-band-reward-fake-touchdownShape approach-to-target rewards as capped ramps with gradient from the starting posture, never as bands or indicators; and compare contact-based terms across simulators, because a policy riding just under a force threshold looks perfect in one and wrong in the other.
Symptom
At iteration 1,000 of the first one-leg run the swing foot never lifted: the policy stood with the "raised" foot resting lightly on the ground. In Isaac the contact-match term paid 96% of full marks; the same policy in MuJoCo pressed that foot on the ground for 450 frames.
Context
The swing-leg goal was "shank folded fully back" (knee 1.5-1.95 rad), rewarded as a binary band: +0.8 inside [1.5, 1.95], zero elsewhere. From knee 0.05 to 1.5 rad the term was flat. Contact is judged at a 5 N force threshold, so a foot carrying less than 5 N counts as lifted. The walk line had hit the same disease with a binary indicator (v4) and fixed it with a capped ramp (knee_swing_amplitude).
Change
swing_knee_fold changed from the binary band to a ramp clamp(|q|/1.5, 0, 1) - a constant gradient capped near 86 deg - and the policy was retrained from scratch (V0r1). After the first real-robot try showed the fold still too low, its weight went 0.8 -> 2.0 (V0.1).
Outcome
V0r1 model_2300 passed the full acceptance 40/40 (swing knee 1.72 rad, about 98.5 deg) and was stamped as oneleg_v0.onnx; the cross-simulator disagreement is recorded as the thing that caught the cheat.
Mechanism
A reward that is flat until the target is reached gives no gradient to approach it, so the policy settles for the nearest state other terms reward - here, a foot that satisfies the contact threshold without lifting; a second simulator with different contact force resolution exposes such threshold-riding.
Applies when
- rewarding a posture target with an in-band / out-of-band indicator
- a contact threshold decides whether a foot counts as lifted
- trainer-side contact terms are near full marks while the video looks wrong
“初版二值带 [1.5,1.95] 在膝 0.05→1.5 全程零梯度,策略停在"卸力虚点地"(Isaac 5N 阈下 contact_match 96% 满分 / MuJoCo 同策略 450 帧实压——跨仿真器互证抓作弊);v4 二值指示同型病,按 knee_swing_amplitude 判例改常数梯度封顶 ramp,从零重训 … **oneleg_v0.onnx = V0r1 model_2300, 40/40 PASS**”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §4 奖励表 swing_knee_fold 行 / §8 核查单 5 The +/-50 mm lateral COM randomization meant to spread the legs coincided with legs pulling IN - rolled back per its own pre-registered contract
com-dr-rollback-on-symptomWhen adopting a DR value that covers no local measurement, write its intent and rollback trigger into the config at adoption time; roll it back as the control arm the moment the symptom contradicts the intent, and promote the symptom's metric into the acceptance battery.
Symptom
After v7 adopted the reference developer's oversized lateral COM randomization (+/-50 mm) explicitly to force leg spread, the real robot's legs narrowed instead - lateral mean 154 mm / closest 107 mm in sim (nominal 214.5), narrower still on hardware with occasional leg contact.
Context
The rollback was clean because the adoption had been honest: the robot.yaml comment recorded the intent AND that the +/-50 value covered no local measurement (only a 16/7 mm measured offset existed; even the prior widening to +/-20 was subjective), plus the reference's own reported side effect (base sway) and the note "这一项要单独跑、 单独归因". When the opposite symptom appeared, v8 returned y to +/-20 mm as the control arm ("要么没起作用、要么帮了倒忙 … 按约定退回做 对照"), kept x/z untouched (a noise-level difference not worth another variable), and named the second suspect: the landing penalty itself, via the reference's own three-link chain (landing penalty -> stance narrows -> spacing penalty needed). A gate lesson was booked in the same table: v7's sim numbers had ALREADY crossed the line (154/107 vs v5's 182/147) - "这个指标本可拦下 v7" - so foot-distance became a standing acceptance row (min >120 mm, zero leg-leg contacts).
Change
base_com_offset_m y: 0.050 -> 0.020 (x/z kept), regenerated through the export tool rather than hand-editing derived files; foot-distance acceptance row added.
Outcome
A borrowed DR lever with no local measurement basis was retired the moment its symptom contradicted its purpose, at single-variable cost; the metric that would have caught it pre-hardware entered the gate.
Mechanism
DR ranges shape behavior through the policy's robustness strategy, which is jointly determined with every reward term; a lever that forces stance width on one robot can be dominated by a stronger narrowing pressure (landing softness) on another. Levers adopted without local measurement must carry their own rollback trigger, because there is no nominal to argue from when they misbehave.
Conflicts
Causality is not fully closed in the source: the narrowing may come from the landing penalty rather than the COM lever ("腿距的第二嫌疑人是 ④ 本身"); the rollback is the pre-agreed control experiment, not a verdict that the lever caused the narrowing.
Applies when
- importing DR ranges or behavioral-forcing randomizations from references
- a DR lever's observed effect contradicts its documented purpose
- a sim metric existed that would have caught a shipped regression
“⑥ 的本意 … 是逼策略把脚分开;真机结果是脚向内收且偶发相碰——要么没起作用、要么帮了倒忙。… 注释当时就写了"这一项要单独跑、单独归因"。现在症状出现了,按约定退回做对照。”
train/WALK_V8_SPEC.md § 3. 改动 C — 质心随机化退回(撤销 v7-⑥ 的 y 项) Every acceptance run records video of the very rollout that produced the numbers - the seated basin, edge-standing feet, tangled legs and the narrow stance were all seen on video before, or instead of, a metric catching them
video-as-acceptance-recordMake video a default output of every acceptance run, rendered from the same rollout the metrics come from (fixed views including the feet), and watch it - posture failures are visible before any gate row exists for them; never let video replace or override the numeric gate.
Symptom
Posture failures in the recovery line kept arriving as things the numbers had no row for: a kneeling W-sit, feet standing on their outer edges, crossed legs, a stance too narrow to hold on hardware.
Context
From 2026-08-09 (user decision) accept_recovery renders offscreen by default, following one env for the whole episode and archiving the clip. The MuJoCo gate's --video renders three views (side, front, feet) of the same rollout the metrics come from, with all plant modelling (delay, push); the older replay-based renderer produced an independent trajectory without delay and was not used for acceptance. Video never gates: if rendering breaks, --no-video keeps the numeric gate running.
Change
Video as a default acceptance artifact, named per policy, category and view, reviewed by the user.
Outcome
The R0.1 prone clip showed the same kneel-sit as R0; the R3.1 failure clip showed the crossed legs; the v2_5 feet view showed edge standing and led to V2.6; the v2_6 videos led the user to order a real-robot A/B between v2_5b and v2_6. Once, the recorder did not start (P1c final acceptance) and the visual material had to be produced separately.
Mechanism
Metrics exist only for failure modes someone anticipated; video shows the unanticipated ones, and rendering the metric rollout itself guarantees the picture and the numbers describe the same episode.
Applies when
- setting up an acceptance pipeline for posture-sensitive skills
- numbers pass but a human reviewer is uneasy
- sim videos are rendered by a separate replay tool
“**验收存视频(用户定 2026-08-09)**:`accept_recovery.py` 默认开 Isaac 离屏 渲染,跟拍一个 env 的整局并归档 … 跪坐这类盆地在数字表出现前肉眼先看见,视频是验收的 定性存档,数字门不受它影响”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §5 R0 验收门(预注册)验收存视频(用户定 2026-08-09) Training at scale 0.4 is NOT the twin of deploying 0.5 at power 0.8 - the algebra matches, the learned policy does not
deploy-scaling-not-training-equivalentNever assume deploy-side scalings can be folded into training-time constants ("burning the crutch into training"): the learned optimum depends on the training-time authority, so treat such conversions as full experiments with pre-registered expectations and a sim2sim gate before any hardware.
Symptom
s1g (S1.6) trained from zero at action_scale 0.4 - meant as the "training twin" of the hardware-proven s1c-at-power-0.8 (0.8 x 0.5 = 0.4) - was all green in Isaac (zero falls, reward 117) yet scored 0/3 across all eight checkpoints and 0/20 at 20 seeds in the MuJoCo gate, falling forward at median 1.57 s with a 2.9x speed overshoot.
Context
The pre-registered expectation (survival gate should pass, since the conviction matrix showed s1c@0.8+delay2 all-survive) was cleanly falsified, and the harness was acquitted by controls: --delay 0 fell identically (not a delay fragility), check_contract all green, and s1c through the same harness survived 2/3. The verdict: "「s1c@0.8 = 0.4 训练孪生」的代数等价不成立" - a policy deployed with a derated output still LIVES in the 0.5 internal model it trained under (its value function, its expectations of its own authority), while a policy that starts training with reduced authority learns a different, clip-hugging gait with zero margin for plant differences ("部署端打折的策略活在 0.5 的内模里,训练起点收权限学出的是贴 clip 的 另一套步态,对 plant 差异零余量"). Result: the policy was withdrawn before hardware ("撤回——不上真机"), the lineage root moved back to the 0.5-contract s1c-5500, and this became the C ladder's cited fact-check ("s1g 是 0/20 证伪出局的那一代").
Change
The amplitude-surgery route abandoned; contract kept at scale 0.5; the deploy-side 0.8 crutch later retired on its own merits when the delay-complete s1e generation ran at full power.
Outcome
One training run bought a clean falsification of a plausible algebraic identity; no hardware time was spent on it because the sim2sim gate caught it.
Mechanism
Output scaling commutes with the network arithmetic but not with learning: the training-time scale shapes which gait solutions are reachable and how much clip headroom the optimum keeps. A derated mature policy retains the wide-authority solution executed softly; a from-zero narrow-authority policy finds a different optimum that saturates its smaller envelope - the two are not the same controller in different units.
Applies when
- proposing to move a deployment derating into a training constant
- a scaled-down contract policy hugs the action clip
- Isaac-green / cross-sim-zero results on a re-scaled lineage
“预注册 a) 证伪——Isaac 全绿(零摔/reward 117)但 MuJoCo --delay 2 八档 checkpoint 扫描全数 0/3、iter6500 20-seed 0/20 … 「s1c@0.8 = 0.4 训练孪生」的代数等价不成立: 部署端打折的策略活在 0.5 的内模里,训练起点收权限学出的是贴 clip 的另一套步态,对 plant 差异零余量。”
train/OMNI_V0_SPEC.md § 3. S1.6 判决(2026-08-07 验收) When training fails repeatedly, inject the target behavior open-loop - stop tuning rewards for an unverified behavior
open-loop-probe-before-reward-tuningAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop on the real plant/sim first, and only resume training once you hold a measured, safe, sign-verified target trajectory.
Symptom
Three sidewalk training rounds failed; hypotheses multiplied (exploration failure / wrong reference waveform / insufficient authority) with no way to pick between them by running more training.
Context
Instead of a fourth reward guess, the team wrote probe_side_ref.py: the candidate reference is injected open-loop on top of a frozen policy's output (bypassing PPO entirely), directly measuring "what happens if the robot literally does this waveform" - separating all three hypotheses in one experiment (5 seeds x 8 s per condition, several waveform families and gains). The probe immediately eliminated the authority hypothesis (full-amplitude execution, 5/5 survival) and localized the problem to the waveform/measurement side. The closing principle was written down after the saga: without a verified target behavior, tuning rewards is "在黑暗里试钥匙" (trying keys in the dark).
Change
Standing method: before opening another training rung on a failing skill, build an open-loop (or task-space) generator of the intended behavior, measure whether the physical system can express it and what it looks like - then train toward a verified, quantified target.
Outcome
The probe chain produced the verified waveform (reversed-sign triangle, half gain), quantified safe amplitude (tilt 8.2 deg at band top, foot distance clear of the wall), exposed the metric bug when probe and training disagreed, and ultimately supplied the feed-forward that made C4 pass in +100 iters.
Mechanism
Training couples exploration, reward design, and feasibility into one opaque outcome; open-loop injection cuts the loop and tests feasibility and waveform alone. A behavior demonstrated open-loop converts the remaining failure into a pure credit-assignment/reward question - and its measured trajectory becomes the reference itself.
Applies when
- repeated training failures on one skill with multiple live hypotheses
- uncertainty whether the platform can physically express the behavior
- a reference trajectory's shape/sign/amplitude is guessed, not measured
“三轮 FAIL 之后不再猜,写 train/probe_side_ref.py 把参考开环注入到策略输出之上(绕过 PPO),直接量「照这个波形做会怎样」,一次分开三个假说:甲 探索 / 乙 波形 / 丙 权限。”
train/C_LADDER_RUN.md § 3f. C4 真因定谳(开环探针) / 3k. 建议的下一步 Never referee a suspect metric with another metric from the same code - they can share the disease
independent-referee-for-metric-disputesTo adjudicate a disputed measurement, compute the quantity by an independent method from raw state; never accept a sibling column from the same pipeline as the tiebreaker.
Symptom
A triple reversal on one question: sidewalk sign diagnosis (correct) was retracted using a second metric from the same script, then the retraction itself had to be retracted when that second metric turned out to be the buggy one - two opposite-direction errors on the same problem in one day, both written into the execution sheet.
Context
The probe's net-displacement metric suggested the sidewalk reference sign was inverted. Worried about yaw-drift pollution of net displacement, the author checked the same table's body-frame vy_mean column (~0.003 everywhere, 20-50x smaller) and retracted the sign diagnosis. But vy_mean came from mj_objectVelocity, which was silently reporting vertical velocity due to a frame bug - the "referee" was the diseased measurement. Re-measured with a truly independent computation (xmat.T @ qvel, world trajectory), the original diagnosis was confirmed: saw -0.5 gave vy +0.058/-0.130 (76%/106%), consistent with the net-displacement values all along (yaw pollution was real but only 10-21 deg, nowhere near reversal-sized).
Change
Lesson written twice, verbatim, as a hard rule: when questioning a measurement, the referee must be an independent algorithm (different code path, different physical derivation), e.g. rotate qvel by the body matrix directly, or inspect the raw world trajectory.
Outcome
With the independent referee in place the frame bug was confirmed, fixed, and the whole C4 line re-scored - revealing sidewalk had been working (see body-frame-velocity-api-audit).
Mechanism
Metrics sharing a code path (or an upstream API) share failure modes; agreement between them is evidence about the code, not the world. Only a measurement with an independent derivation can break the tie, because its errors are uncorrelated with the suspect's.
Applies when
- two metrics of the same quantity disagree
- about to retract a conclusion based on a second readout
- auditing evaluation code after a surprising result
“我用一个坏指标去质疑一个好指标,并把撤回写进了执行单。教训(写死):质疑一个测量时,不能用同一份代码里的另一个测量当裁判 —— 它们可能同源同病。裁判必须是独立算法(这次的裁判应该一开始就是 xmat.T · qvel[:3],或直接看世界轨迹)。”
train/C_LADDER_RUN.md § 3m. 二 我今天犯了两个方向相反的错 / 3n. 五 元教训 A 2-degree joint-zero calibration fix moved the whole runnable envelope - re-test old "cannot run" verdicts after recalibration
zero-offset-calibration-shifts-envelopeDate every hardware verdict with the calibration state; after any zero/mount recalibration, re-test previously condemned policy-power combinations and previously "unexplainable" posture offsets before attributing either to training or model.
Symptom
s1d was on record as "only runs at power 0.7" (kicked wildly at 0.8); after a calibration pass, the same policy ran 12 s at 0.8 with no kicking at all.
Context
The calibration had fixed a 2.08 deg zero offset on r_hip_roll - exactly the constant error source on the dominant joint of the kicking oscillation loop ("恰是乱踢振荡环主导关节的常值误差源"). The three-generation post-calibration hardware sweep also closed a second case: the robot's mysterious "backward lean" disappeared after calibration, and the sim-real posture difference collapsed from opposite-sign 5+ deg to same-sign ~2 deg ("后仰案实质了结") - the lean had been a sensing/zero artifact, not a mass-model error. Booked consequence: if the s1d recovery re-verifies, "真机可跑档整体 上移" - every policy's runnable power envelope shifts up, and downstream lineages' hardware expectations get revised.
Change
Joint-zero and mount calibration promoted from setup chore to a variable that dates hardware verdicts: verdicts about which power/scale levels a policy can run are conditioned on the calibration state they were measured under.
Outcome
One policy rehabilitated at a higher power level; one standing sim-real posture discrepancy closed without touching model or training; a pending re-verification booked rather than asserted.
Mechanism
A constant joint-zero error acts as a persistent disturbance injected at the feedback loop's most-loaded joint; near an oscillation threshold, removing a 2-degree bias is the difference between a stable and an unstable loop. Since the error is additive and machine-side, it shifts every policy's stability envelope simultaneously - which is why verdicts must carry their calibration date.
Conflicts
The s1d rehabilitation awaited one confirming re-run at the time of writing ("待复核一跑坐实") - the offset-as-cause reading is the head suspect, not a closed verdict.
Applies when
- a policy oscillates at a power level others tolerate
- sim and real disagree on a constant posture offset
- deciding whether to re-test old hardware verdicts after maintenance/calibration
“发现①:s1d@0.8 能跑了(旧账「只有 0.7 能跑」)——12s 无乱踢。头号嫌疑 = 标定修正:r_hip_roll offset 修 2.08°,恰是乱踢振荡环主导关节的常值误差源。… 发现②:「后仰」标定后消失 … sim-real 姿态差从反号 5°+ 收敛到同号 2°,后仰案实质了结。”
train/README.md § 真机 @0.8 三代横评(2026-08-07 标定后) Verify changes in the run's resolved config (and checkpoint md5), never in the source you edited
resolved-config-is-source-of-truthAttribution and single-variable claims must be made on the resolved per-run config (and checkpoint hashes), not on source diffs; verify every intended variable landed before burning compute, and verify every rollback byte-level against the historical resolved config.
Symptom
An intended arm-B config change never reached the training run - the run was grid-identical (117/117 cells) to its C2 predecessor - and the burn was only understood afterwards.
Context
The repo's discipline hardened around the logged resolved config (logs/<run>/params/env.yaml) as the only source of truth: (1) the C2 root-cause analysis was performed against the checkpoint's logged env.yaml, not the code ("以真相源 23-19-25/params/env.yaml 核实"); (2) C4 added a pre-flight: grep the landed env.yaml for the new keys, and compare the first checkpoints of the two arms - identical md5 means the variable did not land, stop immediately; (3) the C4 full rollback was accepted only after starting a 1-iter run and byte-comparing its resolved env.yaml against the historical 700-era file (identical except 4 dormant schema fields, each verified to be at its no-op default).
Change
Standing pre-flight and post-change verification: dump/diff the resolved config that the run actually consumed; use checkpoint hash equality as a cheap "variable landed" detector between arms.
Outcome
Caught the not-landed variable class of failure; made the rollback provably equivalent to the historical training state rather than believed-equivalent.
Mechanism
Between edited source and the running experiment sit layered overrides, env-var switches, and registration logic; only the resolved, serialized config reflects their composition. Diffing at that level tests the actual experiment; diffing source tests intent.
Applies when
- launching an A/B pair or any single-variable rung
- rolling back to a historical training state
- a run behaves as if a change was never applied
“开训前先验落盘 cfg(上一轮臂B 的改动没进 run,与 C2 逐格 117/117 相同):grep -E "base_com|joint_friction|push_robot|track_lin_vel_y_exp" logs/<run>/params/env.yaml 另:两臂第一个 checkpoint 的 md5 若相同 = 变量没进去,立刻停。”
train/C_LADDER_RUN.md § 3d. ⚠️ 开训前先验落盘 cfg / 3l. 回退清单(验证) Single-impulse push recovery is a binary chaotic quantity - cross-machine floating-point divergence can flip the outcome
single-impulse-recovery-is-chaoticNever gate or compare single-event recovery outcomes across machines or domains: evaluate disturbances as survival distributions over phases and seeds, compare longitudinally on one machine, and treat any single-point cliff as unconfirmed until it survives the statistical protocol.
Symptom
Mac evaluation found a hard "0.8 N*s cliff" (0/3 survival) that the training machine flatly contradicted: the identical protocol (0.8 impulse at 8 s, cmd 0.2) survived 3/3 there, and a 0.6/0.8/2/4 cross sweep survived everything.
Context
The verdict became a named lesson ("跨机混沌课文"): whether one specific push at one specific phase is survived depends on a trajectory that diverges across machines from floating-point differences alone - "单次冲量恢复是二值混沌量, 跨机浮点发散可翻结局". The boundary was drawn precisely: the 20-seed statistical gates DO agree across machines (established precedent), but that agreement cannot be extrapolated to single-point recovery tests. Protocol amended: disturbance evaluation uses multiple push phases (8/10/12 s), >=10 seeds, and only same-machine longitudinal comparisons; the Mac-side recommendation built on the unreproducible cliff was not adopted, while its directionally-consistent small-impulse data was kept.
Change
Push evaluation redefined from single-event pass/fail to multi-phase multi-seed statistics, with cross-machine comparison banned for event-level results and allowed for distribution-level ones.
Outcome
A false hardware-relevant "cliff" was prevented from steering the ladder (the s2e push rung decisions were made on same-machine statistics); the chaos lesson was cited again when real push tests were restricted to qualitative cross-domain use.
Mechanism
Perturbation recovery near the viability boundary has sensitive dependence on initial conditions; different BLAS/GPU reduction orders yield different trajectories from identical configs, so a binary outcome at one phase is machine-specific noise. Averaging over phases and seeds restores a quantity whose expectation is machine-stable.
Applies when
- a push/disturbance result differs between machines or sim and real
- designing push-recovery acceptance tests
- a sharp pass/fail cliff appears in a chaotic-regime evaluation
“训练机上 Mac 原协议 (0.8 @8s cmd0.2) 3/3 全活 … 与 Mac 的 +0.8 0/3 直接矛盾。定性: 单次冲量恢复是二值混沌量, 跨机浮点发散可翻结局;统计门 (20-seed 八门) 跨机吻合的先例不能外推到单点恢复测试。协议改判: 抗推评测多相位 (push 时刻 8/10/12s) + ≥10 seed + 只做同机纵向比”
train/README.md § s2e 支线终章 (跨机混沌课文) The latency DR range must cover the measured deployment pipeline - 0-20 ms could not even reach the real 1-2 control steps
latency-dr-covers-measured-pipelineMeasure end-to-end action latency in control steps on your own stack (including cross-process queue boundaries), set the DR range to cover it with margin, and never import a delay count without its control frequency.
Symptom
Action latency was randomized over 0-20 ms (0-1 control step at 50 Hz), but the measured deployment path is 1-2 steps: the deploy process writes the target, an independently running BusWorker picks it up on its NEXT cycle, plus CAN round-trip - the training range could not cover the robot's actual latency at all.
Context
Fix: widen action_latency_s to 0-0.06 (0-3 steps). The external reference's "uniform 6 steps" was explicitly NOT copied - that number depends on his unknown control frequency; locally, a sweep at 0/1/2/3 steps showed walk_v5 survives all with insensitive metrics, so 6 steps "在我们这里没有依据" (has no local basis). The range was set from the measured pipeline with margin, not from a foreign constant.
Change
action_latency_s (0, 0.02) -> (0, 0.06), justified by pipeline analysis (writer/worker cycle boundary + bus time) and bounded by the local latency sweep.
Outcome
The DR band now brackets the true deployment latency; the policy trains against the delay it will actually face instead of a fictional sub-step world.
Mechanism
Latency DR only immunizes against delays inside its support; a range below the physical pipeline guarantees an untrained distribution shift at deployment. The correct range comes from tracing the pipeline's worst case (queueing boundaries + transport), and foreign step-counts are meaningless without the control rate they were measured at.
Applies when
- setting or auditing action-delay randomization
- deployment uses a separate bus/worker process from the policy loop
- importing delay-modeling numbers from other projects
“现行 0~20 ms = 0~1 个 50Hz 控制步, 而实测部署链路是 1~2 步(deploy 写 STATE.target 后, 独立跑的 BusWorker 下一轮才取走下发, 再加 CAN 往返)——现在的区间覆盖不到真机的实际延迟。… 不照抄参考来源的"统一 6 步": 那取决于他的控制频率(未知), 而我们扫过 0/1/2/3 步 … 6 步在我们这里没有依据。”
train/WALK_V7_SPEC.md § ⑤ action_latency_s 0~0.02 → 0~0.06 Three same-shaped judging errors - task metrics (survival, tracking, displacement) cannot stand in for posture metrics
task-metrics-vs-posture-metricsKeep validated posture-class rows (tilt, per-joint L/R asymmetry, temperature) in every acceptance battery alongside task rows; when operator feel contradicts the gates, suspect the metric class before the operator - and never build a new skill on what is actually an asymmetry defect.
Symptom
The C2 product judged "full pass" on task metrics (A800: turn-gap 7 pp, vx+0.30 19/20) felt WORSE in the operator's hands than the half-pass 700: A800 tilted up to 12.90 deg (700: 6.64), drifted left while standing, showed larger per-joint asymmetries, and ran its hip_rolls 5 degC hotter.
Context
The re-judgment catalogued three same-type metric errors in one campaign: (1) stand judged by SURVIVAL - missed 0.5-1.4 m wandering; (2) stand ranked by DISPLACEMENT - ordering was opposite to real feel (tilt ordering matched); (3) chirality judged by wz-tracking GAP - measured turning symmetry while the robot's actual disease was postural left/right asymmetry, "两个不同的东西,且结论相反". Common pattern named: "我一直用「任务指标」当判据,而真机手感对应的是「姿态 指标」… 任务类指标不能替代它". The fix was already in the data: the per-joint left/right asymmetry table (printed identically by sim2sim and deploy) agreed with hardware in direction on every row - "判据可用、有预测力,我只是没把它写进 PASS 条件". Shipping decision followed the posture read: product reverted to 700 ("又一次「买到 精度、卖掉别的」"), and the C4 root moved to 700 as well, with the sharpest line of the episode: A800's left-drift "like sidewalking" is probably its frontal-plane asymmetry defect, not a capability - "在缺陷上建能力是危险的".
Change
Two posture quantities with demonstrated real-robot predictive power promoted into every PASS battery: tilt-max median and per-joint left/right asymmetry (both sim-computable, deploy-homologous); motor-temperature readout added to session close-out.
Outcome
Deployment flipped to the posture-better checkpoint; the hip_roll temperature table (43-48 degC vs 25-28) confirmed the earlier 90%-of-heat account; the run-line acceptance battery inherited the posture rows from birth ("任务类替代不了姿态类").
Mechanism
Task metrics measure goal attainment under the evaluator's episode definition; posture metrics measure the body state trajectory that operators, motors, and downstream skills actually experience. The two can rank candidates oppositely because task success tolerates postural pathology - so a battery without posture rows is blind to exactly what hardware feel reports first.
Applies when
- hardware feel disagrees with a green acceptance table
- choosing between checkpoints that split task vs posture metrics
- selecting the root for a skill that resembles an existing defect
“共同模式:我一直用「任务指标」(存活 / 跟踪率 / 位移)当判据,而真机手感对应的是「姿态指标」(倾角、逐关节左右不对称)。→ 验收判据里必须有姿态类指标,任务类指标不能替代它。… A800 的「左飘像 side walk」很可能 … 是它更大的额平面不对称的表现 —— 在缺陷上建能力是危险的。”
train/README.md § C2 选点改判 (2026-08-09): 手性判据第三次选错指标 IMU observation age cut 52-68 ms to ~4 ms by moving AHRS onto the MCU - as a single variable
imu-age-move-fusion-downstreamAudit observation age end-to-end and move time-critical fusion as close to the sensor as possible - and when you fix a latency, change only that one variable so the gain is attributable.
Symptom
IMU-derived observations reaching the policy were 52-68 ms old because attitude fusion ran in Python on the loaded host computer - stale attitude is a direct feedback-loop delay the policy was not trained with.
Context
The fix was scoped deliberately narrowly: move the AHRS computation from Python to the STM32 H7 (MC02). CAN topology explicitly unchanged, so the change is a clean single variable.
Change
AHRS fusion relocated Python -> H7. Before/after - IMU age: 52-68 ms -> ~4 ms; CAN timing: unchanged; Python load: high -> ~0.
Outcome
IMU age reduced by an order of magnitude with no confound; host CPU headroom recovered ("把计算单元搬在stm32上, 这样imu有剩余").
Mechanism
Sensor age is pipeline latency, not sensor quality: fusing on the MCU next to the sensor removes host scheduling jitter and interpreter overhead from the critical path. Keeping the bus topology fixed makes the improvement attributable to the relocation alone.
Applies when
- measured sensor-to-policy age far exceeds sensor sample period
- attitude fusion or filtering runs on a loaded host CPU in an interpreted runtime
- planning infrastructure changes during a sim2real campaign
“AHRS 搬到 H7——这个不改 CAN 拓扑,只是把一段计算从 Python 挪到 MC02,单变量:IMU age 52–68 ms → ~4 ms / CAN 时序 不变 / Python 负载 高 → ≈0”
Experience.md § AHRS 搬到 H7 (lines 28-35) Teleop fed the sidewalk axis a command beyond its training band - feet clipped; give each axis its own speed setting
teleop-command-band-per-axisGive every command axis its own teleop scale, clamped to that axis's training band, and reproduce any hardware incident in sim with the exact deployed command values before touching training.
Symptom
Robot stepped on its own foot when sidewalking left under teleop - and only when going left.
Context
The teleop tool used one speed setting for all axes: --teleop-speed 0.20 applied to A/D sent cmd_vy = 0.20, above the training band's top (0.08-0.18) where foot-spacing margin is thinnest. Sim reproduction of the incident (product policy, pw0.8, 5 seeds x 20 s, true collision threshold = single foot width 104 mm): at vy 0.20 the minimum foot distance was 111-115 mm - 7-11 mm from self-collision - vs 147 mm at vy 0.10. Left was 4x more dangerous than right (25% vs 6% of time inside the 160 mm soft wall at vy 0.10), matching the left-only symptom; the margin did not degrade over time (pressing more just lengthened exposure).
Change
deploy_policy gained --teleop-side (default 0.10), separating the lateral speed from the forward speed so each axis's teleop command sits inside its own trained band.
Outcome
Command now inside the band with 43 mm margin at default; the incident became a quantified, reproduced, closed account rather than a mystery.
Mechanism
The policy's competence envelope is the training command distribution per axis; teleop mappings that share one scalar across axes silently command out-of-band inputs on the weakest axis. Asymmetric risk (left vs right) came from the policy's own chirality bias, so a symmetric command produced an asymmetric hazard.
Applies when
- wiring a joystick/teleop layer over a learned policy
- a hardware incident occurs on one command direction only
- training bands differ across command axes
“A/D 一直与 W/S 共用速度档,所以按 A 下发的是 vy = 0.20 —— 既超训练带(0.08~0.18)上沿 … 0.20(遥控实际值)| 111~115 mm | 7~11 mm … 且左比右危险 4 倍 … 处置:deploy_policy 新增 --teleop-side(默认 0.10),侧移与前进档分开。”
train/C_LADDER_RUN.md § 3p. 一 向左走踩到自己 → --teleop-speed 0.20 同时喂给了 vy A stand gate judged by survival passes a robot that wanders a meter - judge posture instead
stand-gate-posture-not-survivalFor every gate, ask what behavior the metric is a proxy for and validate its ordering against real observations; replace metrics whose ordering disagrees with reality, and demote them explicitly rather than silently.
Symptom
Real robot "standing" drifted 0.5-1.4 m across the floor while the sim stand gate scored a clean 20/20 - because the gate's metric was episode survival, which wandering does not violate.
Context
C2's stand condition was originally written as "survival regression <=2/20". A real-robot counter-example on 2026-08-08 forced the re-judgment: wandering robots survive. Cross-checking candidate sim metrics against real-robot feel showed max-tilt median ordering agreed with hands-on ranking, while displacement ordering was actually OPPOSITE to real impressions - so displacement was demoted to a reference quantity, not a gate.
Change
Stand PASS criterion rewritten from survival to posture: "stand tilt max median <= root baseline +1.5 deg"; displacement kept only as reference. Applied to all subsequent rungs (C4 and redo levels inherit it).
Outcome
Later rungs gated stand on tilt (e.g. C4 product: 7.7 deg vs parent 7.1 deg, +0.6 deg PASS); the wandering failure mode became visible to the battery instead of hidden by survival.
Mechanism
A gate metric is a proxy for an intended behavior; survival is a proxy for "did not fall", not "stood still". Metric choice must be validated against ground truth (real-robot feel/measurement), and a proxy whose ordering disagrees with reality on real data is worse than no metric - it steers selection backwards.
Applies when
- writing PASS conditions for stand/idle/hold behaviors
- a gate passes policies that visibly misbehave on hardware
- choosing between candidate metrics for an acceptance battery
“stand 必须用位姿判,不能用存活判(2026-08-08 真机反证改判):站着乱走 0.5~1.4 m 时存活照样 20/20 —— C2 的 stand 条件原写「存活退化 ≤2/20」,选错了指标。改为: stand 倾角 max 中位 ≤ 根基线 +1.5°(倾角与真机手感排序一致;位移排序与真机相反,降为参考量)”
train/C_LADDER_RUN.md § 3b. PASS 条件 ⚠️ stand 必须用位姿判 Run acceptance under measured contact parameters - honest condim/torsional-friction flipped a false PASS into a real-matching FAIL
eval-plant-honesty-contact-paramsPin the evaluation plant's contact model to measured values (contact dimension, torsional/rolling friction, mu) before trusting any gate that involves slip, impact, or drift - a gate can only fail a policy for physics its simulator contains.
Symptom
walk_v5 passed the old acceptance battery yet failed on the real robot (footfall force, drift, kicking) - the evaluation plant was flattering the policy.
Context
The battery was re-run under "honest contact parameters" - condim 4 (adding torsional contact) with measured torsional friction 0.035 - and v5 then FAILED exactly the rows corresponding to its real problems: heading 185 deg (limit 30), support-foot yaw slip 284 deg (limit 80), landing force 1.72x (limit 1.5x), suspended tilt 45.9 deg (limit 10). The slip physics depends on torsional friction, which the default contact model (condim 3) does not even simulate - a slip problem is invisible to an evaluator that cannot represent yaw friction at the foot. Term-sizing measurements for the new rewards were likewise taken under the same honest parameters (cmd 0.45, skipping the 5 s start transient).
Change
Acceptance harness pinned to condim 4 / torsion 0.035 (measured); verdicts issued under defaults declared non-citable for these rows.
Outcome
Sim acceptance verdicts began agreeing with hardware ("现在失败, 与真机一致"); the v6 fixes could be developed and validated against an evaluator that could actually see the disease.
Mechanism
An evaluator is a plant model too: contact dimensionality and friction values decide which failure modes exist in the simulation at all. Evaluating under default contact parameters tests the policy in a world where its real failure is physically impossible, producing structurally false PASSes.
Applies when
- sim acceptance passes policies that fail on hardware
- slip/drift/impact gates run under default simulator contact settings
- setting up a cross-simulator evaluation harness
“accept_v2.py 已加三条判据, walk_v5 在诚实的接触参数下(--condim 4 --torsion 0.035)现在失败, 与真机一致:直行 15s 航向累计 <30° | 185° ✗ … 落脚力峰值 <1.5× 体重 | 1.72× ✗”
train/WALK_V6_MINIMAL.md § 5. 验收 The real robot's right-leg kicking was over-trained-delay times loop gain - irreducible pipeline latency is plant, model it fully from day one
pipeline-latency-is-plant-not-drMeasure the end-to-end action pipeline delay and build it into the nominal plant and every acceptance gate from day one; treat power/scale deratings that "fix" oscillation as gain-reduction crutches flagging an unmodeled delay, and expect higher-feedback-gain policies to be MORE delay-fragile.
Symptom
On hardware, s1c/s1d at action scale 1.0 always kicked wildly with the right leg (s1c only ran as SOTA at power 0.8; s1d only at 0.7) - while sim showed nothing under default evaluation.
Context
Sim reproduced the incident item by item once the real pipeline delay was injected: s1d@1.0 with --delay 1 fell at 10.2 s, --delay 2 at 5.2 s; s1c@1.0 stressed (r_hip_roll saturation 5 -> 16%; "右脚" = the policy's chirality makes the right leg its high-gain leg); and the combos that worked on hardware (s1c@0.8+delay2, s1d@0.7+delay2) all survived in sim. Mechanism: the real pipeline is ~1-2 ticks (BusWorker next-cycle pickup + CAN round trip) but S1.1 trained only to 1 tick - "超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定" (delay beyond training x full loop gain = oscillation; the power derating had been buying stability by compressing loop gain). s1d was MORE fragile than s1c because its yaw 3-layer stack had learned higher feedback gain - higher gain, lower delay tolerance. Three changes: latency DR widened to cover reality; acceptance gates and smoke runs moved permanently to --delay 2 ("门必须在真机条件下预测 真机"); and the doctrine written twice-paid: "不可约的管线属性(延迟、 限速)不是'随机化选项',是 plant 本体,第一天就该全额建模" - S1's nominal-then-robust staging falsified by hardware for the second time. The later s1e hardware run at power 1.0 (no kicking, normal force) closed the loop: "0.8 = 旧代拐杖" - the derating had been a crutch for the under-modeled delay, not a real requirement.
Change
Latency modeled as plant from day one of any lineage (measured 1-2 ticks covered, bridge-layer rate limits likewise modeled by default); every gate and smoke evaluation issued under --delay 2.
Outcome
Kicking reproduced, explained, and eliminated in the s1e generation at full scale and full power; the deploy-side crutches (0.7/0.8) retired for the new lineage.
Mechanism
Feedback oscillation onset is a product of loop gain and phase lag; a policy trained below the real delay learns gains that sit past the real stability margin, and any output derating masks it by scaling gain down. Since pipeline delay is deterministic hardware property - not an uncertainty - it belongs in the nominal plant, and every evaluation must include it or the gate predicts a robot that does not exist.
Applies when
- hardware oscillation/kicking that sim only reproduces with added delay
- a policy only runs on hardware at reduced power/scale
- defining what belongs in the nominal plant vs the DR list
“真实链路延迟 ~1~2 拍 … S1.1 只训到 1 拍——超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定。s1d 比 s1c 更脆 = yaw 三层栈学出更高反馈增益,增益越高延迟容忍越低。… 教训入账:S1「先标称后鲁棒」第二次被真机证伪——不可约的管线属性(延迟、限速)不是"随机化选项",是 plant 本体,第一天就该全额建模。”
train/OMNI_V0_SPEC.md § 3. S1.4(真机右脚乱踢事故强制) The best checkpoint to SHIP is not the best checkpoint to CONTINUE FROM - maturity is capital against adaptation shock
root-maturity-vs-product-qualityDecide shipping points and fork roots separately: gates rank products, but a root candidate must prove itself by surviving a continuation under the next rung's shift (dual-arm if in doubt) - and prefer the more-trained point as root when product metrics conflict with maturity.
Symptom
A band re-audit found s1e-300 beat the incumbent root s1e-500 on nearly every quality gate (stepping 19/20 vs 13/20 with historically-best 26.9 mm swing, speed gate 14/20 vs 2/20, heading 26 vs 54 deg/20 s) - suggesting the root had been mis-picked and the younger point should take over.
Context
The dual-arm control settled it the other way: continuing the S2 PD rung from s1e-500 adapted smoothly (3/3 smoke throughout), while the b300 control arm (same config, from s1e-300) fell into a survival valley under the PD shock (+100 iters: 1/3 -> 0/3), never climbed out within budget, and its 800-iter product scored 13/20 survival - eliminated. Verdict: "幼年点自身指标再好也扛不住新 DR 适应冲击, 成熟度是本钱,s1e-500 根被数据背书" - a young point's own metrics, however good, do not survive new-DR adaptation shock; maturity is capital. The audit still yielded value: the band scan (200-1000, per-100) mapped the lineage's arc (200 dragging -> 300 peak -> 400+ decay -> 900+ drift blowout), and 300 remains the better PRODUCT answer for shipping-as-is questions.
Change
Selection doctrine split into two questions with different answers: best-product point (quality gates at the point itself) vs best-root point (survives adaptation shocks; more training age = more capital), each decided by its own evidence - and root claims settled by a dual-arm continuation test, not by point metrics.
Outcome
s1e-500 kept the root role with data behind it; the S2e ladder built on it passed rung after rung, while the b300 line was closed at the cost of one control arm.
Mechanism
Early checkpoints sit near sharp optima with less accumulated robustness structure; their headline metrics reflect the narrow training distribution, not resilience to distribution shifts. A continuation rung is itself a distribution shift, so the root property being selected for is shock tolerance - observable only by actually continuing, never by static gates.
Applies when
- a younger checkpoint outscores the current root on quality gates
- choosing the base for a robustification or command ladder
- a continuation run stalls in an early survival valley
“b300 对照臂 … PD 冲击下存活谷(+100 起 1/3→0/3),预算尽未爬出,800 档 20-seed 存活 13/20 出局——幼年点自身指标再好也扛不住新 DR 适应冲击,成熟度是本钱,s1e-500 根被数据背书”
train/README.md § omni_s2e_pd (b300 对照臂) / s1e 选点重审 When a contract default changes, old policies must run under a pinned legacy profile - a silent clock swap is out-of-distribution on hardware
legacy-profile-pinningTreat every trained policy as bound to the contract values of its training era: version the deployment profiles, pin old policies to their era's profile in every command, and never let a changed default silently apply to an old artifact.
Symptom
The walk profile's gait clock moved from 0.40 s to 0.50 s for new training, but versions v5-v9 were all trained at 0.40 s - running them under the updated default would silently feed a 25% slower phase clock to policies that never saw one.
Context
The re-test runbook hard-codes --policy-profile legacy_walk_040 into every command for the old versions, with the warning not to omit the flag: the mismatch is invisible (no error, no crash) but puts the policy out of distribution on hardware, where the same file had already documented that off-clock operation collapses gait quality.
Change
Deployment profiles versioned per training era; historical policies permanently associated with their era's profile; runbooks write the profile flag explicitly rather than relying on defaults.
Outcome
Old policies stayed runnable and comparable after the contract moved on; the silent-mismatch failure mode was closed by convention.
Mechanism
Changing a shared default rebinds every old artifact to a contract it was not trained under; unlike a schema break, a value change produces no error - only degraded, unexplainable behavior. Version-pinned profiles make the binding explicit and permanent.
Applies when
- changing any default in the deployment contract (clock, scales, gains) while old policies remain in use
- writing runbooks that mix policy generations
- a re-tested old policy behaves worse than its era's records
“2026-08-02 起 walk profile 的时钟改为 0.50(WALK_V10_SPEC §3)。v5~v9 全是 0.40 训的,本文件所有命令已改带 --policy-profile legacy_walk_040 ——不要省掉这个 flag,否则是拿慢 25% 的相位时钟静默喂旧策略(分布外,真机危险)。”
train/REAL_SWEEP_V5_V8.md § 1. 预检 ⚠️ 时钟改为 0.50 Low-friction robustness traced to kd DR bandwidth, not friction training - by digging resolved params across 8 lineages, 3840 cells
kd-bandwidth-mu-law-attributionAttribute capability differences by tabulating every lineage's resolved training params and eliminating zero-variance and non-aligned columns first; never let an eval-side override knob serve as the explanation axis, and never write a mechanism into a law before it survives a targeted test.
Symptom
Lineages differed wildly in low-ground-friction survival, and the intuitive explanation - "some trained ground friction, some didn't" - was about to steer the ladder toward a ground-mu training rung.
Context
The attribution ran as a full parameter-vs-result cross: 8 lineages x 4 eval kd levels x 6 mu levels x 20 seeds = 3840 cells, with each lineage's RESOLVED training params dug out and compared item by item. First kill: all 8 lineages had ground mu pinned at (1.0,1.0) - zero variance - so low-mu differences cannot come from friction training at all. The only training parameter aligned with the mu score was kd DR bandwidth: narrow (<=0.24) lineages scored 19.9/19.5/19.5, wide (>=0.40) scored 17.1/15.2/14.6/14.2/12.8 - the two groups completely non-overlapping. Every rival was excluded item by item (kd center no; kp band no; COM small-beneficial non-driving; friction rung a clean double null 19.5->19.5 and 15.2->14.6; iteration count non-monotonic), and the one clean single-variable causal link confirmed it: the s2e-3 kd surgery (0.7,1.3)->(1.08,1.32) moved the score 17.1->19.5. Counter-proof against "each best at its own operating point": the narrow-band lineage evaluated OUT of band (18.2) still beat the wide-band lineage at its own band center (9.2). Two axes were ordered never to be conflated (the first attribution's own error): training kd bandwidth is a parameter axis / lineage property; the eval-side --kd-scale knob is a plant axis (more damping physically helps on slippery floors for ALL policies) - "plant 轴只能当部署缓解,不能当 归因". A tempting mechanism story ("drag vs step attractor") was tested and falsified, and explicitly kept OUT of the law: "机制未定, 不入定律".
Change
The planned ground-mu training rung was recommended closed ("建议 不开") in favor of a kd band-narrowing rung (0.8,1.2)->(0.9,1.1) centered on the deployed value - with a pre-registered risk that the law demands "bandwidth = measured dispersion" and the real robot's kd dispersion was not yet measured; if it exceeds +/-10%, narrowing sacrifices real coverage and the rung must yield.
Outcome
A whole training rung was deleted from the ladder by attribution alone (the second S2 pass dropped mu and push, 5 rungs -> 3); floor material became a deployment-selection input (mu <~0.6 -> deploy the kd1.2 gain profile) rather than a training target.
Mechanism
Cross-lineage performance differences must be attributed over the actual training-parameter table, not over eval knobs or plausible stories: eval knobs act on the plant for every policy (a physical effect), while lineage properties come only from training-time parameters. Zero-variance columns are free eliminations, and one clean single-variable rung is worth more than any correlation.
Applies when
- explaining why lineages differ on a robustness axis
- an eval-side knob (gain scale, power) changes results and invites misattribution
- deciding whether to open a DR rung for an axis never actually varied in training
“8 血统地面 μ 训练带全部钉 (1.0,1.0) 零方差,低 μ 差异与「训没训地面摩擦」无关,是 kd DR 带宽的副产物 … 宽 ≤0.24 → 19.9/19.5/19.5;宽 ≥0.40 → 17.1/15.2/14.6/14.2/12.8, 两组完全不重叠。… 训练 kd 带宽 = 参数轴/血统属性;评测部署 --kd-scale = plant 轴 … plant 轴只能当部署缓解, 不能当归因。… 机制未定, 不入定律。”
train/OMNI_V0_SPEC.md § 4. 地面 μ 鲁棒性 = kd DR 带宽的副产物 (2026-08-08) Model CAN polling skew - joint observations are 6-9 ms stale by read order
can-timing-skew-modelingIf joints are read sequentially over a shared bus, reproduce the per-group observation staleness in sim (or randomize it over the measured range) - synchronous observations are a privileged fiction.
Symptom
Policies trained with synchronous joint observations degrade on hardware where motors are polled sequentially over CAN - hip data is already 6-9 ms old by the time ankle data arrives.
Context
Menlo's core sim2real finding on a leg platform of almost identical mass to Lucen's. CAN is a sequential bus: one poll cycle reads motors in a fixed order, so the observation vector mixes timestamps. Lucen runs 6:6 motors on a dual CAN split, so the problem transfers one-to-one.
Change
Explicitly model CAN timing skew in sim: give the joint groups different observation delays matching physical read order. Menlo went further - running real firmware in the loop with a motor simulator between MuJoCo and firmware injecting 0.4-2 ms uniformly distributed delay.
Outcome
Reported by the reference team as their core sim2real enabler on a same-scale platform; recorded in Lucen's experience log as directly applicable ("这个问题一模一样").
Mechanism
A policy exploits any cross-joint temporal coherence present in training observations; when hardware breaks that coherence per bus position, the learned feedback acts on inconsistent state estimates, injecting phase error exactly at the control bandwidth.
Conflicts
Second-hand episode: outcome numbers are the reference team's report, not a Lucen-run experiment; Lucen adopted the requirement but the corpus has no Lucen A/B of skew-on vs skew-off.
Applies when
- robot polls actuators sequentially over CAN/RS485 or similar shared bus
- sim2real degradation appears as jitter or oscillation not seen in sim
- designing the observation/delay model before a training run
“电机走 CAN 是顺序轮询的,髋部电机的数据到踝部电机上报时已经陈旧了 6-9 ms,他们直接在仿真里显式建模了 CAN 时序偏斜,按读取顺序给三组关节不同的观测延迟。… 注入 0.4–2 ms 的均匀分布延迟。你们是 6:6 双 CAN 分总线,这个问题一模一样”
Experience.md § 执行器 + 时序建模 (line 6)