Training Coach
Doctrine
A report may cite any of these as doctrine-N.
doctrine-1Contract freeze and fingerprint disciplineThe policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.
Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).
Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.
doctrine-2Attribution by resolved training params - never eval-override knobsCapability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.
Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).
Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.
doctrine-3PASS gates become constraints; FAIL gates become objectivesOnce a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.
Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.
Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".
doctrine-4One variable per ladder rung - counted against what the checkpoint sawA rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.
Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.
Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.
doctrine-5Pre-register risks, readings, and stop criteria before the ladderBefore a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.
Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).
Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.
doctrine-6Plant parameters are measured, never inventedEvery plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.
Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).
Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.
doctrine-7Sim2sim gate before sim2real - under deployment conditionsEvery checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).
Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).
Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.
doctrine-8Observation honesty - the actor's inputs are a hardware contractThe actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.
Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).
Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.
doctrine-9Reward economics are audited in realized currencyReward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.
Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).
Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.
doctrine-10The zero-cost option must be the desired behaviorFor every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.
Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).
Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.
doctrine-11Measurement discipline: independent referees, signs, distributionsA disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.
Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).
Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.
doctrine-12The deployment pipeline is plantIrreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.
Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).
Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.
doctrine-13DR budget is finite; its distribution is the measured supportRobustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.
Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).
Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.
doctrine-14Gates measure what hardware feels: posture, margins, stripped assistsAcceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.
Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).
Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.
doctrine-15Fork and root selection: recoverability, maturity, frozen rewardsChoose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.
Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).
Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.
doctrine-16Curricula: verified engagement, lineage counters, disease-phase gatingAutomatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.
Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).
Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.
doctrine-17Probe before training: feasibility first, hypotheses in tablesAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.
Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).
Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).
doctrine-18External advice is recomputed locally; values transfer as ratiosEvery external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.
Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).
Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.
doctrine-19Hardware sessions are scripted experiments, not tuning sessionsReal-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.
Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).
Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.
doctrine-20Close questions in writing; restart when the debt is structuralAudited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.
Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).
Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.
doctrine-21Name the quantity in the space it lives inA goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.
Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).
Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.
doctrine-22Continuation needs a live gradient; a release is chosen by a scanContinue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.
Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).
Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.
Experience cards
170 cards matching “slew-anchor-is-an-integrator”.
Fall recovery was defined as the whole chain - any fallen pose, a stable stand, a clean hand-back to walking - and built as a second policy behind a deploy-side switch, not folded into the walking PPO
recovery-two-policies-and-a-state-machineDefine a recovery skill by the whole chain it must complete, including the hand-back to the next controller; if it is built as a separate policy, make the switching logic and its handoff contract a deliverable of their own, and keep the recovery observation contract a subset of the locomotion one so a unified policy stays possible later.
Symptom
A walking robot that falls needs a human to stand it back up. The design question on 2026-08-09 was whether to teach getting up inside the existing omni walking policy or beside it.
Context
The user set the goal as "any fallen pose -> stand up alone -> stand stably", and the spec named the real difficulty as the full chain fall -> recovery -> stable stand -> correctly initialised walking history and clock -> walking, making the deploy state machine a first-class deliverable. A unified single policy had a real-robot precedent (arXiv:2605.18611, a state-dependent gate near 37 deg tilt) but was deferred until a recovery policy and an omni policy were each reliable. The line ran on its own branch and worktree with every walk/stand/omni/run config path untouched. The development path copied the G1 learned get-up logic (arXiv:2502.12152): first find any feasible get-up (ugly accepted), then add smoothing, torque and real-robot constraints. The recovery contract kept the base 45-dim observation (command slice held at 0, no gait phase, no frame history - their reasons do not apply to a skill without a clock or a velocity task), so it stays a prefix of the 215-dim omni contract and a later merge is not foreclosed.
Change
Two policies and a deploy-side switch instead of one retrained walking policy; recovery got its own minimal contract (45 dims, full-range action, later the beta-anchored profile) and its own acceptance battery.
Outcome
The split held for the whole line: on 08-14 deploy_policy gained a second (PolicyIO, ONNX) pair behind --recovery-policy, each loaded under its own manifest contract, and the runbook runs stand_v1b or omni_c4_ff800 as the locomotion side with recovery_v3_1p1c. The literature scan of 08-10 found that every verified get-up implementation deploys one end-to-end policy (or softly gated experts) and stages only on the training side - so the runtime state machine here is the walk/recovery switch, not a staged get-up.
Mechanism
A separate policy keeps each reward table single-purpose and lets a proven walking lineage stay byte-frozen; the cost moves to the handoff, where every piece of state one policy leaves behind (history, clock, last action, command) must be reset for the other.
Applies when
- adding fall recovery or get-up to a robot that already walks
- choosing between one unified policy and a switched pair of policies
- designing the observation/action contract of a secondary skill
“先做 recovery policy + omni policy 两个策略,部署侧状态机切换;不把 recovery 硬塞进现有 omni PPO。 … 任务定义:**任意跌倒姿态 → 自己站起来 → 稳定站立**。真正的难点不只是"起身", … omni walk**(§6 部署状态机是本 spec 的一等公民,不是附录)”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §0 目标口径与架构判决(用户 2026-08-09 定) Under continuous 3-axis uniform sampling, pure straight-line walking is a zero-measure event the policy never trained
zero-measure-commands-need-mode-samplingEnumerate the exact command points users will actually issue (straight, stop, in-place turn) and give each explicit probability mass via mode sampling with off-axes pinned to zero - never assume a continuous sampler covers its measure-zero subsets.
Symptom
"The robot drifts even in sim when told to walk straight" persisted across reward tunings - because with commands drawn as vx in [0.15,0.5] x vy ~ U(+/-0.2) x wz ~ U(+/-0.6), the event vy=0 AND wz=0 has probability zero: pure straight-line walking was never sampled even once.
Context
Restart evidence item #3: "纯直行是零测度点 … 'sim 里直行就漂'是分布的 必然,不是 reward 没调好" - the drift metric was legitimately drowned by commanded turning (v11's own comment self-documented this). The structural fix is discrete mode sampling: a custom ModeVelocityCommand that first draws a mode by share (stand/forward/back/turn/side/mixed), then draws values only on that mode's axes with all others pinned to exact zero - which is also what preserves single-variable discipline in the C ladder (native 3-axis uniform "采不出'离散模式桶' … 把 C1~C4 的单变量纪律直接毁掉"). The mixed mode later got an ellipsoid constraint rather than a cube for the same reason in reverse - corner combinations of a cube are unrepresentative extremes.
Change
Command generation moved from independent per-axis uniforms to mode-bucket sampling with pinned-zero off-axes (plus 20% rel_standing); acceptance likewise evaluates per mode.
Outcome
Straight-line behavior became a trained, testable mode instead of a measure-zero hope; the C ladder could add one mode per rung with provable isolation.
Mechanism
A policy optimizes expected reward under the command distribution; events of probability zero contribute nothing to the objective, so exact-zero-command behaviors (straight walk, stand, in-place turn) are only learned if the sampler gives them mass. Product-of-uniforms distributions concentrate mass on mixtures and give none to the pure behaviors users actually command.
Applies when
- a "simple" command (straight, stop) underperforms mixtures in sim
- designing command distributions for velocity-tracking tasks
- a ladder needs per-mode isolation for attribution
“纯直行是零测度点:最终 command 为 vx∈[0.15,0.5] × vy∈U(±0.2) × wz∈U(±0.6) 连续均匀,vy=0∧wz=0 从未被专门采样 —— "sim 里直行就漂"是分布的必然,不是 reward 没调好 … Isaac 原生 UniformVelocityCommand 是三轴各自 uniform,采不出"离散模式桶"”
train/OMNI_V0_SPEC.md § 0. 为什么从零 (3) / 三件前置 (1) Before training a one-leg stand, the accounts and a probe showed the default gains could not hold it at all - kp 20 needs 0.39 rad of error to carry the static roll moment, more than the whole adduction range - so per-joint gains came first, and thermal limits set the session length
single-support-gain-authority-probeBefore training a posture that loads one joint statically, compute the steady tracking error load/kp and the series stiffness against m*g*h, and prove with a simple hand-written controller that the posture can be held under the deployment gains - change the gains first if it cannot; then size session length from the thermal account.
Symptom
The one-leg line (standing on one foot, the other folded back, no hopping) had to decide whether the existing gain profile could hold single support before any reward was designed.
Context
Hardware accounts (9.792 kg, COM 0.234 m high, 170 x 80 mm feet, legs 80% of the mass): moving the COM over one foot needs 107 mm of shift and the 20 deg hip-roll adduction range gives 131 mm - geometrically enough. The static frontal moment is 7.8-9 N*m, within RS02's 17 N*m - torque is enough. But at kp 20 carrying 7.8 N*m needs 0.39 rad of tracking error, more than the entire adduction range, and the real robot had already shown it: commanded +0.17, actual -0.04 (0.21 rad droop) under load, 0.0008 rad hanging - load, not the motor. A probe (probe_oneleg.py) then showed open-loop PD cannot hold single support on physics grounds, so the criterion became "an equilibrium exists and a hand-written 4-gain COM feedback can hold it": single-support roll stiffness is hip and ankle in series and must exceed m*g*h_com = 22.5 N*m/rad; ankle kp 12 in series with hip kp 80 gives only 10.4 (open loop 16/16 fell), ankle 60 with hip 80 gives 34.3 (52% margin).
Change
A per-joint gain profile (rl_oneleg: hip_roll kp 80, ankle_roll kp 60, the rest as rl_default) - which needed per-joint gain support in robot.yaml, the bridge, deploy and the trainer's actuator groups - decided before training. Thermal account: single support makes hip_roll the dominant heat load (about 7.8 N*m against a 7 N*m continuous rating), so acceptance and demos run in segments of at most 60 s with a temperature check.
Outcome
Under rl_oneleg the hand-written feedback held six cells cleanly for 6 s (hip_roll steady torque 2.1-3.4 N*m, half the thermal budget); under rl_default the same feedback on the same cells fell 0/4. The trained V0 policy then passed its 40-cell acceptance.
Mechanism
With PD position control, the steady error needed to carry a static load is load/kp; when that error exceeds the joint's range the posture is unreachable whatever the policy does, and series compliance between joints lowers the effective stiffness below the gravity stiffness that single support demands.
Applies when
- single-support, crouched or one-arm-load postures on PD actuators
- a joint "droops" under load on hardware but tracks well when hanging
- deciding whether a new skill needs its own gain profile
“但 kp=20 时撑住 7.8 N·m 需要 **0.39 rad 跟踪误差 > 整个内收行程**。真机已实测: 命令 +0.17 实际 −0.04(droop 0.21 rad),悬挂时 0.0008 rad——是负载不是电机。 … 单支撑滚转是 hip/ankle **串联**刚度,必须 > m·g·h_com = 22.5 N·m/rad;ankle kp12 串 hip80 只有 10.4(开环 16/16 全摔),60 串 80 = 34.3(裕 52%) … **rl_default 同反馈同格 0/4 全摔**(增益档必要性对照)”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §1-1 单脚站: 几何可行,卡点是 hip_roll 增益权限 / §2 A 线增益 / §5 probe 定谳 Changing the gait clock silently flipped a hardwired threshold's meaning - write derived constants as expressions
derived-constants-must-track-their-baseBefore changing any base parameter (clock, control rate, scale), enumerate every constant derived from it and every constant that must NOT change; convert derived literals into expressions of the base so the next change cannot silently flip a term's meaning.
Symptom
Slowing the clock 0.40 -> 0.50 s would have silently inverted the feet_air_time threshold's semantics: the 0.25 s threshold was hardwired, so at ct 0.40 the swing window (~0.20 s) sat below it (constant pressure to lengthen strides), while at ct 0.50 the window (~0.25 s) equals it - the term's meaning flips from "push longer" to "neutral" with no code error anywhere.
Context
The clock change audit walked every dependent quantity: most followed automatically (joint_pos_ref / clearance / contact_number cycle_time params, gait_phase observation, deploy/sim2sim/policy_io, export) - wiring confirmed, zero hand edits; the air_time threshold was the one hardwired constant, fixed by preserving the RATIO: 0.25 -> 0.3125 = 0.625 x ct, with the recommendation to commit it as the expression 0.625*ct "一劳永逸" (solved once and forever). The same audit also listed what must NOT follow the clock (50 Hz control rate, physics dt/decimation, 47-dim contract, action_latency absolute seconds, PD/torque limits) - the change's blast radius stated in both directions.
Change
feet_air_time threshold re-expressed as a fraction of cycle_time; auto-following vs must-not-change lists written into the spec for the clock migration.
Outcome
The clock migration (v10, repeated in v11) carried no silent semantic flips; the expression form removed the trap for every future clock change.
Mechanism
Constants derived from a base parameter encode a ratio at their birth; storing the evaluated number severs the dependency, so changing the base leaves stale semantics with no failing test. Expressions preserve the intent; and an explicit both-directions dependency list (follows / must-not-follow) is what makes a base-parameter change reviewable.
Applies when
- changing gait clock, control frequency, or units
- a reward threshold interacts with a phase/window duration
- config audit finds literals that encode ratios
“feet_air_time 阈值 0.25 是写死的,不跟 ct 走——0.40 时摆动窗 ~0.20s<0.25(恒拉长压力),0.50 时摆动窗 ~0.25s≈阈值(语义翻转)。按比例保原压力:0.25 → 0.3125(=0.625×ct;建议直接写成 0.625 * ct 表达式,一劳永逸)。”
train/WALK_V10_SPEC.md § 3. T —— 慢时钟 (训练侧必做一件) Measure yaw rate by integrating heading, not by averaging body-frame angular velocity - the two differed 15x
heading-integral-not-body-rateFor any secular rate (turn gain, drift), integrate the world-frame angle over the window; never average instantaneous body-frame rates during oscillatory motion - and when code comments warn about a measurement, believe them before re-measuring.
Symptom
Two measurements of the same turn gain disagreed by a factor of ~15: time-averaged body-frame omega_z gave -0.05 while the sim2sim harness's heading-angle integration gave +0.473.
Context
The harness code comment had already documented and predicted the failure: during gait the torso oscillates (body-frame omega_z std up to 0.7); projecting world angular velocity onto a swaying body axis and then averaging biases the estimate systematically - "实测体系均值 −0.04 而实际在以 +0.15 转" (measured body-frame mean -0.04 while actually turning at +0.15). The author's own -0.05 measurement was declared void and the training machine's 1.58/2.45 turn gains confirmed valid.
Change
Measurement doctrine fixed: yaw rate for evaluation = net heading change by integration over the window; instantaneous body-frame rates are unusable for averaged directional statistics during legged gait.
Outcome
Subsequent friction sweeps and turn-gain accounting were all conducted in the heading-integral currency, making cross-simulator comparisons (MuJoCo vs Isaac 1.04/1.02) meaningful.
Mechanism
Averaging a vector quantity expressed in an oscillating frame couples the frame's oscillation into the mean (a rectification bias); the heading integral is computed in the world frame where the gait oscillation integrates to ~zero, leaving the secular component.
Applies when
- measuring turn gain, heading drift, or any secular angular rate
- a body-frame-averaged statistic disagrees with trajectory-level truth
- writing evaluation code for oscillating platforms
“我用体坐标系 ωz 的时间均值测,得 −0.05;sim2sim 用航向角积分,得 +0.473。差 15 倍。… 步态中躯干摇晃(体系 ωz std 可达 0.7),把世界角速度投到摇摆的体轴上再取均值会系统性偏掉 … 结论:偏航率必须用航向积分,体系瞬时角速度取均值不可用。”
train/WALK_DIAGNOSIS.md § ③ 转向增益 —— 我的测法是错的,训练机的 1.58/2.45 成立 Oversized lateral COM randomization (+/-5 cm) deliberately forces leg spread
com-randomization-forces-leg-spreadDR ranges can be behavior-shaping tools, not just robustness padding: oversize a randomization axis to force a strategy the reward struggles to express - and expect a compensating behavior to appear as the cost.
Symptom
Feet drift toward the centerline and even collide; policy has no incentive to keep a lateral support base.
Context
COM randomization ranges were chosen asymmetrically by axis: lateral +/-5 cm ("比常规大,故意的" - larger than usual, on purpose), fore-aft +/-2 cm, vertical +/-2 cm. The oversized lateral range is not robustness padding but a behavioral forcing function. Lucen logged it as directly relevant to its own roll-channel / sideways leg-kick symptom.
Change
Set COM randomization to lateral +/-5 cm, fore-aft +/-2 cm, vertical +/-2 cm, with the lateral band intentionally oversized to make narrow stances fail during training.
Outcome
Effective at separating the feet on the reference robot; side effect - the base began swaying left-right, which then required a foot-centerline distance penalty (see reward-chain-foot-height-landing-spacing).
Mechanism
Randomizing COM laterally makes narrow-stance policies fall for some draws, so PPO discovers wide stances as the only strategy robust across the band - DR used as an implicit reward. The sway side effect appears because the policy hedges against unknown COM by active lateral correction.
Applies when
- feet too close / self-collision in a learned gait
- roll-axis instability suspected to come from narrow stance
- choosing COM or mass-offset DR ranges
“两脚太近甚至互撞 → 先试质心横向随机化 ±5 cm,逼迫策略把脚分开;有效但引发新问题——基座开始左右摇摆 … 横向 ±5 cm(比常规大,故意的,用来逼出分腿)/ 前后 ±2 cm / 垂直 ±2 cm”
Experience.md § 质心随机化范围 (lines 75, 84-86) The walking lines' safety setting, power-scale 0.8, broke the recovery policy's full-range contract - it cut the ends of the joint travel (4/50 could not get up) and left the torque spikes untouched; a kp x 0.9 gain profile inside the trained kp band did the job
power-derating-cuts-full-range-contractA deployment derating knob means something only relative to the action contract: before reusing a line's "safe setting" on a new skill, check what it does to that skill's reachable range and to the term that makes the spikes, prefer a gain change inside the band the policy was randomized over, verify it in simulation, and re-decide when the contract changes.
Symptom
After the violent first real get-up (2026-08-09), the recovery policy needed a gentler setting for its next hardware test, and the walking and omni lines' standard derating - deploying at power-scale 0.8 - was the obvious candidate.
Context
The V0 recovery contract maps actions to absolute targets over the full joint range: a = +/-1 lands exactly on the URDF limits, and standing puts the knee at the clip. Candidates were compared on R3.1 in MuJoCo (5 categories x 10 seeds) on 2026-08-10 before any hardware time was spent.
Change
A new gain profile, rl_kp090 (kp x 0.9, kd unchanged), recorded in robot.yaml as the recovery hardware-test setting, with power-scale 0.8 explicitly banned for recovery.
Outcome
kp x 0.9: 48/50 got up; median torque demand on hip_pitch/knee fell from 120-125% to 100-104% of the deployment limit; leg-leg contact frames 2,152 -> 1,095; the change sits inside the +/-10% kp randomization the policy trained with. power-scale 0.8: 4/50 could not get up, because under the full-range contract it removes the ends of the travel (the deep squat's tucked legs, the straight standing knee), and the torque spikes (kp x error) did not fall at all. When the line moved to the beta-anchored contract, rl_kp090 was declared a V0-era choice that does not fit (beta is calibrated at kp 30) and deployment returned to rl_default; the deploy switch applies power scaling to the walking side only.
Mechanism
A power scale multiplies the action, which under an absolute full-range mapping shrinks the reachable workspace instead of softening the actuator; the spikes come from the proportional term on large errors, which only a gain change reduces - and a gain change inside the trained randomization band stays in distribution.
Conflicts
The undated operator runbook still carries an R3.1 "B comparison" command at power-scale 0.8 beside the rl_default baseline; the sources do not say whether it was written before the ban or was ever run.
Applies when
- reusing a power, torque or action scale from one skill on another
- a policy whose actions map to absolute targets over the full joint range
- choosing a gentler setting for a first or second hardware trial
“kp×0.9 / kd 不动 —— recovery_r3_1 成功 48/50, τ 需求中位 hip_pitch/knee 120~125% -> 100~104% 部署限, 腿-腿接触 2152 -> 1095 帧; ±10% 在训练 kp DR 带内. ⚠️ power-scale 0.8 对 recovery **禁用**: 全 ROM 契约下 0.8 砍的是行程 端点 (深蹲收腿/站直够不到), 实测 4/50 起不来, 且尖峰 (kp·err) 一点不降 —— 它是 walk/omni 的安全档, 不是 recovery 的.”
git:Lucen-recovery@origin/recovery:robot.yaml § gain_profiles 注释: recovery 真机测试安全档 (2026-08-10) / rl_kp090 The shipped checkpoint was chosen by scanning checkpoints on the full gate - neighbours 100 iterations apart failed 1 and 38 cells, late checkpoints degraded - never by taking the last one, and training stopped on signals, not on a schedule
checkpoint-choice-is-a-full-gate-scanChoose a release checkpoint by running the full acceptance battery over a band of checkpoints (including the transfer axis), stop training on measured signals rather than a fixed iteration count, and expect adjacent checkpoints to differ sharply.
Symptom
Gate results moved sharply and non-monotonically between checkpoints of the same run, and the last checkpoint was often not the best.
Context
One-leg V0r1: the 2,000 neighbourhood was best; from 2,500 on the nominal gates degraded (late overtraining); 2,000 itself had one real micro-hop (17.7 mm over 5 frames); 2,300 was all green and shipped. V0r2: failed cells per checkpoint 2,000:19, 2,100:38, 2,200:1, 2,300:3, 2,400:27, 2,500:12, 3,000:18 - 2,200 shipped. The recovery line learned the same from the other side: stopping v2_6 early at a scheduled point left a policy whose re-fall rate had spiked to 9-22% before consolidation healed it ("stop on signals, not on the schedule"), and a continuation's transfer decayed checkpoint by checkpoint while Isaac stayed perfect.
Change
The acceptance rule "scan checkpoints, do not look only at the last one" is written into the one-leg gates (called the S1 discipline); release candidates are chosen from the scan.
Outcome
Both one-leg releases were mid-run checkpoints (2,300 and 2,200) chosen by the full 40-cell battery.
Mechanism
PPO keeps changing the policy after the gates saturate; with no gradient toward the gate's conditions, later checkpoints wander, so gate quality is a noisy function of iteration.
Applies when
- picking which checkpoint of a run to export and stamp
- a run is stopped at a fixed iteration budget
- final-checkpoint results are worse than mid-run smoke tests
“Isaac 侧 S1 纪律: 验收扫 checkpoint,不是只看最后一个。 … 扫描判决: 2000 邻域最优——2500+ 标称面退化(⑤③② 散挂, 晚期过训), 2000 有一例真微跳(L s100 μ1.2, 17.7mm/5帧), 2300 全绿。”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §6 验收门 / §8 核查单 5 与 7 When hardware underperforms, audit deployment knobs before prescribing retraining
deploy-knob-attribution-before-retrainingBefore any "retrain it" decision, reproduce the symptom in sim under the exact deployment configuration; if the symptom follows the deployment knob rather than the checkpoint, fix the knob or randomize it in training - never top-up-train the skill.
Symptom
Real-robot feedback after the C4 deployment - "turning is weak" - with two retraining options on the table: top up turn training, or restart from the s1e root.
Context
The sim account showed the policy turned well (75-81% at pw1.0); the robot was deployed at power-scale 0.8. The 3-6 pp difference between C2 and C4 policies at the same power was noise; the 40-50 pp difference between power levels was the entire effect. Both proposed retraining paths would have burned budget on a non-existent training gap, and restarting from s1e would additionally have discarded the sidewalk skill that took four rungs and a coordinate-bug hunt to obtain.
Change
Decision: retrain nothing. (1) Try pw1.0 on hardware first - sim says net gain; (2) only if 1.0 is unacceptable (heat/feel), the correct training fix is power/torque randomization in the S2 plant line (one variable, fixes turn and backward together) - not skill top-up; (3) restart-from-root explicitly ranked worst.
Outcome
The "weakness" was fully explained by the deployment knob; the sim/real signatures matched the earlier power-derating law verbatim ("与 C2 时代 power 衰减主要伤非前进轴 逐字吻合").
Mechanism
The policy's competence is defined under its training plant; deployment knobs (power scale, teleop mapping, command bands) silently define a different plant. Attributing a deploy-plant effect to a training gap produces exactly the wrong fix - more training on the wrong variable.
Applies when
- real robot underperforms a skill that sim says is fine
- proposals on the table include retraining or re-rooting
- deployment uses any override the trainer never saw (power scale, remapped commands, different control rate)
“正确的训练修法不是补训转向,而是训练时加 power/力矩随机化让策略在 0.8 下自己补偿 —— 单变量,属 S2 plant 线,一次同时修好转向与后退;从 s1e 重训是最差选项:丢掉四轮 + 一个指标 bug 才换来的侧走,而 C2 的转向本来就没问题。”
train/C_LADDER_RUN.md § 3p. 三 处置顺序(回答「补训转向 还是 回 s1e 重训」:都不该) Decompose the offending quantity by channel first - then penalize the failure event, not the joints
penalize-the-slip-not-the-jointBefore penalizing motion to fix a side effect, measure which channels actually carry the offending quantity; prefer penalties conditioned on the failure event that are exactly zero for healthy behavior - and do not medicate behaviors that measurement shows are not sick.
Symptom
Heading drift with support-foot yaw slip (v5: 212-284 deg accumulated over 15 s); the previous v6 draft had attacked it by penalizing lateral joints (a roll 4.0 / yaw 2.0 "home" group) - which collapsed training into the standing basin.
Context
Before choosing the penalty target, the yaw angular momentum was decomposed by joint group with MuJoCo subtree_angmom weighted by real walking joint velocities: pitch-class joints (hip_pitch + knee) carry 95.3%, hip_roll 3.3%, hip_yaw 1.4%. The failed "home" group had been taxing 2.7/step to manage a 4.7% channel. The replacement, feet_yaw_slip (-0.2, |support-foot yaw rate| while in contact), targets the failure event itself and - decisively - costs a non-slipping gait exactly zero, which "横向回家组做不到". The same rung's do-not-do table applied the complementary principle to foot spacing: measured 196-214 mm, stable, no crossing - "没病不吃药" (no disease, no medicine).
Change
Removed joint-usage penalties for the drift problem; added the event-conditional slip penalty (-0.2, realized tax 0.141/step = 12% of tracking) alongside the existing linear-slip term.
Outcome
Turn-gain left/right difference improved 70% -> 19% and heading 185 -> 60.3 deg by v6 without a standing-basin collapse; the 2.7/step lateral tax never returned.
Mechanism
Penalizing joints taxes every use of a channel including healthy use, and if the channel carries little of the offending quantity the tax buys nothing while pushing the optimum toward immobility. An event-conditional penalty (slip while in contact) prices only the failure, leaving the healthy gait's cost surface untouched - and the channel decomposition tells you in advance whether a joint-side fix can even work.
Applies when
- choosing a penalty target for drift/slip/impact problems
- a proposed penalty taxes joints or motions rather than failure events
- a previous joint-penalty attempt collapsed the gait
“pitch 类 (hip_pitch + knee) 占偏航角动量 95.3% … hip_yaw 1.4% … 压 hip_yaw 是管 1.4% 的通道收 2.7/步 的税 —— 上一轮正是这样把策略推进了站立盆地。滑移项不惩罚走路: 不打滑的步态代价为零, 这是横向"回家"组做不到的。”
train/WALK_V6_MINIMAL.md § ① / ② 新增 feet_yaw_slip Narrowing the speed range to stop high-speed falls entrenched crouch-shuffling - judge gait quality at the speed that demands a gait
low-speed-commands-reward-draggingSet command ranges to include speeds that physically demand the target behavior, and evaluate behavior-quality gates at those speeds; when a restriction is added to suppress a failure, check what new optimum it creates at the remaining commands.
Symptom
After the command range was narrowed to (0.15, 0.35) m/s (to treat walk_v1's 134% overspeed and 8.3 s fall at 0.5), the policy settled into crouched foot-dragging; tracking rose monotonically with speed (63% at cmd 0.2, 76% at 0.3, 87% at 0.45), showing low speeds were where the degenerate gait was optimal.
Context
The narrowing advice was the author's own and is retracted in the file: it treated the symptom (falls at speed) while reinforcing the root cause (at 0.15-0.35 m/s, shuffling in a crouch is globally optimal - the Froude number is so low that even humans would not lift their feet). A zero-cost experiment confirmed the flip side: at cmd 0.5 the same policy met BOTH tracking (81%) and clearance (23.0/23.2 mm) standards.
Change
Speed range widened back toward (0.15, 0.5) - upper bound deliberately slightly above the mechanically feasible ~0.44 m/s so the policy finds the boundary itself; acceptance re-pointed: gait-quality criteria (tracking, clearance) judged at 0.45-0.5 m/s, low speed kept only as a survival check.
Outcome
v4 -> v6 progression under the widened range delivered 87% tracking with 34 mm clearance; the "low command = drag" account was confirmed by the monotone tracking-vs-speed curve.
Mechanism
Command distribution is part of the reward: physics prices gaits per speed, and at very low speed the energetic optimum is no swing phase at all. Restricting training to that regime makes the degenerate gait the correct answer to the posed problem - and grading a gait at a speed that does not require stepping measures nothing.
Applies when
- a gait degenerates after a command-range restriction
- quality metrics improve monotonically toward the range boundary
- writing acceptance criteria for gait quality vs survival
“现在看那个建议可能起了反作用:0.15~0.35 m/s 下蹲着蹭就是全局最优,抬腿反而亏。收窄治的是"高速摔倒"的症状,却强化了拖地的病根。… 验收标准里的 cmd 0.2 本身就是拖地速度(Froude 数极低,人在那个速度下也不抬脚)。accept_v2 应把速度跟踪与 clearance 的判定点改到 0.45~0.5 m/s”
train/WALK_DIAGNOSIS.md § ② 放宽速度区间 / ① 零成本实验 Ideal PD is not enough - add a delay buffer and fit armature/friction/delay per joint
actuator-delay-buffer-fittingNever ship ideal PD to hardware: add a measured delay (in control steps) and per-joint armature/friction fitted from step and sine responses, and treat remaining actuator mismatch as your standing largest sim2real residual.
Symptom
Standard ideal PD actuator model transfers poorly; sim assumes targets take effect instantly and joints reach arbitrary acceleration.
Context
A developer with a successful on-hardware Isaac Lab biped modified the actuator model in two ways and calibrated it against the real robot: step-response plus sine-sweep tests (positive step, negative step, sine tracking), overlaying sim curves on measured curves and hand-tuning.
Change
(1) Delay buffer: action targets take effect after a uniform 6 time-step delay on all joints; (2) acceleration limiting so the actuator cannot reach arbitrary acceleration; (3) per-joint fit of armature / friction / delay - different joints genuinely needed different values.
Outcome
Hip joints fit worst, knee best; the developer rated the result "not perfect, the best I could do" and still listed actuator-model improvement as next work - i.e. even the fitted model remained the dominant residual.
Mechanism
Real actuation is a lagged, bandwidth-limited system; a delay buffer and acceleration cap are the two cheapest structures that reproduce its phase and magnitude response. Per-joint differences come from differing load, wiring, and friction states, so a single global constant underfits.
Applies when
- actuator model in sim is ideal PD with no delay
- step-response of real joint visibly lags or overshoots the sim's
- budgeting which sim2real gap to attack first
“标准 ideal PD actuator 不够用,他改了两处:延迟缓冲:目标不是立即生效,全部关节统一 6 个 time step 延迟 / 加速度曲线:执行器不能瞬间达到任意加速度 … 用 armature / friction / delay 三个参数逐关节拟合,标定方法是阶跃响应 + 正弦扫描 … 髋部关节偏差最大,膝关节最好。”
Experience.md § 执行器建模 —— 最值得抄的一条 (lines 50-59) A constant-value plant rung passed every binary gate with record scores - and shipped 60% thinner posture margins that hardware exposed
constant-value-dr-overfits-marginRandomize deployment-critical axes over a narrow band spanning the measured real support - never a single value, never a fictitious tail - and report graded margin quantities (tilt margin) next to binary gates, because saturated gates rank thin-margin and thick-margin policies identically.
Symptom
s2_lag1 (trained at constant 1-frame latency) posted the strongest sim gate sheet in history (20/20 everywhere) yet was unstable on hardware, while s1e (trained across the full 0-3 frame band) was the every-run-stable SOTA at the same power.
Context
The sim autopsy (new --delay-jitter harness modeling the BusWorker's time-varying phase drift): 18 runs across constant and time-varying delays ALL survived - time variation alone does not kill - but the tilt-margin ordering reproduced hardware exactly: s1e 7.7-9.1 deg (thickest) < s2_lag1 10.7-15.0 < s1c 16.2-18.2. Attribution: constant-value training permits precise specialization to that one value; s1e's band diversity forced cross-value robustness - "恒定 1 帧训练 vs s1e 的 0~3 帧全带——分布多样性逼出跨值鲁棒,恒定值允许精确 特化" - so the constant-rung policy's margins were ~60% thinner, fine in sim's clean world, pushed over the line by real-world disturbances. Tool lesson booked: "存活门二值饱和后掩盖裕度差" - binary survival gates saturate and hide margin differences; graded margin columns (tilt-max) belong in the report. The synthesis with the opposite failure (wide tails cause drag-glide): the proposed resolution was a NARROW uniform band (0.02, 0.04) covering exactly the real 1-2 ticks - diversity inside the measured support, no tail, no single point. s1e's root selection later leaned on the same property: its full-band latency training "预装" the delay rungs and delivered "全工况稳定裕度" that survived power derating.
Change
DR-on-an-axis design refined to a three-way distinction: no wide fictitious tails (drag), no single constant values (thin margins), but a narrow band spanning the measured real support; acceptance reports gained graded margin columns alongside binary gates.
Outcome
The tilt-margin column entered the standard report; the s1e root (band-trained) carried the C ladder while the constant-value branch was archived with its three contributions credited.
Mechanism
Robustness margins are shaped by the diversity of the training distribution, not just its support: a point-mass distribution lets the optimizer trade margin for on-point performance, while a band forces solutions that keep margin across the band - and binary survival metrics cannot see the difference until the margin is spent on hardware.
Conflicts
The narrow-band (0.02,0.04) resolution was a pending recommendation ("裁决建议(待用户)") at the time of writing; the lineage instead moved root to s1e whose full-band training predated the staged ladder - the deterministic-staging card and this card record the two failure modes the final design must avoid simultaneously.
Applies when
- a rung trained at a fixed plant value aces sim but wobbles on hardware
- binary acceptance gates are all saturated across candidates
- choosing between constant, banded, and wide DR on one axis
“18 跑全活,时变性单独不足以击杀;但 tilt_max 裕度排序完整复现真机:s1e 7.7~9.1°(最厚)< s2_lag1 10.7~15.0 … 恒定 1 帧训练 vs s1e 的 0~3 帧全带——分布多样性逼出跨值鲁棒,恒定值允许精确特化 … 存活门二值饱和后掩盖裕度差(s2_lag1 sim 门 20/20 史上最强却真机不稳)”
train/README.md § s2_lag1 真机不稳 × s1e 稳的 sim 对拍(2026-08-07,时变延迟实验) Push-test protocol - positive side first, fragile side spotted, axes aligned in the log, and cross-machine push counts stay qualitative
push-test-chirality-protocolOrder disturbance tests from the robust side to the fragile side with protection scaled to sim-measured asymmetry, align and log frame conventions before testing, and treat cross-domain disturbance counts as qualitative evidence only.
Symptom
Hand-push testing on hardware risked falls on a side sim had already flagged as fragile, and push counts invited apples-to-oranges comparison with sim numbers.
Context
Sim chirality was explicit: descendants were far more fragile in -y (fric-3000@kd1.2: +6 N*s survived 15/20 vs -6 N*s only 3-9/20) while the s1e control was perfectly symmetric (40/40). The protocol therefore: push the positive direction first, keep a spotter for the negative side; before any push, record which real-robot side corresponds to sim's +y in the log ("上机前对一次坐标"); and - citing the chaos lesson ("混沌课文") - real push results are used only as qualitative corroboration, never compared numerically with sim survival counts across machines.
Change
Push testing became a scripted, chirality-aware protocol with frame alignment as a logged precondition and an explicit epistemic limit on cross-domain count comparison.
Outcome
The fragile side was tested with protection informed by sim's quantified asymmetry; logs stayed interpretable because the frame correspondence was recorded before the first push.
Mechanism
Disturbance-response chirality is a real, quantifiable lineage property, so test order should follow measured fragility; and perturbation outcomes are chaotic in the details (divergent trajectories from tiny differences), so counts do not transfer across domains even when qualitative rankings do.
Applies when
- planning push/disturbance tests on hardware
- sim shows directional asymmetry in disturbance survival
- someone proposes comparing real push counts to sim counts
“先正向后负向, 负向留人扶 —— sim 手性明确: 后代在负 y 向显著更脆 (fric-3000 @kd1.2: +6 N·s 15/20 vs −6 N·s 3~9/20), 而 s1e@0.8 两向 40/40 完全对称。上机前对一次坐标 … 跨机不做二值结论 (混沌课文): 真机推力只作定性对照, 不与 sim 计数对比。”
train/REAL_RUN_S2.md § 3. 抗推 (可选, 人手推; 做则按此协议) The fallen-state reset was designed, not sampled from SO(3) - fixed category shares with jitter, a low drop that settles physically, equal left/right shares for mirror augmentation, and a numeric check before training
fallen-pose-reset-distributionBuild a fallen-start distribution from named, physically plausible categories with jitter and a settle phase, keep mirrored categories at equal probability, and check the realized shares and penetration numerically before spending a training run on it.
Symptom
A get-up policy can only learn from the fallen states its resets produce; uniformly random orientations produce ground-penetrating and limit-jammed states the robot can never be in.
Context
R0 reset_root_fallen: supine 30%, prone 30%, side_l 15%, side_r 15%, mid (random axis 50-125 deg) 10%, +/-15 deg jitter, full yaw, dropped from 0.28-0.40 m and left to settle under physics, joints uniform inside the soft limits with a 5% margin plus small random velocities. Random SO(3) was rejected (the advisor agreed). side_l and side_r must have equal probability because mirror augmentation turns a left fall into a right fall. The advisor had proposed supine and prone only for R0; the spec included side and mid because the feasibility accounts showed physical solutions for all of them, and wrote "narrow back to supine+prone" down as the first fallback. With no display on the training box the reset was checked numerically instead of by eye.
Change
Category mix as above; realized shares, settle height and penetration measured over 512 envs before the first run. A fallen-state bank (real falls, settled and stored) was pre-registered for R2.
Outcome
Realized shares 29.3/31.6/16.4/17.8% against the config, settle +0.262 m, final penetration 0/512 (a 0.10 m peak at the write instant, ankle links only, pushed out within 80 ms because the 0.28 m drop floor is shorter than a fully extended leg). R0's failure was a reward basin, not a reset artifact. The fallen-state bank stayed unbuilt through V3.1 (checklist item open); R0.3 later re-sliced the prone share into roll_l/roll_r bands, which is what forced the acceptance distribution to be frozen separately.
Mechanism
A category-structured, physically settled start distribution keeps training on states the robot can actually occupy, and equal mirrored shares keep mirror augmentation a pure doubling of data rather than a bias.
Applies when
- designing reset distributions for get-up, recovery or multi-contact skills
- mirror/symmetry augmentation is on and the task has chiral start states
- no viewport is available to inspect resets on the training machine
“角度 jitter ±15°、yaw 全域、0.28~0.40 m 低空放下由物理沉降,关节软限位内 均匀(留 5% 余量)+ 小随机速度。**不用 random SO(3)**(会采出穿地/极限卡死 等现实不可能状态,顾问同判) … side_l/side_r **概率必须相等**(镜像增强的样本同分布前提)”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §3 R0 任务定义 / §9 核查单 Push DR helped one lineage and hurt another at the same dose - robustness budget is conserved and gets borrowed, not created
push-dr-conditional-budget-conservationBefore opening a disturbance-DR rung, measure whether the untrained policy already meets the spec; if training it anyway, expect the benefit to be conditional on the lineage's existing DR load, grade the intensity, and audit retained margins - budget spent elsewhere will be borrowed back.
Symptom
The push rung's outcome flipped with the lineage: direct +/-0.6 m/s push failed outright on first attempt (base walking collapsed - kd1.2 scan 0/3 from iter 3300, sim2sim self-falls with pushes OFF - no PASS point existed); staged +/-0.3 then gave the narrow-kd single-working-point lineage real gains (push survival 1/5 -> 4/5) while the SAME dose made the dual-working-point balanced-band lineage WORSE (20-seed survival 18 -> 12/20 plus across-the-board push regression).
Context
The four-ladder verdict ("四梯定案", s2e/s2f at both intensities) named the pattern: "push DR 收益条件性" - the benefit is conditional on how much robustness budget the lineage has already spent. The law candidate: "DR 总预算守恒, 平衡带鲁棒性从抗扰余量借" - total DR budget is conserved; a lineage already covering a wide plant band pays for push tolerance out of its disturbance margin. Both S2 ladders therefore closed at the friction rung, with the decisive numerator: untrained push tolerance already met the 4-6 N*s requirement, so the rung was not needed at all ("⑥ push 不训(收益条件性,免训 ±0.6 已达 标)"). The same accounting later justified the C-before-S2 ordering ("push/μ 两轮已实证 DR 预算有限且会被重分配") and trimmed the second S2 pass to three rungs.
Change
Push removed from the standing ladder; graded intensity retained as the method IF a lineage ever needs push training; "does the untrained policy already meet the disturbance spec" instituted as the first check before opening any disturbance rung.
Outcome
Two rungs (push, ground mu) deleted from the second S2 pass on measured grounds; the ladder's real yield was re-stated honestly as precision, not robustness (speed gate 0 -> 20/20, zero-command drift 0.98 -> 0.06 m, but push 159 -> 125/160).
Mechanism
A fixed-capacity policy allocates representation and margin across the training distribution; adding a disturbance axis to a lineage that already spans a wide plant family forces reallocation - the new tolerance is bought with existing margins. Lineages with narrow plant coverage have free budget, so the identical DR dose lands as gain. Benefit is a property of (dose x lineage state), never of the dose alone.
Applies when
- proposing push/perturbation training on a hardened lineage
- the same DR rung helped one lineage and hurt another
- accounting where a ladder's robustness gains actually came from
“push DR 收益条件性 —— s2e⑥a (单工作点血统 kd 窄带) ±0.3 得抗推 1/5→4/5; s2f⑥ (双工作点平衡带血统) 同档反而 20-seed 存活 18→12/20 且抗推全面倒退。规律候选: DR 总预算守恒, 平衡带鲁棒性从抗扰余量借。两阶梯均以 ⑤ 摩擦级收官 … 抗推 4~6 N·s 免训已达标。”
train/OMNI_V0_SPEC.md § 4. ⑥ push 四梯定案 (2026-08-07) Four in-lineage attempts to widen the standing stance failed - remove a tax, add a joint-space knife, change the target, add a task-space metric penalty - because the stance was the end state of the get-up path; trained from scratch with the right terms it grew right from day one
stance-decided-by-get-up-pathA posture a skill ends in is shaped by the path the policy takes to reach it; if several single-variable edits to the terminal-phase reward cannot move it, stop editing that phase and retrain with the terminal constraint present from the start.
Symptom
v2_6c stood with its feet 0.159 m apart (task-space) and its hips yawed 45-47 deg the same way, which split on the real robot. Standing-phase reward edits did not move it.
Context
V2.7-A removed the flat-feet tax on compensated stances (stance unchanged); V2.7b added a hip-roll lower-bound hinge (+5 deg in 3,000 iterations, yaw ratchet); V2.8 changed the stand_pose target to a wide flat stance (stance unchanged, yaw not unwound, feet nearly overlapping, mu 0.4 transfer 2%); V2.9 penalized lateral spacing in metres (the policy parked just outside the penalty's gate in a lunge, 0% success). The v2_6c get-up goes through a split and closes the feet together as it rises.
Change
In-lineage stance surgery was formally closed. V3.1 trained from scratch with task-space stance terms present from the first iteration (and, after P1, a positive width band instead of a penalty).
Outcome
V3.1 P1b: lateral stance 0.364 m, foot tilt 0.0 deg, all four categories 100%, MuJoCo mu 1.0 and 0.4 both 100% - with a symmetric toe-out the kinematic audit had not enumerated. P1c (with a yaw guard): 0.355 m, all six acceptance criteria passing, mu 1.0-0.4 all 100%; it became the product.
Mechanism
A converged policy does not rebuild the path that produced its terminal posture; a standing-phase gradient only finds the nearest hack around the posture the get-up delivers.
Applies when
- the final posture of a transition skill is wrong and resists terminal-phase shaping
- repeated continuation rungs produce hacks instead of the intended posture
- deciding between another in-lineage fix and a from-scratch retrain
“窄站距 + yaw 扭是 v2_6c 起身策略(劈叉起身 → 双脚并拢收势)的**结构性 终态**,不是站立段的孤立参数 —— 站立形态由起身路径决定,在血统内只动 站立段奖励改不动它。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 结果:V2.8 判 FAIL —— 血统内站姿手术第三次证伪 Knee swing collapsed because it directly trades against the slip penalty - price the conflict explicitly and clamp what reward cannot hold
knee-swing-vs-slip-pricingWhen a behavior collapses as another metric improves, look for the term pair trading them and set their price ratio deliberately (with escalation and reverse tripwires pre-registered); where the policy actively spends action budget to undo your target, stop paying more reward and clamp the target space structurally.
Symptom
Knee peak-to-peak swing collapsed across generations - v5 33 deg, v10 26-30, v10b 7-8, v11 6.5-8.6 - and rolling the clock back did not recover it, acquitting the clock; the collapse tracked the gated slip penalty instead: "屈膝与不打滑在当前奖励里直接对抗" - v10b's excellent 93 deg slip was purchased with knee amplitude.
Context
Reward-side flexion fixes had failed three times: raising reference amplitude backfired twice (v9/v11), and v11's deep-squat default was actively fought by the policy - it spent 0.68 of action budget pulling the squat straight ("被策略花 0.68 动作拉直反杀"). v12's design accepted the conflict as real and attacked on two tracks: (1) ECONOMICS - a direct knee_swing_amplitude reward (+0.3, target 0.55 rad, capped at 0.6/step = 55% of tracking), explicitly opposed to the slip penalty by design ("显式对立——这正是设计:v12 就是这场对抗的定价实验"), with an escalation ladder (K +0.3 -> +0.5, then slip -0.5 -> -0.3, one layer at a time) and a reverse tripwire (slip telemetry back at v10 levels -> slip weight to -0.8, accept ~20 deg knee compromise); (2) STRUCTURE - knee target bounds [0.2, 0.9] rad so full straightening is physically impossible (straightest 11.5 deg) and the 0.68 fighting budget is released. A bonus falsifiable prediction was attached: phase-lock strength tracks amplitude (v9_probe 48 deg locked 2.5 Hz; v11 low-amplitude 1.36 Hz unlocked), so if K works, hardware phase-lock should return - one change, two verdicts.
Change
knee_swing_amplitude reward + knee target clamp + pre-registered escalation/reverse levers; the failed reward-side-only approach retired.
Outcome
The lineage was frozen before v12 trained (strategic reset), but the diagnosis stands as the walk line's clearest example of two reward terms trading a behavior between them, with the pricing experiment and structural clamp fully designed and calibrated.
Mechanism
When two terms price opposite aspects of one motion (swing amplitude creates yaw momentum that becomes slip), the optimizer settles wherever the price ratio puts it - patching one side moves the equilibrium, not the conflict. Explicit pricing makes the trade a designed quantity; structural clamps remove the regions where the policy spends budget fighting the designer.
Conflicts
The pricing experiment (K vs slip) was designed and calibrated but never trained - the 2026-08-05 reset suspended v12; the collapse attribution table and the 0.68-action counterattack are measured, the remedy's效果 is untested.
Applies when
- one gait quality degrades in lockstep with another's improvement
- the policy visibly fights a default pose or reference
- repeated reward-side fixes for the same behavior have failed
“膝摆塌在 v10→v10b,头号嫌疑是门控滑移罚(四代实测膝 p2p:v5 33° / v10 26~30° / v10b 7~8° / v11 6.5~8.6°;退时钟没救回 → 非时钟)——"屈膝"与"不打滑"在当前奖励里直接对抗 … 奖励侧修屈膝已三败 … v11 深蹲 default 被策略花 0.68 动作拉直反杀”
train/WALK_V12_SPEC.md § 0. 定位 / 2. K —— 膝摆经济(与滑移罚的对偶) Bracket a real-robot A/B with a repeated reference run - battery drain is the confound
battery-bracketed-real-abOrder hardware A/B sessions as A-B-A: repeat the first condition at the end, and void the comparison if the bracket runs disagree - never let battery or venue drift ride on the second condition.
Symptom
In a two-policy teleop A/B on hardware, the second policy is measured on a lower battery voltage than the first - a systematic bias that would be read as a policy difference.
Context
C2 real A/B (checkpoint 700 vs A800, same floor, same day) was scripted as 700 -> A800 -> 700-rerun, with the explicit note that a teleop session drains the pack and the trailing policy "naturally suffers".
Change
Protocol: run the reference policy first AND last; if the two reference runs differ noticeably, declare the whole session battery/floor-polluted and void the A/B ("结论作废重来"). Also log electricity per run.
Outcome
Called out as the round's only systematic confound, closed by one extra command ("这是本轮唯一的系统性混淆源,一条命令就能堵掉").
Mechanism
Battery voltage scales available torque, and torque loss hits behavior asymmetrically (see power-scale-hurts-nonforward-axes), so drain masquerades as policy regression; a head/tail reference pair converts the unobserved drift into a measured control.
Applies when
- comparing two policies or settings on hardware in one session
- any sequential hardware evaluation where the plant drifts (battery, temperature, floor wear)
“为什么要 700 复跑:一次遥控 session 下来电池会掉压,第二枚天然吃亏。头尾各跑一次 700,若两次 700 明显不同,说明这轮 A/B 被电量污染,结论作废重来。这是本轮唯一的系统性混淆源,一条命令就能堵掉。”
train/C_LADDER_RUN.md § 3c. A-3 真机 A/B(同一段地板、同一天、电量记账) When the eval proxy has a known systematic bias, gate on within-proxy differences, not absolutes
relative-metrics-survive-proxy-biasWhere the evaluator is known-biased, design gates as within-evaluator contrasts (left vs right, A vs B, pre vs post) that cancel common-mode error; reserve absolute thresholds for quantities whose proxy calibration has been checked.
Symptom
MuJoCo systematically overestimated yaw turn gain (1.58/2.45 where Isaac measured 1.04/1.02), and friction alignment recovered only part of the gap (mu 1.0 -> 0.6 pulled it to 1.43/2.27) - so any absolute turn-gain acceptance threshold in MuJoCo would be judging the proxy's bias, not the policy.
Context
The acceptance criterion was rewritten to use only the left/right difference of the turn gain (standard: <20%): both directions pass through the same biased proxy, so the bias largely cancels in the difference while the policy's chirality - the thing being gated - survives. The absolute-gain row was dropped: "转向只判左右差,不判绝对值". A parallel task was still opened to align MuJoCo contact parameters to Isaac's (mu, restitution), since drift proved highly friction-sensitive (straight-line drift -65 deg at mu 1.0 vs +1 deg at mu 0.4) - bias reduction and bias-robust metrics proceeded together.
Change
Gate metric changed from absolute turn gain to left-right gain difference; proxy-alignment work scheduled separately rather than blocking acceptance.
Outcome
Turn acceptance became meaningful across proxy versions (v5 70% -> v6 19% difference measured the real improvement) while the absolute bias question was pursued without holding the ladder hostage.
Mechanism
A systematic multiplicative or additive proxy bias applies to both arms of a mirrored measurement; differencing (or ratioing) mirrored conditions cancels the common-mode bias to first order, leaving the asymmetry signal. Metrics built this way remain valid while the proxy is imperfect - which it always is somewhere.
Applies when
- a sim proxy disagrees with the trainer or hardware in absolute terms
- writing acceptance thresholds for direction-paired skills
- proxy calibration work would otherwise block a ladder
“转向只判左右差,不判绝对值 —— MuJoCo 的偏航增益系统性高估(实测 1.58/2.45 vs Isaac 1.04/1.02),摩擦只能解释一部分(μ 1.0→0.6 仅拉到 1.43/2.27)。差值是相对量。”
train/WALK_V5_SPEC.md § 6. 验收 Continuing a converged policy on a change that carried no new gradient drifted its transfer from 100/98% to 80/28% over 3,000 iterations while every Isaac gate stayed perfect - scan every checkpoint on the second simulator's friction axis
converged-continuation-is-poisonBefore continuing a converged policy, check that the change creates a live gradient; if it does not, cap the budget at a few hundred iterations, and in every continuation scan each checkpoint on the second simulator's transfer axis (for example low friction) - trainer-side gates can stay perfect while transfer decays.
Symptom
V2.7-A (swap the flat_feet term for a compensated version, continue from v2_6c) finished with the line's best Isaac score (100%) and a MuJoCo transfer collapse: mu 1.0 98 -> 80%, mu 0.4 98 -> 28%; the stance it was meant to widen had not moved.
Context
The new term's calibration run showed a near-zero tax from the start: the policy already satisfied it, so the reward landscape offered nothing new. A checkpoint scan on MuJoCo mu {1.0, 0.4} located the damage: +100 iterations 100/98% (better than the baseline), then 86/54, 60/38, 80/28 - monotonic decay with training length, while entropy and action noise rose (7.77 -> 8.18, 0.588 -> 0.612): drift, not sharpening.
Change
Rule written in: with no new gradient, a continuation budget is short (at most a few hundred iterations) and the MuJoCo transfer axis enters every checkpoint scan. The next rung (V2.7b, a live stance-width gradient) was budgeted at 1,000 iterations with mu {1.0, 0.4} scans every 100 and a stop-on-signal rule.
Outcome
V2.7b kept transfer at the same depth (mu 1.0 98% / mu 0.4 92% at +1,000, where A had already rotted to 86/54) and at +3,000 (100/96%): a live gradient preserved transfer. V2.8 then broke that pattern (mu 0.4 2%): the gradient must also be compatible with the policy's existing form.
Mechanism
On a converged reward landscape PPO keeps updating without a signal to follow, and the random walk is pulled toward whatever the training plant rewards idiosyncratically - invisible in the trainer's own gates.
Conflicts
The drift mechanism is the spec's reading of one decay series plus one contrasting run; V2.8 is recorded as an exception to "live gradient keeps transfer".
Applies when
- fine-tuning a converged policy with a small reward change
- a continuation run's trainer-side metrics improve while real or cross-sim results worsen
- choosing which checkpoint of a continuation to ship
“**checkpoint 扫定死因**(μ1.0/μ0.4):**29500(+100 iter)= 100/98%** (优于基线!)→ 30400 = 86/54 → 31400 = 60/38 → 32398 = 80/28 —— **迁移随续训长度单调衰减**。 … **教训入库:收敛均衡上的长续训是毒药 —— 无新梯度时 续训预算须短(≲数百 iter),且 MuJoCo 迁移轴必须进 checkpoint 扫描。**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §43 结果:V2.7-A 判 FAIL —— 换刀本身无罪,毒在续训预算 A curriculum ramp keyed to the process step counter re-fires on every resume - and shipped policies never saw the penalty
curriculum-counter-lineage-stepsKey every curriculum/ramp schedule to lineage-cumulative progress, not per-process counters; and audit where your shipped checkpoints sit relative to every ramp - a penalty that no product ever experienced is not part of your training.
Symptom
vx+0.30 died at a fixed relative time in every resumed run: resume at 500 -> slide at 1100, zero at 1300; resume at 700 -> slide at 1300, zero at 1500 - absolute depths offset by exactly the resume offset, relative timetable identical.
Context
ramp_reward_weight (the saturation penalty ramp) read env.common_step_counter, which restarts at 0 for every run including --resume. So start_step=600 meant "600 iters after THIS resume", not "lineage iteration 600". The A/B arm pair was the clean proof: their env.yaml differed only in log_dir, only resume point distinguished them, and the omni CurriculumManager had exactly one active term - nothing else could produce that timetable. Second consequence: every shipped checkpoint (s1e-500 at +500, C2-700 at +200, A800 at +100) was selected before its run's +600, so the saturation penalty weight was 0.000 for every product ever shipped - explaining saturation 33% and raw |action| 1.9 against clip 1.0 (hip_roll in bang-bang), i.e. half the heat budget.
Change
Two independent recommendations recorded: (1) make the ramp count lineage steps (add the checkpoint's iteration offset on resume) or pin terminal weights in downstream rungs instead of ramping; (2) give the saturation penalty its own rung - never mixed into a skill-learning rung (that would be two variables again).
Outcome
Explained the recurring +600 death of the highest-amplitude command and the persistent actuator saturation of all shipped products with one root cause; honest caveat booked (at +800 the weight is only -0.086, small, but vx+0.30 is the command demanding the largest action amplitude, so it is squeezed first).
Mechanism
Resumable training splits "the lineage" from "the process"; any schedule keyed to process-local counters silently re-applies its transient to every descendant run, and any product-selection habit that picks checkpoints early systematically samples the pre-ramp regime - the curriculum exists in the config but never in any shipped policy.
Applies when
- resumed/forked training with any scheduled reward or DR ramp
- a metric dies at a fixed offset after each resume
- shipped policies show behavior a late-schedule penalty should prevent
“ramp_reward_weight 读的是 env.common_step_counter,它每个 run 从 0 开始,--resume 也不例外。… 原始 run / 臂B | 500 | 1100 = +600 | 1300 = +800;臂A | 700 | 1300 = +600 | 1500 = +800 … 所有出品其实从没见过饱和罚。… 这解释了 sat_max_pct 33%、raw |a| 最大 1.9(clip 是 1.0)—— hip_roll 一直在 bang-bang,而罚它的那一项权重恒 0。热账的一半在这里。”
train/C_LADDER_RUN.md § 3g. 系统性问题:saturation_ramp 每次 resume 归零 Training-log reward values and fixed-command eval values live on different distributions - comparing them once claimed a 44% improvement that was really 6-10%
same-distribution-reward-comparisonQuote reward-term values only with their distribution attached (command range, DR on/off, environment), and compare across runs only when those match; re-measure in a common environment before claiming any improvement percentage.
Symptom
A v6-era analysis concluded slip had dropped 44% by comparing the training log's Episode_Reward against values calibrated in a fixed-command play environment; a same-condition re-measurement showed the true improvement was 6-10%.
Context
The training log's reward is an expectation over the training command distribution (vx 0.15-0.5, yaw +/-0.6, with pushes and domain randomization); play-environment calibrations are taken at a single fixed command with DR off. Subtracting one from the other compares apples to oranges - the warning was written into the v7 pre-flight: "奖励数值只能在同一指令分布下比较 … 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次)".
Change
Rule adopted: any before/after reward-term comparison must hold the command distribution, DR state, and evaluation environment fixed; training-log values compare only against training-log values of runs with identical command/DR configs.
Outcome
The phantom 44% improvement was retracted; later term-level accounting (e.g. the C4 ignore-floor work) consistently specified its distribution before quoting numbers.
Mechanism
A reward term's expectation depends on the visited-state distribution as much as on the policy; changing the command distribution or DR moves every term's baseline. Cross-distribution differences therefore measure the distributions, not the policy change.
Applies when
- comparing reward telemetry across training runs or vs play evals
- claiming improvement percentages from training logs
- term-level reward accounting for diagnosis
“奖励数值只能在同一指令分布下比较。训练日志的 Episode_Reward 是在训练指令分布上算的(vx 0.15~0.5 / 偏航 ±0.6 / 带推力与域随机化), 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次: 据此以为滑移降了 44%, 同条件对拍只有 6~10%)。”
train/WALK_V7_SPEC.md § 3. 开训自查 ⚠️ Ungated phase shaping made standing 42x more expensive than stepping - and the stepping was cooking the hip motors
moving-gate-42x-stand-taxGate every phase/clock-driven shaping term on the command that justifies motion, price the cmd=0 case explicitly during design - and when a reward flaw is sim-only, book it with trigger conditions instead of operating immediately on a working lineage.
Symptom
At cmd = 0 the sim policy never stood still - it stepped in place and crept 0.98 m per 20 s; on the real robot the same lineage's stepping made hip_roll motors run 20 degC hotter than every other joint (43-48 degC vs 25-28).
Context
The arithmetic closed it: the three phase-shaping terms (joint_pos_ref 1.6, feet_contact_number 1.2, feet_clearance_swing 1.6) are driven by the gait clock with NO command gating, so standing at cmd=0 forfeits 3.32/step of shaping while honest standing earns only 0.078 of tracking - stepping wins 42x. The heat chain: perpetual stepping = perpetual single support = one hip_roll stalled at ~4.3 N*m (25% of torque limit) carrying the torso's frontal-plane moment - two hip_rolls = 90% of whole-machine steady-state I2R; measured stand-vs-step comparison: total heat -71% when actually standing. The suspended test acquitted the actuator (0.21 rad sag -> 0.0008 rad in air) and mechanics vetoed the easy fix ("降 kp 救不了热" - equilibrium torque equals the external load regardless of kp). The fix (moving_gate: hard-gate the three shaping terms on |cmd| > eps) was designed - then DEFERRED by the user because the real robot at the time stood fine: "真机不表现该问题, 为真机不存在的病改奖励表不划算", with written trigger conditions (C-ladder stand row persistently failing, or real robot starting to step/drift at cmd=0) and the known hazard tag (this is exactly the reward-change class that triggers the B-arm signature). When the real robot later DID step and cook, the booked trigger fired and moving_gate moved from debt to to-do with its benefit re-priced: "停止空烧 hip_roll,稳态发热降 ~71%".
Change
moving_gate designed with the gate_by_cmd convention; deferral, triggers, and expected heat recovery all pre-registered instead of patching the reward for a then-sim-only symptom.
Outcome
The cmd=0 stepping went from mystery to closed arithmetic; the thermal measurement (hip_roll +20 degC) quantitatively confirmed the 90%-of-heat prediction; the reward change waited for real-world justification instead of spending a risky revision early.
Mechanism
Clock-driven shaping terms define a perpetual-motion bounty unless gated by command; the resulting idle gait is not a training bug but the table's optimum. Its cost surfaces on hardware as stall-torque heating set by statics (mass x lateral offset), which no gain change can remove - only removing the motion (gating) or widening the stance can.
Applies when
- the policy steps in place or creeps at zero command
- specific joints run hot in idle behaviors
- deciding when a known reward flaw justifies a risky mid-lineage fix
“塑形合计 −3.32/步 … 净: 站定亏 42 倍 … 两颗 hip_roll 4.29 / 4.09 N·m(各占限扭 25%),占全机稳态 I²R 的 90% … 吊挂实测 hip_roll 跟踪误差 0.21 rad → 0.0008 rad … 降 kp 救不了热 … 真机不表现该问题, 为真机不存在的病改奖励表不划算”
train/README.md § C1 FAIL 节 (cmd=0 的 sim/真机分歧记账) / C2 真机 A/B 三b 发热定性 A soft joint-limit penalty charged the standing pose itself - the geometric-zero knee sat on its hard limit, so stand_v1 bent its knees to dodge 0.419 per step and leaned 4.1 deg forward; excluding the knee gave 0.24 deg
soft-limit-penalty-charges-nominal-poseBefore training, evaluate every penalty at the nominal pose; if a joint's soft limit sits inside the pose the task requires (a straight knee on its hard stop), exclude that joint from the soft-limit penalty and let the action clip enforce the hard limit.
Symptom
stand_v1 (retrained after the default pose moved to the CAD geometric zero and mirror augmentation was added) fixed left/right asymmetry (6.8 -> 0.0 deg) but settled at a 4.1 deg forward lean, where pure PD at the same default settled at 0.1 deg - the policy was actively pushing itself forward, which is exactly the real robot's failure direction.
Context
soft_joint_pos_limit_factor = 0.9 shrank the knee's soft limit to -/+0.1047 rad, while the geometric-zero default has the knee at q = 0, exactly on the hard limit. Standing in the nominal pose therefore paid 0.2094 x 2.0 = 0.419 per step in dof_pos_limits (alive earned only 0.5). The policy's way out was to bend the knees to -/+0.1013 rad, and the cost was the forward lean.
Change
stand_v1b: the knees excluded from dof_pos_limits (a straight knee IS the standing pose; the hard limit is still enforced by the action clip). No other change.
Outcome
stand_v1b: max tilt 0.3 deg, steady tilt 0.24 deg, asymmetry 0.1 deg, height 0.384 m - exactly nominal - with knees at -0.0007 / +0.0005 rad. It became the standing release used on the real robot, and later the standing side of the recovery switch. The same exclusion was carried into the recovery contract (knee at the clip in the standing pose) and the one-leg reward table.
Mechanism
A limit penalty whose soft boundary lies inside the nominal pose turns the nominal into a taxed state, and the policy buys its way out with whatever posture change is cheapest - here a knee bend paid for with lean.
Applies when
- a standing or default pose has a joint at or near its hard limit
- a policy settles in a small steady tilt that pure PD does not show
- soft-limit factors shrink limits uniformly across joints
“`soft_joint_pos_limit_factor=0.9` 把膝软限位内缩到 ∓0.1047,而几何零位 default **膝盖 q=0 正好压在硬限位上** ⇒ 站在标称姿态每步白扣 `0.2094 × 2.0 = 0.419` (alive 才 +0.5)。策略只能屈膝到 ∓0.1013 躲罚,代价是躯干前倾 —— 恰好是真机的 失效方向。 … 修掉"软限位罚标称姿态"后重训(`dof_pos_limits` 排除膝盖)。**前倾问题彻底消失** … 不再屈膝躲惩罚,高度正好落回标称 0.3840。”
train/README.md § 三期: 镜像对称增强 + 站立 v1 (2026-07-28) / stand_v1b (2026-07-28): 站立定版 Privileged signals (true velocity, foot force, foot height) go to the critic only
observation-honesty-critic-onlyTreat the actor observation vector as a hardware contract: every element must exist on the real robot with realistic noise; privileged simulator truths belong in the critic only.
Symptom
Policies trained on ground-truth base linear velocity work in sim and fail on hardware, where only a drifting IMU and encoders exist - the policy has learned to depend on a signal that does not survive deployment.
Context
Many open-source locomotion stacks feed simulator ground-truth linear velocity to the actor. The reference team refused: the real robot has no ground-truth velocity. Asymmetric actor-critic keeps the training benefit of privileged information without deploying the dependency.
Change
Route ground-truth velocity, foot contact forces, and foot heights to the critic only; the actor observes exclusively signals that exist on hardware (IMU-derived quantities, encoders, commands, previous actions).
Outcome
Recorded as adopted doctrine in Lucen's experience log; the trained actor's input contract matches what the real robot can actually produce.
Mechanism
The critic is discarded at deployment, so it may consume any privileged state to reduce value-estimation variance; the actor's observation set is a deployment contract - anything in it that hardware cannot supply (or supplies with different noise/drift) becomes a train/deploy distribution shift the policy was never trained to handle.
Applies when
- designing actor/critic observation spaces
- reviewing a config where the actor sees base_lin_vel or contact forces
- sim policy is strong but real robot drifts, oscillates, or falls without obvious actuator cause
“很多开源代码库把真值线速度喂给策略,Asimov 团队没有,因为真机上没有真值速度,只有会漂的 IMU 和编码器;用完美速度训练出来的策略会依赖它,然后在硬件上失效。真值速度、足底力、足高统统只给 critic”
Experience.md § 观测空间的诚实性 (line 7) Real robot walked at half the sim clock for two generations - resolved by racing a reward-side and a plant-side evidence line, not by guessing
period-doubling-evidence-raceFor a hardware-only pathology, refuse to guess: pre-register one probe per side of the sim2real boundary (can the reward mechanism change it on hardware? can fitted plant parameters reproduce it in sim?) and let the first positive result direct the next version.
Symptom
The number-one sim2real gap: on hardware v6/v7 stepped at 1.23-1.32 Hz - almost exactly half the 2.50 Hz gait clock they were trained and simulated at; sim never reproduced it, two generations running.
Context
Instead of committing training budget to a guess, v8 pre-registered two mutually controlled evidence lines and kept the clock OUT of the training variables: (a) reward-side - if the v8 saturation fix revives joint_pos_ref (the term that pins the gait to the clock), re-run hardware and see whether frequency returns to 2.5 Hz (hypothesis: v7's frozen actions meant NO reward was pinning the gait to the clock, and the real plant - with armature and friction making high frequencies expensive - slid down to the leg's pendulum natural frequency ~1.1 Hz); (b) plant-side - record suspended joint data (fit_actuator), fit armature/friction, load the fitted values into sim2sim and see whether the 1.25 Hz reproduces IN SIM. Decision rule fixed in advance: "谁先给出阳性结果谁定 v9 的方向 (奖励侧 vs plant 侧)" - whichever line goes positive first sets the next version's direction.
Change
Period-doubling excluded from the v8 change set; both diagnostic lines scheduled in parallel as non-blocking work; frequency reported factually in acceptance with no pass/fail attached ("倍周期是否消失 不设判定,它是 §9 的关键证据").
Outcome
The gap was routed into a decisive-experiment structure rather than a speculative retrain; the plant-side line pointed at exactly the unmodeled armature/friction that were later measured and installed as the plant baseline. Resolution (era-2c full-plant retest): the family had TWO causes - v8's low-speed period-doubling vanished once measured armature+friction were installed (1.30 -> 2.50 Hz, bifurcation-edge machine sensitivity), while v7's stood untouched at 1.20 Hz (saturation-freeze-driven policy property) - both evidence lines paid off, one per case.
Mechanism
A behavior appearing only on hardware has candidate causes on both sides of the sim2real boundary; changing training to fix it tests only one side per expensive cycle. Two cheap parallel probes - one intervening on the reward mechanism, one making sim reproduce the real behavior - localize the cause to a side before any training money is spent, and sim-reproduction of a real pathology is itself the strongest form of plant validation.
Applies when
- a gait pathology appears on hardware but never in any simulator
- deciding whether a sim2real gap is reward-side or plant-side
- tempted to change the gait clock/reward to chase a hardware symptom
“倍周期(真机 1.23~1.32 Hz ≈ 时钟一半,v6/v7 连续两代;sim 从不出现):两条证据线互为对照——(a)… 真机重跑看频率是否回 2.5 Hz(假说:v7 没有任何奖励把步态钉在时钟上,真机 plant 有 armature/摩擦、高频贵,自由滑落到复摆自然频率 ~1.1 Hz);(b)真机吊挂录 fit_actuator.py … 看能否在仿真里复现 1.25 Hz。谁先给出阳性结果谁定 v9 的方向。”
train/WALK_V8_SPEC.md § 9. 平行线 (倍周期) The hip_roll (l+r) asymmetry scalar predicted real-robot lateral drift - promote validated sim scalars into the gate
hip-roll-sum-predicts-lateral-driftHunt for cheap sim scalars that predict real-robot behaviors, validate them on direction AND ordering across multiple policies, then promote them into the acceptance battery; treat later violations as debt to justify in writing, not noise to ignore.
Symptom
A persistent hip_roll left/right asymmetry row in the sim2sim symmetry table had been dismissed as "calibration or mechanical asymmetry" noise; meanwhile real deployments drifted sideways by policy-dependent amounts.
Context
Forward-kinematics analysis reframed the scalar: both hip_rolls move the feet in +y for positive angle, so a same-signed (l+r) sum IS a lateral translation mode - the scalar is a direct lateral-drift bias estimate. Checked against real deployments: s1e (l+r = -0.0178, smallest magnitude) was the steadiest with least drift; 700 (+0.0253) drifted mildly left; A800 (+0.0267) drifted clearly left with the largest tilt 12.9 deg. Direction correct 3/3, ordering correct 3/3 (the log's heading calls it "四枚四中", four-for-four).
Change
The scalar was promoted into the acceptance battery as a posture-class criterion alongside tilt-max median: "hip_roll 左右不对称 |l+r| 不得比父代大" - doubling as a heat proxy (error ~ torque ~ heating).
Outcome
Used at every later gate; when the C4 product exceeded it by +0.005 rad (~+0.3 deg vs parent), the criterion was not silently waived - it was booked as explicit debt with a mechanism argument (the increment is task-required, far smaller than the sidewalk amplitude +/-2.2 deg) plus a related account (stand saturation 32.4% -> 37.2%).
Mechanism
A policy's static joint-angle bias in a translation-producing mode integrates into real-world drift; sim can measure that bias precisely and cheaply. A sim scalar earns gate status exactly when its predictions are validated against hardware in both direction and ordering - and a validated gate may only be exceeded with a written mechanism-level justification, never silently.
Conflicts
The log's heading says "四枚四中" (4/4) but the evidence table lists three policies and the text says "方向 3/3、排序 3/3"; the fourth instance is not shown in this file.
Applies when
- a real robot drifts or leans in a policy-dependent way
- deciding which sim measurements deserve gate status
- a validated gate criterion is marginally exceeded by a new product
“s1e | −0.0178(绝对值最小)| 微右、最不飘 | 三者中最稳、飘最小 ✓ … A800 | +0.0267 | 左、最飘 | 明显左飘、倾角最大 12.9° ✓ 方向 3/3、排序 3/3。 → 正式纳入验收表(与「倾角 max 中位」并列为姿态类判据)。”
train/C_LADDER_RUN.md § 3e. 顺带:hip_roll 左右不对称 (l+r) 就是横移偏置 —— 四枚四中 Before training a one-leg stand the spec named the cheapest cheats - hopping on the support foot, a raised foot resting unloaded, a leg tripod - and gave each a countermeasure and a gate; one still appeared and was caught by exactly those gates
enumerate-cheapest-cheats-before-trainingBefore training, list the cheapest behaviours that would satisfy each reward term without doing the task, give each a countermeasure in the reward and a gate in acceptance, and prove the intended behaviour is reachable with a probe - then treat any gate the policy games as evidence about the reward, not the gate.
Symptom
The literature's single-leg benchmark reports eight state-of-the-art general policies holding a clean one-leg stand 0 times out of 90 - they survive by sneaking steps and hops - so the task's first adversary was the policy's own cheating.
Context
The spec's shape self-check ("what is the zero-cost option?") listed, for the one-foot bucket: the cheapest cheat, a foot resting on the ground without load, countered by a 5 N contact threshold plus positive swing income; the second cheapest, small hops on the support foot to reset balance, countered by a continuous support-air penalty plus a gate of zero support-foot flight segments. The probe that preceded training had already seen a third: early low-lift postures "survived" by pressing the swing foot at 78-95 N, a leg tripod, removed by folding the shank back. The two-foot bucket was checked too: its zero-cost behaviour is ordinary standing, with no odd base state.
Change
Countermeasures and gates written before training: swing-contact and support-air penalties, gate 2 (zero swing-foot contact frames above 5 N), gate 3 (zero support-foot flight segments).
Outcome
The first run still found the unloaded-foot cheat (a binary reward band gave it no gradient to lift) - and it was caught, by the contact gates and the cross-simulator comparison, not discovered on hardware. The retrained V0 passed all gates 40/40, including zero support-foot flight after the flight detector was corrected.
Mechanism
A policy optimizes the reward, not the intent; the cheapest behaviours that satisfy the reward are predictable from the reward's structure, and a gate written for each before training turns a silent cheat into a failed row.
Applies when
- designing rewards for balance, contact or "hold still" tasks
- benchmark policies are known to cheat the task
- writing acceptance gates for a new skill
“文献里 8 个 SOTA 通用策略在单脚站基准上 0/90 干净保持, 全靠偷步偷跳活命,这是本任务的第一反作弊对象 … 单脚桶下最便宜的作弊是"脚虚放地上不受力"——接触判定 >5 N 力阈(沿用),配 swing_height_band 正收入拉开。 … 第二便宜是"支撑脚小跳重置"——support_air_penalty 连续罚 + 验收门支撑脚腾空段=0 双保险。”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §0 目标口径 / §5 形状自检(零成本选项是什么) Torque caps cannot soften footfalls - impact is falling-mass momentum, only the reward can treat it
landing-impact-not-fixed-by-torque-capsClassify each hardware symptom by the physics that sets it: quantities fixed by ballistic momentum at contact must be treated through the policy's trajectory (reward terms on approach velocity/force), never through actuator caps - and size such penalty weights against your own tracking reward, not a lighter robot's.
Symptom
Footfalls slammed at 1.78x body weight in sim baseline (human walking: 1.2-1.5x); the tempting hardware-side fix was cutting actuator torque limits.
Context
Measured directly: scaling torque limits from x1.0 down to x0.4 left peak landing force essentially unchanged (1.75 -> 1.78x body weight) - the impact force comes from the momentum of the falling mass at touchdown, not from motor effort. The fix has to change the trajectory, i.e. the policy, i.e. the reward: feet_contact_forces penalty above a threshold of 113 N (= 1.2x the 9.58 kg robot's weight), clipped, weight -0.005. The weight was sized locally, not copied: the reference robot's -0.001 would amount to 0.9% of tracking reward on this robot ("策略不会理它" - the policy would ignore it); -0.005 gives 4.4%.
Change
Added threshold-type contact-force penalty (-0.005, threshold 1.2x body weight) as one of v6-minimal's three changes; hardware torque cuts explicitly rejected as a footfall treatment.
Outcome
Landing force 1.72x -> 1.55x by v6 (target <1.5x, missed by 3% - progress booked honestly); the torque-cap dead end was documented so it would not be retried.
Mechanism
At touchdown the ground stops a ballistic mass; the impulse is set by approach velocity and effective inertia, which motors can no longer influence in the final instant. Only earlier trajectory choices (approach velocity, timing) reduce it - and those are selected by the reward, not by actuator limits.
Applies when
- footfall impact or landing noise on hardware
- proposals to derate torque as a softness fix
- importing contact-force penalty weights from another robot
“⚠️ 硬件限扭降不了落脚力 —— 砸地力来自下落质量的动量: 实测 tau ×1.0→×0.4, 落脚力 1.75→1.78× 体重纹丝不动。只有这条奖励能治。… ⚠️ 权重不能用 Pi 的 −0.001 —— 实测在我们身上只占跟踪奖励的 0.9%, 策略不会理它 (Pi 6.94 kg 更轻)。−0.005 给到 4.4%。”
train/WALK_V6_MINIMAL.md § ③ 新增 feet_contact_forces Removing a hand trim re-exposed the plant offset it had been silently compensating - and a slope scan told bias from sensitivity
hand-trims-hide-plant-offsetsTreat hand-tuned trims as undocumented plant measurements: before deleting one, find what it compensates and re-house that knowledge in the model or the reward budget; diagnose posture errors with a sensitivity sweep to distinguish constant bias from gain problems.
Symptom
After switching from the old hand-trimmed default to the clean geometric zero, the retrained stand policy's only regression was torso lean: 1.8 deg -> 4.1 deg backward.
Context
The old default's ankle-pitch trim (-0.0489/+0.0628) had been pre-compensating a fore-aft COM mismatch; removing the trim removed the hidden compensation, and the posture reward alone was too weak to win it back. A COM sensitivity scan settled what kind of problem this was: sweeping base COM offset -50 to +50 mm gave nearly identical slopes for old and new policies (~0.026 deg/mm) - "不是质心敏感度问题, 是恒定偏置" (not a sensitivity problem, a constant bias). Fix landed in stand_v1b: posture corrected to +0.24 deg while keeping symmetry (<=0.1 deg) and low effort (0.259), disturbance rejection better than both predecessors. Model credibility was checked the honest way: v0's sim prediction at the real COM position (-22 mm) was -2.31 deg lean vs real measured 2.2-3.1 deg - "预测精准命中" - which is what licensed trusting v1b's -0.52 deg prediction. (Side flag from the same file: a sign convention had been documented wrongly in early comments - gravity_base[0] > 0 is forward lean.)
Change
Trims retired in favor of explicit modeling: symmetric geometric default plus a posture-reward budget sized to carry the real COM offset; the offset itself known (real COM ~22 mm behind model).
Outcome
stand_v1b passed acceptance as the standing lineage's final version; the walk-line requirement "加大躯干姿态惩罚权重" was upgraded from suggestion to mandatory, since walking amplifies what standing tolerates (real walk_v1 hit 26 deg lean vs sim 7.4).
Mechanism
Hand trims are plant knowledge stored in the wrong place - invisible, asymmetric, and stale after recalibration; removing them re-exposes the raw plant error. A sensitivity sweep separates the two possible diagnoses (slope change = control problem; parallel offset = constant plant bias), each with a different fix.
Applies when
- cleaning up hand-tuned offsets/trims in defaults or calibration
- a posture bias appears after a default or calibration change
- deciding whether a lean is a COM-sensitivity or constant-offset issue
“两者斜率几乎相同(≈0.026°/mm),v1 只是整体多后仰约 2.4° —— 不是质心敏感度问题,是恒定偏置。成因:旧 default 的踝俯仰 trim(−0.0489/+0.0628)本就预补偿了前后质心偏差,换成零位 default 后这份补偿没了 … v0 在真机质心处(−22 mm)的 sim 预测为 −2.31° 后仰,真机实测 2.2~3.1° 后仰 —— 预测精准命中。”
train/RETRAIN_v2.md § 4b. stand_v1 独立验证结果 / 4c. stand_v1b 验收结果 Write each config's expected hardware signature before the session - and if reality disagrees, change the books, not the conclusion
preregistered-real-expectationsBefore hardware runs, write per-config expected signatures and the disagreement rule (hardware outranks sim; discrepancies get recorded, not reconciled); validate the harness by checking it reproduces at least one known real behavior.
Symptom
Hardware impressions are easily narrated after the fact; without written expectations, any real-robot outcome can be made to "match" the sim story.
Context
The S2 acceptance sheet carried a section titled "sim 侧预注册预期 (事后核对, 不许事后改)" - per-configuration behavioral signatures written before the session: s1e@0.8 the disturbance king (push 159/160, zero chirality, all-mu 20/20) at the cost of speed gates 0/20 and zero-command wander ~0.98 m with -29.5 deg/20 s rotation; fric-3000 "walks accurately but is easier to push over"; fric-2400 neither. Credibility check included: the sim harness had reproduced the already-recorded real behavior (pace in place + right drift + net rotation -30 deg/20 s), which "提高本单全部预期的可信度". The anomaly clause fixed the epistemics in advance: if results systematically disagree with sim, "不改结论改账" - don't massage the conclusion, write the discrepancy into the books, and per the earlier zero-warning lesson, hardware wins.
Change
Every hardware session ships with a pre-registered expectation table (signature per config), a baseline-match credibility check, and a written precedence rule for disagreement.
Outcome
The A/B session became falsifiable: agreement confirms the proxy, disagreement is booked as a proxy-bias finding rather than argued away.
Mechanism
Pre-registration converts qualitative hardware sessions into tests of the sim-to-real mapping itself; a reproduced known behavior calibrates trust in the remaining predictions; and fixing "who wins on disagreement" beforehand prevents authority from drifting to whichever source flatters the plan.
Applies when
- planning any hardware acceptance or A/B session
- the sim harness's credibility in this regime is unestablished
- post-session write-ups tempt narrative fitting
“⚠️ sim 复现了真机已记录的「原地踏步 + 右漂 + 净旋 −30°/20s」—— harness 与真机行为对得上, 提高本单全部预期的可信度。… 结果与 sim 系统性不符 → 不改结论改账: 写进 README 该节, 按 「Isaac 指标三次零预警」的教训, 以真机为准。”
train/REAL_RUN_S2.md § 2. sim 侧预注册预期 (事后核对, 不许事后改) / 4. 异常处置 Every new penalty ships with a pre-registered withdrawal clause - if healthy gait must pay above the cap, the term stands down
calibration-threshold-with-withdrawal-clauseIntroduce every new penalty with: the zero-cost-option audit, a replay-calibrated weight formula (healthy pays a fixed small fraction of tracking), and a pre-registered withdrawal condition - and let the clause fire without argument when the calibration says the term cannot be afforded.
Symptom
Three same-shaped crashes had established a failure archetype: v4's clearance, v8a's landing window (weight off by 58x uncalibrated), and v6a's bare hip_yaw suppression all combined a zero-cost "don't move" option with a fee on any motion - a reverse barrier that pushes policies toward standing still.
Context
The v11 landing-window penalty was therefore introduced under a calibration-threshold protocol: (1) shape chosen with the window tightened (h_gate 0.03 -> 0.02, because 0.03 equaled the clearance target and priced the entire descent); (2) weight from a FORMULA, not judgment: measure the term's raw value on healthy replays (v5/v10b), set w = -(0.10-0.15 x tracking reward) / raw_healthy; (3) withdrawal clause pre-registered: if healthy gait must pay >15% of tracking no matter the tuning, the term is withdrawn to the next version rather than forced in - "不硬上". The companion hip_yaw quieting term ran the same protocol (calibrate on replays, healthy pays <=5%) and was later retired entirely when a structural fix (zero action scale) made its shaping tax unnecessary.
Change
Penalty introduction protocol: shape audit (what is the zero-cost option?), replay-based weight formula, healthy-pay cap with a written stand-down condition - all before training.
Outcome
The landing term was in fact withdrawn under its clause (v12 records "P5 落地窗口罚 已撤 … 维持撤下"), demonstrating the protocol firing as designed instead of the fourth same-type crash.
Mechanism
A penalty's damage mode is mispricing healthy behavior; since the healthy price is measurable in advance on replays, both the weight and the go/no-go decision can be computed rather than discovered by a ruined training run. The withdrawal clause converts "make it work" pressure into a clean deferral.
Applies when
- adding any motion-taxing penalty to a working gait
- a proposed term's weight has no measurement behind it
- a previous same-shaped term crashed training
“权重公式而非拍脑袋:先在 v5/v10b 回放上量 h_gate=0.02 的原始值,w = −(0.10~0.15 × 跟踪奖励) / raw_健康;标定门槛:若健康步态无论如何要付 >15% 跟踪,本项撤下留 v12,不硬上 (v4 clearance/v8a-B/v6a 三次同型翻车的教训:代价为零的"不动"选项 + 一动就收费 = 反向壁垒)。”
train/WALK_V11_SPEC.md § 6. P5 —— 落地窗口罚(三代欠账,标定门槛制) Compare the achieved reward to the computed ignore-floor to tell "never learned" from "learned but unprofitable"
ignore-floor-diagnosisFor any skill that trains flat, compute the reward the null policy would earn on that term; achieved==floor means the behavior never paid out (find why: exploration, reward observability, or feasibility) - do not tune weights first.
Symptom
C4 sidewalk failed on both arms; the question was whether the policy had found sidewalk and rejected it as unprofitable, or never found it at all - two diagnoses with opposite fixes.
Context
The theoretical value of the sidewalk tracking term for a policy that completely ignores the command was computable from the command distribution: 0.189. Trained final values landed at 0.187 (arm A) and 0.204 (arm B) - sitting exactly on the ignore-floor - while the honest balance at the optimum actually favored sidewalking (net +0.88/step inside the side bucket). Later the same arithmetic closed the whole saga: for the true reward landscape, doing real sidewalk scored 0.153 vs 0.944 for ignoring - the policy's refusal "是理性最优,不是探索失败" (rational optimum, not exploration failure) under one hypothesis, and under the final measurement-corrected account the policy had "每一次都在 做理性选择" (made the rational choice every time).
Change
Diagnostic rule adopted: compute the ignore-floor for the new term; if the achieved value sits on it, the behavior was never expressed in useful volume (or the reward cannot distinguish it - check both); if the achieved value is above floor but the behavior is absent at deployment, the policy sampled it and priced it out - then the reward balance, not exploration, is the lever.
Outcome
Correctly identified that PPO had not merely under-valued sidewalk; each subsequent hypothesis (waveform sign, regularization cage, exploration form, reward kernel) was tested against this floor arithmetic, which kept the search honest through three reversals.
Mechanism
Every reward term has a computable value under the null behavior; the achieved-vs-floor gap is a one-number audit of whether the optimizer ever monetized the target behavior. It converts "training failed" into one of two mechanistically distinct states with different fixes.
Applies when
- a new skill's tracking reward plateaus early
- deciding between exploration fixes and reward-weight fixes
- post-mortem of a failed curriculum rung
“track_lin_vel_y_exp 训练终值恰好坐在「完全无视指令」的底分上(A 0.187 / B 0.204,理论值 0.189),而终点 balance 明明有利(side 桶内净 +0.88/步)。不是学会了不划算,是根本没学到。”
train/C_LADDER_RUN.md § 3e. 为什么首战 FAIL Training at scale 0.4 is NOT the twin of deploying 0.5 at power 0.8 - the algebra matches, the learned policy does not
deploy-scaling-not-training-equivalentNever assume deploy-side scalings can be folded into training-time constants ("burning the crutch into training"): the learned optimum depends on the training-time authority, so treat such conversions as full experiments with pre-registered expectations and a sim2sim gate before any hardware.
Symptom
s1g (S1.6) trained from zero at action_scale 0.4 - meant as the "training twin" of the hardware-proven s1c-at-power-0.8 (0.8 x 0.5 = 0.4) - was all green in Isaac (zero falls, reward 117) yet scored 0/3 across all eight checkpoints and 0/20 at 20 seeds in the MuJoCo gate, falling forward at median 1.57 s with a 2.9x speed overshoot.
Context
The pre-registered expectation (survival gate should pass, since the conviction matrix showed s1c@0.8+delay2 all-survive) was cleanly falsified, and the harness was acquitted by controls: --delay 0 fell identically (not a delay fragility), check_contract all green, and s1c through the same harness survived 2/3. The verdict: "「s1c@0.8 = 0.4 训练孪生」的代数等价不成立" - a policy deployed with a derated output still LIVES in the 0.5 internal model it trained under (its value function, its expectations of its own authority), while a policy that starts training with reduced authority learns a different, clip-hugging gait with zero margin for plant differences ("部署端打折的策略活在 0.5 的内模里,训练起点收权限学出的是贴 clip 的 另一套步态,对 plant 差异零余量"). Result: the policy was withdrawn before hardware ("撤回——不上真机"), the lineage root moved back to the 0.5-contract s1c-5500, and this became the C ladder's cited fact-check ("s1g 是 0/20 证伪出局的那一代").
Change
The amplitude-surgery route abandoned; contract kept at scale 0.5; the deploy-side 0.8 crutch later retired on its own merits when the delay-complete s1e generation ran at full power.
Outcome
One training run bought a clean falsification of a plausible algebraic identity; no hardware time was spent on it because the sim2sim gate caught it.
Mechanism
Output scaling commutes with the network arithmetic but not with learning: the training-time scale shapes which gait solutions are reachable and how much clip headroom the optimum keeps. A derated mature policy retains the wide-authority solution executed softly; a from-zero narrow-authority policy finds a different optimum that saturates its smaller envelope - the two are not the same controller in different units.
Applies when
- proposing to move a deployment derating into a training constant
- a scaled-down contract policy hugs the action clip
- Isaac-green / cross-sim-zero results on a re-scaled lineage
“预注册 a) 证伪——Isaac 全绿(零摔/reward 117)但 MuJoCo --delay 2 八档 checkpoint 扫描全数 0/3、iter6500 20-seed 0/20 … 「s1c@0.8 = 0.4 训练孪生」的代数等价不成立: 部署端打折的策略活在 0.5 的内模里,训练起点收权限学出的是贴 clip 的另一套步态,对 plant 差异零余量。”
train/OMNI_V0_SPEC.md § 3. S1.6 判决(2026-08-07 验收) Fix a too-deep nominal pose before adding any penalties - the default stance defines the basin training starts in
nominal-posture-before-penaltiesBefore tuning penalties on a degenerate gait, audit the nominal pose and height targets against morphology and published ratios; if the default stance encodes the degenerate behavior, fix it first - and recompute dependent quantities (init height) by FK, not by hand.
Symptom
Policy lived in a crouched shuffle; nominal knee angle was 0.5 rad (28.6 deg) - deeper than published configs (Unitree G1 0.3 rad / 17.2 deg, Booster T1 0.4 rad) - so the policy's starting point and its action-space center both sat inside the crouch basin.
Context
Initially ranked "secondary" in the local diagnosis, this was promoted to co-first priority by the cross-check against published reward tables, which states that with nominal knee flexion above ~0.4 rad, fixing the posture must precede adding any penalties ("改这个之前别加任何 惩罚都是白费"). Companion base-height items: walk profile had weakened base_height_l2 to -5.0 (base class -10, field standard -10 to -20, "the second most common cause of death"), and the height target must be the STANDING height (0.384), not the crouch height.
Change
Nominal knee 0.5 -> 0.3 rad with init_base_height recomputed by MuJoCo FK (0.3739 -> 0.3802); base_height_l2 restored to -10 with standing height target; both bundled as first-priority alongside the clearance term.
Outcome
Part of the v5/v6 package that lifted swing height to 34 mm and tracking to 87%; the crouch basin stopped being the default answer.
Mechanism
The nominal pose is the fixed point every regularizer pulls toward and the point where action=0 lands; if that point is itself the degenerate posture, every penalty fights the geometry. Correcting the attractor is prior to shaping the gradient field around it.
Applies when
- policy converges to a crouched or collapsed posture
- nominal joint angles were chosen for stability rather than gait
- base-height reward targets or weights were locally weakened
“研究明确说"nominal 膝屈超过 ~0.4 rad 必须先改,改这个之前别加任何惩罚"。我们是 0.50,超标。… base_height_l2 在 walk profile 里被减到 −5.0(基类是 −10)。研究说这是"第二常见死因"且应 −10 ~ −20。改回 −10。目标高度用站立高 0.384 是对的(研究要求 target 必须是*站立*高度而非蹲姿)。”
train/WALK_DIAGNOSIS.md § 修正 ①(升级优先级) / 修正 ④ Model CAN polling skew - joint observations are 6-9 ms stale by read order
can-timing-skew-modelingIf joints are read sequentially over a shared bus, reproduce the per-group observation staleness in sim (or randomize it over the measured range) - synchronous observations are a privileged fiction.
Symptom
Policies trained with synchronous joint observations degrade on hardware where motors are polled sequentially over CAN - hip data is already 6-9 ms old by the time ankle data arrives.
Context
Menlo's core sim2real finding on a leg platform of almost identical mass to Lucen's. CAN is a sequential bus: one poll cycle reads motors in a fixed order, so the observation vector mixes timestamps. Lucen runs 6:6 motors on a dual CAN split, so the problem transfers one-to-one.
Change
Explicitly model CAN timing skew in sim: give the joint groups different observation delays matching physical read order. Menlo went further - running real firmware in the loop with a motor simulator between MuJoCo and firmware injecting 0.4-2 ms uniformly distributed delay.
Outcome
Reported by the reference team as their core sim2real enabler on a same-scale platform; recorded in Lucen's experience log as directly applicable ("这个问题一模一样").
Mechanism
A policy exploits any cross-joint temporal coherence present in training observations; when hardware breaks that coherence per bus position, the learned feedback acts on inconsistent state estimates, injecting phase error exactly at the control bandwidth.
Conflicts
Second-hand episode: outcome numbers are the reference team's report, not a Lucen-run experiment; Lucen adopted the requirement but the corpus has no Lucen A/B of skew-on vs skew-off.
Applies when
- robot polls actuators sequentially over CAN/RS485 or similar shared bus
- sim2real degradation appears as jitter or oscillation not seen in sim
- designing the observation/delay model before a training run
“电机走 CAN 是顺序轮询的,髋部电机的数据到踝部电机上报时已经陈旧了 6-9 ms,他们直接在仿真里显式建模了 CAN 时序偏斜,按读取顺序给三组关节不同的观测延迟。… 注入 0.4–2 ms 的均匀分布延迟。你们是 6:6 双 CAN 分总线,这个问题一模一样”
Experience.md § 执行器 + 时序建模 (line 6) A single run's drift direction may be a limit cycle, not a policy bias - check the sign distribution across seeds
multiseed-sign-test-for-driftDistinguish "bias" from "broken symmetry limit cycle" by the sign distribution over many seeds; report drift as (mean, sign split), and never compare single-run drift magnitudes across versions.
Symptom
Net yaw over 15 s appeared to worsen from -41 deg (v2) to -84 deg (v4), inviting the conclusion that the new version drifted more.
Context
The Isaac-side view across 32 environments told a different story: per-env yaw was mixed-sign (20 negative / 12 positive) with mean ~0 - the drift is a limit cycle whose direction depends on initial conditions, not a systematic policy bias. The single MuJoCo run had sampled one draw from that distribution, so its magnitude could not be compared across versions as if it were a property.
Change
Evaluation rule: before classifying drift as systematic, run multiple seeds and examine the sign distribution; single-trajectory drift magnitudes are samples, not properties.
Outcome
The v2-vs-v4 drift "regression" was reclassified as not-established; later drift work (hip_roll l+r bias) used cross-policy, cross-seed evidence instead.
Mechanism
Symmetric dynamical systems can settle into either of two mirrored limit cycles; the selected cycle is decided by noise and initial state. A statistic whose sign is initial-condition-dependent has no meaning as a single sample - only its distribution does.
Applies when
- comparing heading drift or lateral drift across policy versions
- a symmetric-looking behavior shows a consistent direction in one run
- deciding whether to fix "drift" in reward or calibration
“偏航反而变差(−41° → −84°):注意 Isaac 侧 32 env 的逐 env 偏航是正负混合(20/12)、均值 ≈0,说明这是极限环性质(方向随初值)而非策略偏置 —— MuJoCo 单次跑测到的是分布里的一个样本,不能当作系统性偏差。要判断需多种子统计。”
train/WALK_DIAGNOSIS.md § walk_v4 独立验收 读法 (偏航) A sim veto needs real confirmation too - the worst sim cell was scheduled as the most informative hardware run
sim-veto-needs-real-confirmationNever let sim alone both condemn a purpose-built configuration and escape audit: spend one cheap, safeguarded hardware run on the condemned cell, pre-registering what agreement and disagreement would each imply about the proxy.
Symptom
The fric-2400@kd1.0 combination was sim's worst cell across the board (survival 17/20 - the only miss, mu0.4 1/20, push 103/160, zero-cmd 2/20), yet it was the only product specifically trained for the kd1.0 deployment gain - discarding it on sim evidence alone would leave the sim's own validity untested exactly where it mattered.
Context
The team had been burned in the other direction before ("Isaac 指标三次 零预警" - training-side metrics gave zero warning three times), so the symmetric rule was written: sim's rejection also needs hardware confirmation ("sim 判被支配 ≠ 真机被支配 … sim 的否决也要真机确认"). The run was pre-registered with a dual reading: real matches sim -> the S2f ladder closes and the fork root is settled; real clearly better than sim -> the MuJoCo proxy has a systematic bias in the kd1.0/low-margin region, "那比选型本身重要得多" - and every S2f sim acceptance would need re-scoring.
Change
The condemned configuration was kept on the hardware roster (last, spotted, minimal exposure) explicitly as a proxy-validation probe, not as a deployment candidate.
Outcome
Session design captured either result as progress: selection confirmed, or a proxy bias discovered that would re-price the whole ladder's verdicts.
Mechanism
Every sim verdict is a joint statement about the policy AND the proxy; cells where a policy was purpose-trained for the exact condition sim condemns are where proxy error is most likely and most costly. Testing the veto converts a selection decision into a calibration measurement of the evaluator itself.
Applies when
- sim rejects the configuration that targets the actual deployment condition
- the eval proxy's calibration has never been checked in that regime
- deciding which hardware runs are worth their risk
“但它也是唯一为 kd1.0 部署档专门训的产物 —— sim 判被支配 ≠ 真机被支配, 「Isaac 指标三次零预警」的教训反过来同样成立: sim 的否决也要真机确认。… 若真机明显好于 sim → MuJoCo 代理在 kd1.0/低裕度区有系统性偏差, 那比选型本身重要得多。”
train/REAL_RUN_S2.md § 上机名单 note / 2. sim 侧预注册预期 ⑤ The +/-50 mm lateral COM randomization meant to spread the legs coincided with legs pulling IN - rolled back per its own pre-registered contract
com-dr-rollback-on-symptomWhen adopting a DR value that covers no local measurement, write its intent and rollback trigger into the config at adoption time; roll it back as the control arm the moment the symptom contradicts the intent, and promote the symptom's metric into the acceptance battery.
Symptom
After v7 adopted the reference developer's oversized lateral COM randomization (+/-50 mm) explicitly to force leg spread, the real robot's legs narrowed instead - lateral mean 154 mm / closest 107 mm in sim (nominal 214.5), narrower still on hardware with occasional leg contact.
Context
The rollback was clean because the adoption had been honest: the robot.yaml comment recorded the intent AND that the +/-50 value covered no local measurement (only a 16/7 mm measured offset existed; even the prior widening to +/-20 was subjective), plus the reference's own reported side effect (base sway) and the note "这一项要单独跑、 单独归因". When the opposite symptom appeared, v8 returned y to +/-20 mm as the control arm ("要么没起作用、要么帮了倒忙 … 按约定退回做 对照"), kept x/z untouched (a noise-level difference not worth another variable), and named the second suspect: the landing penalty itself, via the reference's own three-link chain (landing penalty -> stance narrows -> spacing penalty needed). A gate lesson was booked in the same table: v7's sim numbers had ALREADY crossed the line (154/107 vs v5's 182/147) - "这个指标本可拦下 v7" - so foot-distance became a standing acceptance row (min >120 mm, zero leg-leg contacts).
Change
base_com_offset_m y: 0.050 -> 0.020 (x/z kept), regenerated through the export tool rather than hand-editing derived files; foot-distance acceptance row added.
Outcome
A borrowed DR lever with no local measurement basis was retired the moment its symptom contradicted its purpose, at single-variable cost; the metric that would have caught it pre-hardware entered the gate.
Mechanism
DR ranges shape behavior through the policy's robustness strategy, which is jointly determined with every reward term; a lever that forces stance width on one robot can be dominated by a stronger narrowing pressure (landing softness) on another. Levers adopted without local measurement must carry their own rollback trigger, because there is no nominal to argue from when they misbehave.
Conflicts
Causality is not fully closed in the source: the narrowing may come from the landing penalty rather than the COM lever ("腿距的第二嫌疑人是 ④ 本身"); the rollback is the pre-agreed control experiment, not a verdict that the lever caused the narrowing.
Applies when
- importing DR ranges or behavioral-forcing randomizations from references
- a DR lever's observed effect contradicts its documented purpose
- a sim metric existed that would have caught a shipped regression
“⑥ 的本意 … 是逼策略把脚分开;真机结果是脚向内收且偶发相碰——要么没起作用、要么帮了倒忙。… 注释当时就写了"这一项要单独跑、单独归因"。现在症状出现了,按约定退回做对照。”
train/WALK_V8_SPEC.md § 3. 改动 C — 质心随机化退回(撤销 v7-⑥ 的 y 项)