Training Coach
Doctrine
A report may cite any of these as doctrine-N.
doctrine-1Contract freeze and fingerprint disciplineThe policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.
Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).
Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.
doctrine-2Attribution by resolved training params - never eval-override knobsCapability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.
Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).
Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.
doctrine-3PASS gates become constraints; FAIL gates become objectivesOnce a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.
Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.
Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".
doctrine-4One variable per ladder rung - counted against what the checkpoint sawA rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.
Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.
Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.
doctrine-5Pre-register risks, readings, and stop criteria before the ladderBefore a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.
Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).
Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.
doctrine-6Plant parameters are measured, never inventedEvery plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.
Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).
Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.
doctrine-7Sim2sim gate before sim2real - under deployment conditionsEvery checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).
Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).
Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.
doctrine-8Observation honesty - the actor's inputs are a hardware contractThe actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.
Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).
Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.
doctrine-9Reward economics are audited in realized currencyReward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.
Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).
Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.
doctrine-10The zero-cost option must be the desired behaviorFor every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.
Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).
Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.
doctrine-11Measurement discipline: independent referees, signs, distributionsA disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.
Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).
Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.
doctrine-12The deployment pipeline is plantIrreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.
Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).
Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.
doctrine-13DR budget is finite; its distribution is the measured supportRobustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.
Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).
Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.
doctrine-14Gates measure what hardware feels: posture, margins, stripped assistsAcceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.
Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).
Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.
doctrine-15Fork and root selection: recoverability, maturity, frozen rewardsChoose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.
Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).
Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.
doctrine-16Curricula: verified engagement, lineage counters, disease-phase gatingAutomatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.
Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).
Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.
doctrine-17Probe before training: feasibility first, hypotheses in tablesAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.
Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).
Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).
doctrine-18External advice is recomputed locally; values transfer as ratiosEvery external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.
Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).
Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.
doctrine-19Hardware sessions are scripted experiments, not tuning sessionsReal-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.
Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).
Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.
doctrine-20Close questions in writing; restart when the debt is structuralAudited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.
Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).
Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.
doctrine-21Name the quantity in the space it lives inA goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.
Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).
Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.
doctrine-22Continuation needs a live gradient; a release is chosen by a scanContinue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.
Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).
Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.
Experience cards
97 cards matching “preregistered-real-expectations”.
Write each config's expected hardware signature before the session - and if reality disagrees, change the books, not the conclusion
preregistered-real-expectationsBefore hardware runs, write per-config expected signatures and the disagreement rule (hardware outranks sim; discrepancies get recorded, not reconciled); validate the harness by checking it reproduces at least one known real behavior.
Symptom
Hardware impressions are easily narrated after the fact; without written expectations, any real-robot outcome can be made to "match" the sim story.
Context
The S2 acceptance sheet carried a section titled "sim 侧预注册预期 (事后核对, 不许事后改)" - per-configuration behavioral signatures written before the session: s1e@0.8 the disturbance king (push 159/160, zero chirality, all-mu 20/20) at the cost of speed gates 0/20 and zero-command wander ~0.98 m with -29.5 deg/20 s rotation; fric-3000 "walks accurately but is easier to push over"; fric-2400 neither. Credibility check included: the sim harness had reproduced the already-recorded real behavior (pace in place + right drift + net rotation -30 deg/20 s), which "提高本单全部预期的可信度". The anomaly clause fixed the epistemics in advance: if results systematically disagree with sim, "不改结论改账" - don't massage the conclusion, write the discrepancy into the books, and per the earlier zero-warning lesson, hardware wins.
Change
Every hardware session ships with a pre-registered expectation table (signature per config), a baseline-match credibility check, and a written precedence rule for disagreement.
Outcome
The A/B session became falsifiable: agreement confirms the proxy, disagreement is booked as a proxy-bias finding rather than argued away.
Mechanism
Pre-registration converts qualitative hardware sessions into tests of the sim-to-real mapping itself; a reproduced known behavior calibrates trust in the remaining predictions; and fixing "who wins on disagreement" beforehand prevents authority from drifting to whichever source flatters the plan.
Applies when
- planning any hardware acceptance or A/B session
- the sim harness's credibility in this regime is unestablished
- post-session write-ups tempt narrative fitting
“⚠️ sim 复现了真机已记录的「原地踏步 + 右漂 + 净旋 −30°/20s」—— harness 与真机行为对得上, 提高本单全部预期的可信度。… 结果与 sim 系统性不符 → 不改结论改账: 写进 README 该节, 按 「Isaac 指标三次零预警」的教训, 以真机为准。”
train/REAL_RUN_S2.md § 2. sim 侧预注册预期 (事后核对, 不许事后改) / 4. 异常处置 Training at scale 0.4 is NOT the twin of deploying 0.5 at power 0.8 - the algebra matches, the learned policy does not
deploy-scaling-not-training-equivalentNever assume deploy-side scalings can be folded into training-time constants ("burning the crutch into training"): the learned optimum depends on the training-time authority, so treat such conversions as full experiments with pre-registered expectations and a sim2sim gate before any hardware.
Symptom
s1g (S1.6) trained from zero at action_scale 0.4 - meant as the "training twin" of the hardware-proven s1c-at-power-0.8 (0.8 x 0.5 = 0.4) - was all green in Isaac (zero falls, reward 117) yet scored 0/3 across all eight checkpoints and 0/20 at 20 seeds in the MuJoCo gate, falling forward at median 1.57 s with a 2.9x speed overshoot.
Context
The pre-registered expectation (survival gate should pass, since the conviction matrix showed s1c@0.8+delay2 all-survive) was cleanly falsified, and the harness was acquitted by controls: --delay 0 fell identically (not a delay fragility), check_contract all green, and s1c through the same harness survived 2/3. The verdict: "「s1c@0.8 = 0.4 训练孪生」的代数等价不成立" - a policy deployed with a derated output still LIVES in the 0.5 internal model it trained under (its value function, its expectations of its own authority), while a policy that starts training with reduced authority learns a different, clip-hugging gait with zero margin for plant differences ("部署端打折的策略活在 0.5 的内模里,训练起点收权限学出的是贴 clip 的 另一套步态,对 plant 差异零余量"). Result: the policy was withdrawn before hardware ("撤回——不上真机"), the lineage root moved back to the 0.5-contract s1c-5500, and this became the C ladder's cited fact-check ("s1g 是 0/20 证伪出局的那一代").
Change
The amplitude-surgery route abandoned; contract kept at scale 0.5; the deploy-side 0.8 crutch later retired on its own merits when the delay-complete s1e generation ran at full power.
Outcome
One training run bought a clean falsification of a plausible algebraic identity; no hardware time was spent on it because the sim2sim gate caught it.
Mechanism
Output scaling commutes with the network arithmetic but not with learning: the training-time scale shapes which gait solutions are reachable and how much clip headroom the optimum keeps. A derated mature policy retains the wide-authority solution executed softly; a from-zero narrow-authority policy finds a different optimum that saturates its smaller envelope - the two are not the same controller in different units.
Applies when
- proposing to move a deployment derating into a training constant
- a scaled-down contract policy hugs the action clip
- Isaac-green / cross-sim-zero results on a re-scaled lineage
“预注册 a) 证伪——Isaac 全绿(零摔/reward 117)但 MuJoCo --delay 2 八档 checkpoint 扫描全数 0/3、iter6500 20-seed 0/20 … 「s1c@0.8 = 0.4 训练孪生」的代数等价不成立: 部署端打折的策略活在 0.5 的内模里,训练起点收权限学出的是贴 clip 的另一套步态,对 plant 差异零余量。”
train/OMNI_V0_SPEC.md § 3. S1.6 判决(2026-08-07 验收) Every rung gets written stop criteria - hit any one, stop; tightened when priors say results should come fast
preregistered-stop-criteria-per-rungFreeze per-rung stop criteria (old-skill floors, new-skill deadline, oscillation signature, known pathology signatures) before training, stop on first hit - and shorten the deadline in proportion to how fast your mechanism says results should appear.
Symptom
Continued training past the point of degradation had previously destroyed capabilities (C1 trained away the root's backward skill); without hard stop rules, sunk-cost reasoning keeps runs alive too long.
Context
Standard rung protocol - eval_c_matrix every 100 iters at 5 seeds with a fixed criteria list, frozen before training ("开训前写死,事后不许改"): (1) stand or vx+0.15 at <=3/5 for two consecutive points; (2) vx+0.15 tracking <70% for two consecutive points (vs recorded root baseline); (3) oscillation signature - survival repeatedly crossing zero between adjacent checkpoints, stop on first occurrence, do not wait for confirmation; (4) new-skill-no-progress deadline (e.g. wz tracking <30% or still same-signed after iter 400); (5) a known lineage pathology signature (s1c arm-B: scatter -> half-recover -> collapse).
Change
Stop budgets scale with prior knowledge: when C4-redo4's feed-forward was already proven open-loop, the no-progress deadline was tightened from +500 to +100 iters ("前馈是开环就给 71~96% 的 … 若 +100 还没有 … 不值得再烧").
Outcome
Multiple rungs were stopped exactly on criterion (C4-redo criterion 4 at iter 1200; redo2/redo3 at absolute 900), converting each into a clean hypothesis test instead of a drifting run; no rung burned its full cap on a dead hypothesis.
Mechanism
Degradation of retained skills is a lineage-level injury that later training often cannot undo, so detection latency is capital loss; and a stop rule written before the run cannot be bent by hope. Tightening deadlines when priors predict immediate results converts "no result yet" into evidence against the mechanism.
Applies when
- starting any resumed/curriculum training rung
- deciding whether to keep training a run that shows early regression
- a mechanism-backed change should produce results immediately
“每 100 iter eval_c_matrix.py --seeds 5,命中任一即停 … 存活在相邻 checkpoint 间反复跨越 0(出现即停)… 收紧版:绝对 iter 900(= +200)时 vy±0.10 跟踪仍 <30% 即停”
train/C_LADDER_RUN.md § 3h. 预注册停梯判据(开训前写死,事后不许改) The first real-robot get-up was "very violent, kicking on the floor, dangerous" - a sim-perfect policy with no reason to be slow, unbounded absolute targets, no domain randomization and a rate limiter that filtered nothing; the task was restated as "safe, slow, transferable"
first-real-get-up-violent-stage-one-policyDo not put a get-up policy on hardware until its action is bounded (hard bound or state-anchored targets), smoothed, randomized and tested at the real pipeline's latency, and say explicitly that the task is "safe, slow and transferable" - a simulation-perfect policy optimizes only "gets up".
Symptom
On 2026-08-09 the user ran a V0-lineage recovery policy on the real robot and stopped it: very violent, kicking on the floor, dangerous. The planned next rung (a heavier torque_headroom) was never started.
Context
The spec had pre-registered that R0/R1 products stay in simulation and that the real-robot precondition was the R3 smoothing rungs plus a bridge-slew check plus a hanging protocol; the robustness (DR) rungs had not run. In simulation the policy passed 100% with a get-up of about a second. Which ONNX, which gain profile and whether a torque/joint log existed were left "to be recorded later" and never were.
Change
The V0 ladder was stopped at its best product (R3.1, sim only) and a re-rooting proposal was put to the user. The spec's four-layer account: style (the reward pays for standing early and nothing pays for slowness - HumanUP's "Stage I" get-up, "fast but unsafe ... infeasible for real-world deployment"); impact (full-range absolute targets with no hard bound, raw |a| up to 4.77, action saturation 100%, a single-step change of 0.306 saturating hip_pitch); transfer (zero DR, friction pinned at 1.0, the learned leg bracing); link (the bridge's RL slew equals vel_limit, 0.2-0.66 rad per step, while the real pipeline has 1-2 steps of time-varying latency and acceptance ran at delay 0).
Outcome
The line was re-rooted twice (training-side rate limit, then the beta-anchored action space) and gained a hang protocol before the next real attempt; on 08-11 a beta-anchored policy produced the line's first real get-up.
Mechanism
A task reward that pays for standing early selects the fastest feasible get-up; with absolute full-range targets every large target jump is a torque impulse bounded only by the clip; zero DR and braced-leg solutions do not transfer; and a limiter set at the velocity limit does nothing at 50 Hz.
Conflicts
The four layers are the spec's reconstruction from simulation probes and the literature; the real run's policy file, gain profile and log were never recorded, so no layer was confirmed against hardware data.
Applies when
- a first hardware trial of a high-effort skill is being scheduled
- sim success is high but the policy saturates actions or torques
- pre-registered hardware preconditions are not all met
“用户真机反馈:**非常猛、地上乱踢、危险**,叫停(R3.3 torque_headroom 加档已选型 weight −0.5→−1.5,未启动)。真机细节(哪个 onnx、什么档、有无 τ/q log)**待补记** … 任务从"能起来"变成 **"安全、慢、可迁移"** … **链路层**:桥层 slew RL 档 = vel_limit(10/20/33 rad/s ≈ 每拍 0.2~0.66 rad), 对 recovery 形同虚设;真机 1~2 拍时变延迟,验收默认 delay 0。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 真机叫停与换根判决(2026-08-09) The hip_roll (l+r) asymmetry scalar predicted real-robot lateral drift - promote validated sim scalars into the gate
hip-roll-sum-predicts-lateral-driftHunt for cheap sim scalars that predict real-robot behaviors, validate them on direction AND ordering across multiple policies, then promote them into the acceptance battery; treat later violations as debt to justify in writing, not noise to ignore.
Symptom
A persistent hip_roll left/right asymmetry row in the sim2sim symmetry table had been dismissed as "calibration or mechanical asymmetry" noise; meanwhile real deployments drifted sideways by policy-dependent amounts.
Context
Forward-kinematics analysis reframed the scalar: both hip_rolls move the feet in +y for positive angle, so a same-signed (l+r) sum IS a lateral translation mode - the scalar is a direct lateral-drift bias estimate. Checked against real deployments: s1e (l+r = -0.0178, smallest magnitude) was the steadiest with least drift; 700 (+0.0253) drifted mildly left; A800 (+0.0267) drifted clearly left with the largest tilt 12.9 deg. Direction correct 3/3, ordering correct 3/3 (the log's heading calls it "四枚四中", four-for-four).
Change
The scalar was promoted into the acceptance battery as a posture-class criterion alongside tilt-max median: "hip_roll 左右不对称 |l+r| 不得比父代大" - doubling as a heat proxy (error ~ torque ~ heating).
Outcome
Used at every later gate; when the C4 product exceeded it by +0.005 rad (~+0.3 deg vs parent), the criterion was not silently waived - it was booked as explicit debt with a mechanism argument (the increment is task-required, far smaller than the sidewalk amplitude +/-2.2 deg) plus a related account (stand saturation 32.4% -> 37.2%).
Mechanism
A policy's static joint-angle bias in a translation-producing mode integrates into real-world drift; sim can measure that bias precisely and cheaply. A sim scalar earns gate status exactly when its predictions are validated against hardware in both direction and ordering - and a validated gate may only be exceeded with a written mechanism-level justification, never silently.
Conflicts
The log's heading says "四枚四中" (4/4) but the evidence table lists three policies and the text says "方向 3/3、排序 3/3"; the fourth instance is not shown in this file.
Applies when
- a real robot drifts or leans in a policy-dependent way
- deciding which sim measurements deserve gate status
- a validated gate criterion is marginally exceeded by a new product
“s1e | −0.0178(绝对值最小)| 微右、最不飘 | 三者中最稳、飘最小 ✓ … A800 | +0.0267 | 左、最飘 | 明显左飘、倾角最大 12.9° ✓ 方向 3/3、排序 3/3。 → 正式纳入验收表(与「倾角 max 中位」并列为姿态类判据)。”
train/C_LADDER_RUN.md § 3e. 顺带:hip_roll 左右不对称 (l+r) 就是横移偏置 —— 四枚四中 The recovery line's real-robot verdicts live in three places that disagree - the spec's "fairly stable, first real stand-up" for v2_6, an undated runbook note that only v3_1p1c works, and a first run whose details were never recorded
write-hardware-verdicts-backA hardware verdict is a dated entry in the authoritative ledger - policy file and stamp, gain profile, floor, battery, tries, log file, what was seen - written back the same day; a note in a command file is a pointer, not a verdict, and a newer verdict that contradicts an older one must say so.
Symptom
Asked "which recovery policy works on the real robot", the sources give different answers, and none of them carries the conditions of the test.
Context
08-09: the first real run was stopped as violent and dangerous; which ONNX, which gain profile and whether a log existed were marked "to be recorded" and never were. 08-11: v2_6 was "fairly stable", the line's first real get-up, with splits after standing; v2_5b's result and both CSVs were "to be reported". 08-14: v3_1p1c was stamped and pushed, with "the real first test still needs the user present"; the spec records no hardware result for it. The operator runbook (undated) puts above the v2_6 and v2_5b floor commands the note that none of the recovery policies below work, only recovery_v3_1p1c - a verdict never written back into the spec, with no date, floor, battery, number of tries or log attached.
Change
None recorded in the sources; this card records the gap.
Outcome
The line's authoritative record ends with v3_1p1c as the product awaiting its first real test, while the operator's note implies it is the only one that works and that v2_6 (recorded as a success) does not.
Mechanism
Verdicts given at the robot travel by word of mouth and command-file comments; without a record carrying the conditions, a later reader cannot tell a changed verdict from a changed floor, battery or stack.
Conflicts
§43 (2026-08-11) records v2_6 as the first successful real get-up ("fairly stable"); the undated runbook says every recovery policy except recovery_v3_1p1c does not work; §49 (2026-08-14) says v3_1p1c's first real test was still pending. The runbook's claim has no date and was never written back to the spec, so it cannot be ordered against §43.
Applies when
- choosing which policy to deploy from an operator's notes
- a hardware session ends without a written result
- two documents disagree about what worked on the robot
“下面的recovery都不行 只有recovery_v3_1p1c.onnx”
RL系统/FOLLOW THIS copy 2.md § #### Recovery Policy (operator runbook, undated) A hardware run without its log is an anecdote - the first real get-up's policy, gain profile and log were never recorded, two CSVs stayed "to be reported", and runbook commands wrote different policies' logs under one copied filename
hardware-log-is-the-attribution-inputMake the log part of the run: name it from the policy and conditions automatically (never by hand-copied filenames), record the policy digest and gain profile inside it, include what the open questions need (torque, joint positions and targets), and treat a session without a collected log as incomplete.
Symptom
The recovery line's oldest open question - whether Isaac or MuJoCo reads torque demand correctly - was waiting on real-robot logs that never arrived, and the verdicts that did arrive could not be tied to files.
Context
deploy_policy writes a CSV per run (--log); the runbook's own analysis snippet reads its joint-position, target and action columns (q_, tgt_, act_), and the recovery hanging checklist asks for torque and joint logs for the whole run, to be compared with simulation. The first real get-up (08-09): policy, gain profile and log "to be recorded later". The first real A/B (08-11): v2_5b's result and both policies' CSVs "to be reported". In the runbook's walking commands, three runs of two different c4 policies log to real_s1e_pw08_teleop_0808.csv, and s1e and s2e_fric runs log to real_c2_700_pw08_teleop_0808.csv - filenames copied from other commands.
Change
None recorded; the spec kept listing the open-loop comparison as waiting for real logs.
Outcome
No real-robot log appears in the recovery spec through §50, so the simulator disagreement stayed unresolved and hardware verdicts stayed unattached to data.
Mechanism
Attribution needs the run's identity (policy digest, profile, conditions) and its signals in one artifact; a filename copied from another command mislabels the file, and a log not collected at the session is rarely collected later.
Applies when
- planning a hardware session whose result should settle a sim question
- log filenames are typed or pasted by hand
- hardware feedback arrives as prose without files
“python tools/deploy_policy.py --policy train/policies/omni_c4_ff800_pj.onnx … --log train/real_logs/real_s1e_pw08_teleop_0808.csv … q,t,a=d[:,c("q_")],d[:,c("tgt_")],d[:,c("act_")]”
RL系统/FOLLOW THIS copy 2.md § Walk 遥控 / S2 / csv 分析片段 (operator runbook, undated) A stand gate judged by survival passes a robot that wanders a meter - judge posture instead
stand-gate-posture-not-survivalFor every gate, ask what behavior the metric is a proxy for and validate its ordering against real observations; replace metrics whose ordering disagrees with reality, and demote them explicitly rather than silently.
Symptom
Real robot "standing" drifted 0.5-1.4 m across the floor while the sim stand gate scored a clean 20/20 - because the gate's metric was episode survival, which wandering does not violate.
Context
C2's stand condition was originally written as "survival regression <=2/20". A real-robot counter-example on 2026-08-08 forced the re-judgment: wandering robots survive. Cross-checking candidate sim metrics against real-robot feel showed max-tilt median ordering agreed with hands-on ranking, while displacement ordering was actually OPPOSITE to real impressions - so displacement was demoted to a reference quantity, not a gate.
Change
Stand PASS criterion rewritten from survival to posture: "stand tilt max median <= root baseline +1.5 deg"; displacement kept only as reference. Applied to all subsequent rungs (C4 and redo levels inherit it).
Outcome
Later rungs gated stand on tilt (e.g. C4 product: 7.7 deg vs parent 7.1 deg, +0.6 deg PASS); the wandering failure mode became visible to the battery instead of hidden by survival.
Mechanism
A gate metric is a proxy for an intended behavior; survival is a proxy for "did not fall", not "stood still". Metric choice must be validated against ground truth (real-robot feel/measurement), and a proxy whose ordering disagrees with reality on real data is worse than no metric - it steers selection backwards.
Applies when
- writing PASS conditions for stand/idle/hold behaviors
- a gate passes policies that visibly misbehave on hardware
- choosing between candidate metrics for an acceptance battery
“stand 必须用位姿判,不能用存活判(2026-08-08 真机反证改判):站着乱走 0.5~1.4 m 时存活照样 20/20 —— C2 的 stand 条件原写「存活退化 ≤2/20」,选错了指标。改为: stand 倾角 max 中位 ≤ 根基线 +1.5°(倾角与真机手感排序一致;位移排序与真机相反,降为参考量)”
train/C_LADDER_RUN.md § 3b. PASS 条件 ⚠️ stand 必须用位姿判 Removing a foot-spacing wall passed every simulated gate and made the feet collide on the real robot - nothing priced stance width in the two-foot phase, the policy narrowed to the simulator's self-collision floor, and real calibration offsets closed the last millimetres; the wall came back with a gate
removed-wall-returns-on-hardwareWhen a constraint is removed, name what will govern that quantity instead and add a gate for it; never let a simulator's collision floor be the margin, and when a gate is exceeded by a hair, record the exact numbers and hand the release decision to a person instead of quietly passing it.
Symptom
On the first real-robot try of oneleg_v0 (2026-09-16) the two feet collided; the user also judged the folded foot not high enough.
Context
The V0 reward table had dropped the feet_lateral_distance wall because it seemed to conflict with the hip adduction single support needs. In the two-foot command bucket no remaining term governed stance width, so the policy drifted narrower until the simulator's self-collision stopped it; the sim acceptance had no foot-spacing gate, so 40/40 said nothing about it. On the robot, calibration offsets consumed the margin.
Change
V0.1: the wall restored (-10, minimum 0.16 m), re-checked against measured numbers (a swing-phase lateral spacing of ~148 mm costs 0.12 per step, acceptable); fold weight 0.8 -> 2.0; a ninth gate: minimum foot spacing >= 100 mm and zero leg-contact frames. The removal was kept on record.
Outcome
oneleg_v0_1 (V0r2 model_2200) passed 39/40 with the spacing gate 40/40. The single miss (a 15.4 deg tilt transient against a < 15 deg limit during a side switch, steady 6.9 deg, everything else green) was recorded with its numbers and released for the user to overrule.
Mechanism
An unpriced degree of freedom drifts to wherever the simulator stops it; if that stop is the simulator's own collision model, the policy's margin on hardware is whatever the calibration error leaves.
Applies when
- dropping a reward term that looked redundant or conflicting
- hardware shows a failure no simulated gate measures
- a release candidate misses one gate row by a small amount
“V0 撤墙被真机证伪(2026-09-16):双脚桶没有任何项管站宽,策略贴 sim 自碰撞底线收窄,真机标定偏差一吃**双脚相碰**。 … min ≥ 100 mm 且腿碰 0 帧(eval_straight 同判据)—— … V0 真机双脚相碰暴露 sim 门未看脚距的缺口 … L s2 标称 tilt 瞬态 15.4°(门限 <15, 超 0.4°, 稳态 6.9°, 该跑其余全绿)——换侧瞬态蹭线, 判定放行留档, 用户可否决。”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §4 feet_lateral_distance 行 / §6 验收门 ⑨ / §8 核查单 7 A 2-degree joint-zero calibration fix moved the whole runnable envelope - re-test old "cannot run" verdicts after recalibration
zero-offset-calibration-shifts-envelopeDate every hardware verdict with the calibration state; after any zero/mount recalibration, re-test previously condemned policy-power combinations and previously "unexplainable" posture offsets before attributing either to training or model.
Symptom
s1d was on record as "only runs at power 0.7" (kicked wildly at 0.8); after a calibration pass, the same policy ran 12 s at 0.8 with no kicking at all.
Context
The calibration had fixed a 2.08 deg zero offset on r_hip_roll - exactly the constant error source on the dominant joint of the kicking oscillation loop ("恰是乱踢振荡环主导关节的常值误差源"). The three-generation post-calibration hardware sweep also closed a second case: the robot's mysterious "backward lean" disappeared after calibration, and the sim-real posture difference collapsed from opposite-sign 5+ deg to same-sign ~2 deg ("后仰案实质了结") - the lean had been a sensing/zero artifact, not a mass-model error. Booked consequence: if the s1d recovery re-verifies, "真机可跑档整体 上移" - every policy's runnable power envelope shifts up, and downstream lineages' hardware expectations get revised.
Change
Joint-zero and mount calibration promoted from setup chore to a variable that dates hardware verdicts: verdicts about which power/scale levels a policy can run are conditioned on the calibration state they were measured under.
Outcome
One policy rehabilitated at a higher power level; one standing sim-real posture discrepancy closed without touching model or training; a pending re-verification booked rather than asserted.
Mechanism
A constant joint-zero error acts as a persistent disturbance injected at the feedback loop's most-loaded joint; near an oscillation threshold, removing a 2-degree bias is the difference between a stable and an unstable loop. Since the error is additive and machine-side, it shifts every policy's stability envelope simultaneously - which is why verdicts must carry their calibration date.
Conflicts
The s1d rehabilitation awaited one confirming re-run at the time of writing ("待复核一跑坐实") - the offset-as-cause reading is the head suspect, not a closed verdict.
Applies when
- a policy oscillates at a power level others tolerate
- sim and real disagree on a constant posture offset
- deciding whether to re-test old hardware verdicts after maintenance/calibration
“发现①:s1d@0.8 能跑了(旧账「只有 0.7 能跑」)——12s 无乱踢。头号嫌疑 = 标定修正:r_hip_roll offset 修 2.08°,恰是乱踢振荡环主导关节的常值误差源。… 发现②:「后仰」标定后消失 … sim-real 姿态差从反号 5°+ 收敛到同号 2°,后仰案实质了结。”
train/README.md § 真机 @0.8 三代横评(2026-08-07 标定后) A sim veto needs real confirmation too - the worst sim cell was scheduled as the most informative hardware run
sim-veto-needs-real-confirmationNever let sim alone both condemn a purpose-built configuration and escape audit: spend one cheap, safeguarded hardware run on the condemned cell, pre-registering what agreement and disagreement would each imply about the proxy.
Symptom
The fric-2400@kd1.0 combination was sim's worst cell across the board (survival 17/20 - the only miss, mu0.4 1/20, push 103/160, zero-cmd 2/20), yet it was the only product specifically trained for the kd1.0 deployment gain - discarding it on sim evidence alone would leave the sim's own validity untested exactly where it mattered.
Context
The team had been burned in the other direction before ("Isaac 指标三次 零预警" - training-side metrics gave zero warning three times), so the symmetric rule was written: sim's rejection also needs hardware confirmation ("sim 判被支配 ≠ 真机被支配 … sim 的否决也要真机确认"). The run was pre-registered with a dual reading: real matches sim -> the S2f ladder closes and the fork root is settled; real clearly better than sim -> the MuJoCo proxy has a systematic bias in the kd1.0/low-margin region, "那比选型本身重要得多" - and every S2f sim acceptance would need re-scoring.
Change
The condemned configuration was kept on the hardware roster (last, spotted, minimal exposure) explicitly as a proxy-validation probe, not as a deployment candidate.
Outcome
Session design captured either result as progress: selection confirmed, or a proxy bias discovered that would re-price the whole ladder's verdicts.
Mechanism
Every sim verdict is a joint statement about the policy AND the proxy; cells where a policy was purpose-trained for the exact condition sim condemns are where proxy error is most likely and most costly. Testing the veto converts a selection decision into a calibration measurement of the evaluator itself.
Applies when
- sim rejects the configuration that targets the actual deployment condition
- the eval proxy's calibration has never been checked in that regime
- deciding which hardware runs are worth their risk
“但它也是唯一为 kd1.0 部署档专门训的产物 —— sim 判被支配 ≠ 真机被支配, 「Isaac 指标三次零预警」的教训反过来同样成立: sim 的否决也要真机确认。… 若真机明显好于 sim → MuJoCo 代理在 kd1.0/低裕度区有系统性偏差, 那比选型本身重要得多。”
train/REAL_RUN_S2.md § 上机名单 note / 2. sim 侧预注册预期 ⑤ Bracket a real-robot A/B with a repeated reference run - battery drain is the confound
battery-bracketed-real-abOrder hardware A/B sessions as A-B-A: repeat the first condition at the end, and void the comparison if the bracket runs disagree - never let battery or venue drift ride on the second condition.
Symptom
In a two-policy teleop A/B on hardware, the second policy is measured on a lower battery voltage than the first - a systematic bias that would be read as a policy difference.
Context
C2 real A/B (checkpoint 700 vs A800, same floor, same day) was scripted as 700 -> A800 -> 700-rerun, with the explicit note that a teleop session drains the pack and the trailing policy "naturally suffers".
Change
Protocol: run the reference policy first AND last; if the two reference runs differ noticeably, declare the whole session battery/floor-polluted and void the A/B ("结论作废重来"). Also log electricity per run.
Outcome
Called out as the round's only systematic confound, closed by one extra command ("这是本轮唯一的系统性混淆源,一条命令就能堵掉").
Mechanism
Battery voltage scales available torque, and torque loss hits behavior asymmetrically (see power-scale-hurts-nonforward-axes), so drain masquerades as policy regression; a head/tail reference pair converts the unobserved drift into a measured control.
Applies when
- comparing two policies or settings on hardware in one session
- any sequential hardware evaluation where the plant drifts (battery, temperature, floor wear)
“为什么要 700 复跑:一次遥控 session 下来电池会掉压,第二枚天然吃亏。头尾各跑一次 700,若两次 700 明显不同,说明这轮 A/B 被电量污染,结论作废重来。这是本轮唯一的系统性混淆源,一条命令就能堵掉。”
train/C_LADDER_RUN.md § 3c. A-3 真机 A/B(同一段地板、同一天、电量记账) Real-robot trials of a new skill were staged by risk - a hanging dry run with the robot posed by hand, then one short try per category on a mat with the hardest last, then the composed behaviour (switch + walking) last - with the user present and a log every time
staged-hang-mat-floor-for-get-upStage a new skill's hardware trials by risk - hanging dry run posed by hand, one short try on a mat per start category with the hardest last, the composed behaviour last - with an operator ready to cut enable and a log for every try; relax a safety ban only for short, attended runs and say so in writing.
Symptom
A get-up policy acts violently near the ground by design, and the first unstaged real run of the line was stopped as dangerous.
Context
The hanging checklist written with the first stamped recovery product (v2_5, 2026-08-11), to be ticked item by item with the user present: both machines on the same commit and firmware torque limits checked; the robot hung from a single point about 0.1 m off the ground; a dry run with the robot posed by hand into supine and prone to watch that the target stream is gentle (the beta contract keeps targets within +/-0.25 rad of the measured pose, so enabling causes no homing fling); the first floor try is supine only, on a mat, once, with torque and joint logs; categories are added one at a time, prone last; any kicking or oscillation cuts enable immediately. For the switch (08-14) the runbook orders: hang with the standing policy as the locomotion side, then on a mat push the robot over and let it recover, and only last swap in the walking policy. An exemption was also written: edge-standing policies stay banned from long or unattended runs, but a short single A/B with the user present, hung or on a mat, is allowed. The one-leg line reused the same order (hang, then floor with a spotter, 60 s segments with a temperature check).
Change
Real trials as a checklist of stages, each gated on the previous one, with the composed behaviour last.
Outcome
The line's first real get-up (v2_6, 08-11) came through this protocol and was reported "fairly stable"; no further hardware outcomes of the switch are recorded in the spec.
Mechanism
Each stage exposes one new risk (commanded targets without contact, a single category with contact, harder categories, then the interaction of two policies), so a failure is attributable and cheap.
Applies when
- first hardware trial of a recovery, jumping or other high-impact skill
- switching between two policies on hardware for the first time
- a policy with a known posture defect needs a comparison run
“吊挂空跑: 手动摆到 supine/prone 姿态, 看目标流是否温和 (β 帽 7.5 N·m, 目标永远贴着当前 q ±0.25 rad —— 使能瞬间无归位甩动, 这是 β 契约附带保证) … 落地首试: supine 一类, 垫子, 单次; τ/q --log 全程记录 … 逐类别扩展 (prone 最后), 每类先单次”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §40 吊挂执行单(真机首试;需用户在场,逐项打勾) The real robot's right-leg kicking was over-trained-delay times loop gain - irreducible pipeline latency is plant, model it fully from day one
pipeline-latency-is-plant-not-drMeasure the end-to-end action pipeline delay and build it into the nominal plant and every acceptance gate from day one; treat power/scale deratings that "fix" oscillation as gain-reduction crutches flagging an unmodeled delay, and expect higher-feedback-gain policies to be MORE delay-fragile.
Symptom
On hardware, s1c/s1d at action scale 1.0 always kicked wildly with the right leg (s1c only ran as SOTA at power 0.8; s1d only at 0.7) - while sim showed nothing under default evaluation.
Context
Sim reproduced the incident item by item once the real pipeline delay was injected: s1d@1.0 with --delay 1 fell at 10.2 s, --delay 2 at 5.2 s; s1c@1.0 stressed (r_hip_roll saturation 5 -> 16%; "右脚" = the policy's chirality makes the right leg its high-gain leg); and the combos that worked on hardware (s1c@0.8+delay2, s1d@0.7+delay2) all survived in sim. Mechanism: the real pipeline is ~1-2 ticks (BusWorker next-cycle pickup + CAN round trip) but S1.1 trained only to 1 tick - "超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定" (delay beyond training x full loop gain = oscillation; the power derating had been buying stability by compressing loop gain). s1d was MORE fragile than s1c because its yaw 3-layer stack had learned higher feedback gain - higher gain, lower delay tolerance. Three changes: latency DR widened to cover reality; acceptance gates and smoke runs moved permanently to --delay 2 ("门必须在真机条件下预测 真机"); and the doctrine written twice-paid: "不可约的管线属性(延迟、 限速)不是'随机化选项',是 plant 本体,第一天就该全额建模" - S1's nominal-then-robust staging falsified by hardware for the second time. The later s1e hardware run at power 1.0 (no kicking, normal force) closed the loop: "0.8 = 旧代拐杖" - the derating had been a crutch for the under-modeled delay, not a real requirement.
Change
Latency modeled as plant from day one of any lineage (measured 1-2 ticks covered, bridge-layer rate limits likewise modeled by default); every gate and smoke evaluation issued under --delay 2.
Outcome
Kicking reproduced, explained, and eliminated in the s1e generation at full scale and full power; the deploy-side crutches (0.7/0.8) retired for the new lineage.
Mechanism
Feedback oscillation onset is a product of loop gain and phase lag; a policy trained below the real delay learns gains that sit past the real stability margin, and any output derating masks it by scaling gain down. Since pipeline delay is deterministic hardware property - not an uncertainty - it belongs in the nominal plant, and every evaluation must include it or the gate predicts a robot that does not exist.
Applies when
- hardware oscillation/kicking that sim only reproduces with added delay
- a policy only runs on hardware at reduced power/scale
- defining what belongs in the nominal plant vs the DR list
“真实链路延迟 ~1~2 拍 … S1.1 只训到 1 拍——超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定。s1d 比 s1c 更脆 = yaw 三层栈学出更高反馈增益,增益越高延迟容忍越低。… 教训入账:S1「先标称后鲁棒」第二次被真机证伪——不可约的管线属性(延迟、限速)不是"随机化选项",是 plant 本体,第一天就该全额建模。”
train/OMNI_V0_SPEC.md § 3. S1.4(真机右脚乱踢事故强制) Push-test protocol - positive side first, fragile side spotted, axes aligned in the log, and cross-machine push counts stay qualitative
push-test-chirality-protocolOrder disturbance tests from the robust side to the fragile side with protection scaled to sim-measured asymmetry, align and log frame conventions before testing, and treat cross-domain disturbance counts as qualitative evidence only.
Symptom
Hand-push testing on hardware risked falls on a side sim had already flagged as fragile, and push counts invited apples-to-oranges comparison with sim numbers.
Context
Sim chirality was explicit: descendants were far more fragile in -y (fric-3000@kd1.2: +6 N*s survived 15/20 vs -6 N*s only 3-9/20) while the s1e control was perfectly symmetric (40/40). The protocol therefore: push the positive direction first, keep a spotter for the negative side; before any push, record which real-robot side corresponds to sim's +y in the log ("上机前对一次坐标"); and - citing the chaos lesson ("混沌课文") - real push results are used only as qualitative corroboration, never compared numerically with sim survival counts across machines.
Change
Push testing became a scripted, chirality-aware protocol with frame alignment as a logged precondition and an explicit epistemic limit on cross-domain count comparison.
Outcome
The fragile side was tested with protection informed by sim's quantified asymmetry; logs stayed interpretable because the frame correspondence was recorded before the first push.
Mechanism
Disturbance-response chirality is a real, quantifiable lineage property, so test order should follow measured fragility; and perturbation outcomes are chaotic in the details (divergent trajectories from tiny differences), so counts do not transfer across domains even when qualitative rankings do.
Applies when
- planning push/disturbance tests on hardware
- sim shows directional asymmetry in disturbance survival
- someone proposes comparing real push counts to sim counts
“先正向后负向, 负向留人扶 —— sim 手性明确: 后代在负 y 向显著更脆 (fric-3000 @kd1.2: +6 N·s 15/20 vs −6 N·s 3~9/20), 而 s1e@0.8 两向 40/40 完全对称。上机前对一次坐标 … 跨机不做二值结论 (混沌课文): 真机推力只作定性对照, 不与 sim 计数对比。”
train/REAL_RUN_S2.md § 3. 抗推 (可选, 人手推; 做则按此协议) action_rate weight is the sim2real bandwidth knob - re-tune it whenever a rate limiter is removed
action-rate-weight-vs-bandwidthSet action_rate weight relative to real actuator bandwidth, and re-tune it any time another smoothing/limiting element (filter, slew limiter, gain) changes - reward weights are load-bearing parts of the actuator model.
Symptom
With a low action_rate_l2 weight the policy learns fast actions; the unmodeled part of the actuator response is then excited hardest, and sim2real "直接崩" (collapses outright). With too high a weight, actions become so slow the robot cannot maintain balance.
Context
The reference developer called action_rate_l2 the single most important reward for transfer, with side-by-side video evidence that the high-penalty, slower policy is clearly better on hardware. Lucen context: the team had just removed the SOFT_SPD=1.0 velocity limiter, which had been an implicit actuator-bandwidth constraint - leaving action_rate as the only remaining constraint on action speed.
Change
Decision recorded: after removing SOFT_SPD, re-evaluate the action_rate weight rather than keep the old value, since its effective role changed from "additional smoother" to "sole bandwidth constraint".
Outcome
Logged as a priority follow-up ("重新评估 action_rate 权重 - 拆掉 SOFT_SPD 之后这一项的作用变了"); the failure mode it guards against is training high-frequency actions the real actuators cannot track.
Mechanism
Slower actions stay inside the frequency band where the ideal-PD sim actuator and the real actuator agree; fast actions probe the band where unmodeled delay, inductance, and bandwidth limits dominate, so model error is amplified in exact proportion to action speed. Any removed external rate limit transfers that constraint's entire job onto the action_rate penalty.
Conflicts
The low/high tradeoff evidence is the external developer's report (with video); the Lucen-side entry is a pre-registered risk and decision, not yet an on-robot A/B at the time of writing.
Applies when
- removing or adding an action filter, slew limiter, or low-level speed cap
- real robot shows high-frequency chatter or overheating absent in sim
- tuning smoothness rewards before a hardware deployment
“权重低 → 动作快 → 执行器模型不准的部分被放大,sim2real 直接崩 / 权重高 → 动作慢 → 好迁移,但可能慢到无法维持平衡 … 我们刚拆掉 SOFT_SPD=1.0 的限速器,等于把执行器带宽约束整个移除了。action_rate 惩罚现在是唯一还在约束动作速率的东西,需要重新评估权重”
Experience.md § action_rate_l2 是他认为最关键的 reward (lines 61-70) When hardware underperforms, audit deployment knobs before prescribing retraining
deploy-knob-attribution-before-retrainingBefore any "retrain it" decision, reproduce the symptom in sim under the exact deployment configuration; if the symptom follows the deployment knob rather than the checkpoint, fix the knob or randomize it in training - never top-up-train the skill.
Symptom
Real-robot feedback after the C4 deployment - "turning is weak" - with two retraining options on the table: top up turn training, or restart from the s1e root.
Context
The sim account showed the policy turned well (75-81% at pw1.0); the robot was deployed at power-scale 0.8. The 3-6 pp difference between C2 and C4 policies at the same power was noise; the 40-50 pp difference between power levels was the entire effect. Both proposed retraining paths would have burned budget on a non-existent training gap, and restarting from s1e would additionally have discarded the sidewalk skill that took four rungs and a coordinate-bug hunt to obtain.
Change
Decision: retrain nothing. (1) Try pw1.0 on hardware first - sim says net gain; (2) only if 1.0 is unacceptable (heat/feel), the correct training fix is power/torque randomization in the S2 plant line (one variable, fixes turn and backward together) - not skill top-up; (3) restart-from-root explicitly ranked worst.
Outcome
The "weakness" was fully explained by the deployment knob; the sim/real signatures matched the earlier power-derating law verbatim ("与 C2 时代 power 衰减主要伤非前进轴 逐字吻合").
Mechanism
The policy's competence is defined under its training plant; deployment knobs (power scale, teleop mapping, command bands) silently define a different plant. Attributing a deploy-plant effect to a training gap produces exactly the wrong fix - more training on the wrong variable.
Applies when
- real robot underperforms a skill that sim says is fine
- proposals on the table include retraining or re-rooting
- deployment uses any override the trainer never saw (power scale, remapped commands, different control rate)
“正确的训练修法不是补训转向,而是训练时加 power/力矩随机化让策略在 0.8 下自己补偿 —— 单变量,属 S2 plant 线,一次同时修好转向与后退;从 s1e 重训是最差选项:丢掉四轮 + 一个指标 bug 才换来的侧走,而 C2 的转向本来就没问题。”
train/C_LADDER_RUN.md § 3p. 三 处置顺序(回答「补训转向 还是 回 s1e 重训」:都不该) A real-robot verdict is (policy x deployment stack) - when the stack changes materially, old verdicts expire
stale-verdicts-under-old-stackDate every hardware verdict with the deployment-stack version it was measured under; after any material stack change, re-test before trusting old condemnations or old praises - with the interpretation of each possible result written down first.
Symptom
walk_v5 stood condemned as "kicks wildly" and walk_v6 as "cannot walk unassisted" - but those verdicts were issued under an earlier deployment stack (--heading did not exist yet, several fixes had just landed); only v7 had ever run under the current unified stack.
Context
New sim evidence sharpened the doubt: a same-harness four-version sweep showed v6 was the HEALTHIEST archive at the real operating point (cmd 0.15: dominant frequency locked at 2.50, steps 35:35 perfectly symmetric, foot distance 203/190 mm best of four, mean tilt 4.2 deg, saturation 15%). A full re-test under the unified stack was scheduled with per-version questions and a pre-filled interpretation table ("判读表(预填假设,回来对号)"): e.g. v6 walks + frequency ~2.5 -> old verdict was the stack's fault, v6 becomes the comparison champion; v6 walks but at ~1.25 -> period-doubling on hardware = confirmed plant gap, actuator fitting promoted to mainline; v5 no longer kicks -> the kicking was an old-stack artifact.
Change
All four versions re-queued on hardware under one stack (same torque limits, slew profile, heading loop, logging), with the version-specific legacy profile pinned; verdicts held provisional until re-issued.
Outcome
The re-test design separated policy properties from stack artifacts before any policy was permanently written off - and turned each outcome into a specific conclusion via the pre-filled table.
Mechanism
A deployed behavior is produced by the policy plus everything between it and the motors (heading loop, slew limits, torque caps, clock); verdicts implicitly condition on that whole stack. Fixing the stack invalidates the conditioning, so old failures may be stack artifacts and old successes may not survive either.
Applies when
- deployment tooling (limits, filters, loops) changed since a policy was last judged
- deciding which historical policy is the rightful baseline
- a sim sweep contradicts an old hardware verdict
“只有 v7 在完整的今日部署栈下上过真机 … v5"左右乱踢"、v6"未能自主"的判决全部来自更早的栈(--heading 尚不存在, 部分修复刚落地)——判决已过期。且 2026-08-02 四代同机仿真横测翻出了新证据:v6 在真机工况(cmd 0.15)下是四代里最健康的仿真档案”
train/REAL_SWEEP_V5_V8.md § 0. 为什么重测 A constant-value plant rung passed every binary gate with record scores - and shipped 60% thinner posture margins that hardware exposed
constant-value-dr-overfits-marginRandomize deployment-critical axes over a narrow band spanning the measured real support - never a single value, never a fictitious tail - and report graded margin quantities (tilt margin) next to binary gates, because saturated gates rank thin-margin and thick-margin policies identically.
Symptom
s2_lag1 (trained at constant 1-frame latency) posted the strongest sim gate sheet in history (20/20 everywhere) yet was unstable on hardware, while s1e (trained across the full 0-3 frame band) was the every-run-stable SOTA at the same power.
Context
The sim autopsy (new --delay-jitter harness modeling the BusWorker's time-varying phase drift): 18 runs across constant and time-varying delays ALL survived - time variation alone does not kill - but the tilt-margin ordering reproduced hardware exactly: s1e 7.7-9.1 deg (thickest) < s2_lag1 10.7-15.0 < s1c 16.2-18.2. Attribution: constant-value training permits precise specialization to that one value; s1e's band diversity forced cross-value robustness - "恒定 1 帧训练 vs s1e 的 0~3 帧全带——分布多样性逼出跨值鲁棒,恒定值允许精确 特化" - so the constant-rung policy's margins were ~60% thinner, fine in sim's clean world, pushed over the line by real-world disturbances. Tool lesson booked: "存活门二值饱和后掩盖裕度差" - binary survival gates saturate and hide margin differences; graded margin columns (tilt-max) belong in the report. The synthesis with the opposite failure (wide tails cause drag-glide): the proposed resolution was a NARROW uniform band (0.02, 0.04) covering exactly the real 1-2 ticks - diversity inside the measured support, no tail, no single point. s1e's root selection later leaned on the same property: its full-band latency training "预装" the delay rungs and delivered "全工况稳定裕度" that survived power derating.
Change
DR-on-an-axis design refined to a three-way distinction: no wide fictitious tails (drag), no single constant values (thin margins), but a narrow band spanning the measured real support; acceptance reports gained graded margin columns alongside binary gates.
Outcome
The tilt-margin column entered the standard report; the s1e root (band-trained) carried the C ladder while the constant-value branch was archived with its three contributions credited.
Mechanism
Robustness margins are shaped by the diversity of the training distribution, not just its support: a point-mass distribution lets the optimizer trade margin for on-point performance, while a band forces solutions that keep margin across the band - and binary survival metrics cannot see the difference until the margin is spent on hardware.
Conflicts
The narrow-band (0.02,0.04) resolution was a pending recommendation ("裁决建议(待用户)") at the time of writing; the lineage instead moved root to s1e whose full-band training predated the staged ladder - the deterministic-staging card and this card record the two failure modes the final design must avoid simultaneously.
Applies when
- a rung trained at a fixed plant value aces sim but wobbles on hardware
- binary acceptance gates are all saturated across candidates
- choosing between constant, banded, and wide DR on one axis
“18 跑全活,时变性单独不足以击杀;但 tilt_max 裕度排序完整复现真机:s1e 7.7~9.1°(最厚)< s2_lag1 10.7~15.0 … 恒定 1 帧训练 vs s1e 的 0~3 帧全带——分布多样性逼出跨值鲁棒,恒定值允许精确特化 … 存活门二值饱和后掩盖裕度差(s2_lag1 sim 门 20/20 史上最强却真机不稳)”
train/README.md § s2_lag1 真机不稳 × s1e 稳的 sim 对拍(2026-08-07,时变延迟实验) Privileged signals (true velocity, foot force, foot height) go to the critic only
observation-honesty-critic-onlyTreat the actor observation vector as a hardware contract: every element must exist on the real robot with realistic noise; privileged simulator truths belong in the critic only.
Symptom
Policies trained on ground-truth base linear velocity work in sim and fail on hardware, where only a drifting IMU and encoders exist - the policy has learned to depend on a signal that does not survive deployment.
Context
Many open-source locomotion stacks feed simulator ground-truth linear velocity to the actor. The reference team refused: the real robot has no ground-truth velocity. Asymmetric actor-critic keeps the training benefit of privileged information without deploying the dependency.
Change
Route ground-truth velocity, foot contact forces, and foot heights to the critic only; the actor observes exclusively signals that exist on hardware (IMU-derived quantities, encoders, commands, previous actions).
Outcome
Recorded as adopted doctrine in Lucen's experience log; the trained actor's input contract matches what the real robot can actually produce.
Mechanism
The critic is discarded at deployment, so it may consume any privileged state to reduce value-estimation variance; the actor's observation set is a deployment contract - anything in it that hardware cannot supply (or supplies with different noise/drift) becomes a train/deploy distribution shift the policy was never trained to handle.
Applies when
- designing actor/critic observation spaces
- reviewing a config where the actor sees base_lin_vel or contact forces
- sim policy is strong but real robot drifts, oscillates, or falls without obvious actuator cause
“很多开源代码库把真值线速度喂给策略,Asimov 团队没有,因为真机上没有真值速度,只有会漂的 IMU 和编码器;用完美速度训练出来的策略会依赖它,然后在硬件上失效。真值速度、足底力、足高统统只给 critic”
Experience.md § 观测空间的诚实性 (line 7) Run acceptance under measured contact parameters - honest condim/torsional-friction flipped a false PASS into a real-matching FAIL
eval-plant-honesty-contact-paramsPin the evaluation plant's contact model to measured values (contact dimension, torsional/rolling friction, mu) before trusting any gate that involves slip, impact, or drift - a gate can only fail a policy for physics its simulator contains.
Symptom
walk_v5 passed the old acceptance battery yet failed on the real robot (footfall force, drift, kicking) - the evaluation plant was flattering the policy.
Context
The battery was re-run under "honest contact parameters" - condim 4 (adding torsional contact) with measured torsional friction 0.035 - and v5 then FAILED exactly the rows corresponding to its real problems: heading 185 deg (limit 30), support-foot yaw slip 284 deg (limit 80), landing force 1.72x (limit 1.5x), suspended tilt 45.9 deg (limit 10). The slip physics depends on torsional friction, which the default contact model (condim 3) does not even simulate - a slip problem is invisible to an evaluator that cannot represent yaw friction at the foot. Term-sizing measurements for the new rewards were likewise taken under the same honest parameters (cmd 0.45, skipping the 5 s start transient).
Change
Acceptance harness pinned to condim 4 / torsion 0.035 (measured); verdicts issued under defaults declared non-citable for these rows.
Outcome
Sim acceptance verdicts began agreeing with hardware ("现在失败, 与真机一致"); the v6 fixes could be developed and validated against an evaluator that could actually see the disease.
Mechanism
An evaluator is a plant model too: contact dimensionality and friction values decide which failure modes exist in the simulation at all. Evaluating under default contact parameters tests the policy in a world where its real failure is physically impossible, producing structurally false PASSes.
Applies when
- sim acceptance passes policies that fail on hardware
- slip/drift/impact gates run under default simulator contact settings
- setting up a cross-simulator evaluation harness
“accept_v2.py 已加三条判据, walk_v5 在诚实的接触参数下(--condim 4 --torsion 0.035)现在失败, 与真机一致:直行 15s 航向累计 <30° | 185° ✗ … 落脚力峰值 <1.5× 体重 | 1.72× ✗”
train/WALK_V6_MINIMAL.md § 5. 验收 Order hardware runs by sim risk, gate each stage on the last, and put the fragile cell last with a spotter
risk-ordered-real-deploymentScript hardware sessions as a risk ladder: baseline first, sim-riskiest last with a spotter, suspended smoke before ground, each stage gated on the previous, environment (floor mu) recorded as a selection input - and stop at the stage that misbehaves.
Symptom
Five policy-x-gain combinations had to go on hardware in one session, with sim survival ranging from 20/20 down to 17/20 (and zero-command survival down to 2/20) - an unordered session risks breaking the robot on an avoidable run.
Context
The execution sheet fixed the order as sim-risk low to high, control baseline first (current SOTA establishes the floor reference), the fragile cell (fric-2400@kd1.0) last with a person spotting throughout. Stage gating: suspended smoke (feet off ground, 10 s each, all five pass before anything touches down) -> suspended with IMU and forward command (gait forms in the air) -> grounded runs -> speed raise only for combos that survived the previous stage -> zero-command tests only with a spotter, ordered by sim zero-cmd survival, with the 2/20 cell skipped by default. Preconditions include recording the floor material and estimating mu (if mu <~0.6, sim says pick the kd1.2 gain as main), port/CAN self-check, calibration frozen. Any stage failing stops the session at that stage: "任一段出问题就停在那一段, 不要跳到下一段".
Change
Session structured as a risk ladder with per-stage gates instead of a flat checklist; per-combo sim survival numbers written into the run table as the ordering key.
Outcome
The session design localized any failure to the cheapest stage that could reveal it, kept the robot safe for the informative fragile run, and made the control baseline available before any comparison run.
Mechanism
Hardware sessions consume a shared budget (robot integrity, battery, floor time); ordering by predicted risk means information is bought cheapest-first, and stage gates convert an expensive failure into a cheap earlier one. Baselines run first because every later reading is relative to them.
Applies when
- taking multiple policies/configs to hardware in one session
- a candidate is known-fragile in sim but must be measured
- writing a deployment runbook for a new robot
“跑序 = sim 风险从低到高, 最险的放最后 (依据 = 存活门/零指令存活) … ⑤ 是 sim 里最脆的一格 … 放最后跑, 全程留人扶, 起步即给 cmd, 零指令不做。… 任一段出问题就停在那一段, 不要跳到下一段。”
train/REAL_RUN_S2.md § 上机名单 / 全部命令 Ungated phase shaping made standing 42x more expensive than stepping - and the stepping was cooking the hip motors
moving-gate-42x-stand-taxGate every phase/clock-driven shaping term on the command that justifies motion, price the cmd=0 case explicitly during design - and when a reward flaw is sim-only, book it with trigger conditions instead of operating immediately on a working lineage.
Symptom
At cmd = 0 the sim policy never stood still - it stepped in place and crept 0.98 m per 20 s; on the real robot the same lineage's stepping made hip_roll motors run 20 degC hotter than every other joint (43-48 degC vs 25-28).
Context
The arithmetic closed it: the three phase-shaping terms (joint_pos_ref 1.6, feet_contact_number 1.2, feet_clearance_swing 1.6) are driven by the gait clock with NO command gating, so standing at cmd=0 forfeits 3.32/step of shaping while honest standing earns only 0.078 of tracking - stepping wins 42x. The heat chain: perpetual stepping = perpetual single support = one hip_roll stalled at ~4.3 N*m (25% of torque limit) carrying the torso's frontal-plane moment - two hip_rolls = 90% of whole-machine steady-state I2R; measured stand-vs-step comparison: total heat -71% when actually standing. The suspended test acquitted the actuator (0.21 rad sag -> 0.0008 rad in air) and mechanics vetoed the easy fix ("降 kp 救不了热" - equilibrium torque equals the external load regardless of kp). The fix (moving_gate: hard-gate the three shaping terms on |cmd| > eps) was designed - then DEFERRED by the user because the real robot at the time stood fine: "真机不表现该问题, 为真机不存在的病改奖励表不划算", with written trigger conditions (C-ladder stand row persistently failing, or real robot starting to step/drift at cmd=0) and the known hazard tag (this is exactly the reward-change class that triggers the B-arm signature). When the real robot later DID step and cook, the booked trigger fired and moving_gate moved from debt to to-do with its benefit re-priced: "停止空烧 hip_roll,稳态发热降 ~71%".
Change
moving_gate designed with the gate_by_cmd convention; deferral, triggers, and expected heat recovery all pre-registered instead of patching the reward for a then-sim-only symptom.
Outcome
The cmd=0 stepping went from mystery to closed arithmetic; the thermal measurement (hip_roll +20 degC) quantitatively confirmed the 90%-of-heat prediction; the reward change waited for real-world justification instead of spending a risky revision early.
Mechanism
Clock-driven shaping terms define a perpetual-motion bounty unless gated by command; the resulting idle gait is not a training bug but the table's optimum. Its cost surfaces on hardware as stall-torque heating set by statics (mass x lateral offset), which no gain change can remove - only removing the motion (gating) or widening the stance can.
Applies when
- the policy steps in place or creeps at zero command
- specific joints run hot in idle behaviors
- deciding when a known reward flaw justifies a risky mid-lineage fix
“塑形合计 −3.32/步 … 净: 站定亏 42 倍 … 两颗 hip_roll 4.29 / 4.09 N·m(各占限扭 25%),占全机稳态 I²R 的 90% … 吊挂实测 hip_roll 跟踪误差 0.21 rad → 0.0008 rad … 降 kp 救不了热 … 真机不表现该问题, 为真机不存在的病改奖励表不划算”
train/README.md § C1 FAIL 节 (cmd=0 的 sim/真机分歧记账) / C2 真机 A/B 三b 发热定性 Cross-simulator gate (Isaac Lab to MuJoCo) comes before any hardware attempt
sim2sim-gate-before-sim2realGate every policy through a second simulator with an independently built plant before hardware; treat sim2sim failure as a contract or overfitting bug, and sim2real failure after a sim2sim pass as a plant/actuator gap.
Symptom
A policy that only ever ran in its training simulator carries untested dependencies on that simulator's solver, contact model, and defaults; the first place those dependencies surface should not be the real robot.
Context
Standing order of operations for the whole Lucen program, recorded as the opening line of the experience log: train in Isaac Lab, gate in MuJoCo, only then go to hardware. The MuJoCo side is the same plant used for evaluation batteries, so a sim2sim pass also validates the exported policy + contract (obs ordering, scales, defaults) outside the training stack.
Change
Pipeline rule adopted - every checkpoint must pass the MuJoCo evaluation battery (sim2sim) before it is considered for real deployment (sim2real).
Outcome
Used generation after generation as the cheap filter; hardware sessions only ever received policies that had already survived a second simulator.
Mechanism
Two simulators disagree exactly where a policy is overfit to simulator-specific artifacts (contact softness, integrator, default parameters, obs conventions); a cross-sim transfer catches contract bugs and solver overfitting at zero hardware risk, so real-robot failures that remain are attributable to genuine plant/actuator gaps.
Applies when
- planning the path from training to first hardware trial
- exported policy behaves differently outside the training framework
- triaging whether a real-robot failure is contract vs plant
“先sim2sim - 从isaaclab 到mujoco / 再sim2real”
Experience.md § opening lines (1-2) Real robot walked at half the sim clock for two generations - resolved by racing a reward-side and a plant-side evidence line, not by guessing
period-doubling-evidence-raceFor a hardware-only pathology, refuse to guess: pre-register one probe per side of the sim2real boundary (can the reward mechanism change it on hardware? can fitted plant parameters reproduce it in sim?) and let the first positive result direct the next version.
Symptom
The number-one sim2real gap: on hardware v6/v7 stepped at 1.23-1.32 Hz - almost exactly half the 2.50 Hz gait clock they were trained and simulated at; sim never reproduced it, two generations running.
Context
Instead of committing training budget to a guess, v8 pre-registered two mutually controlled evidence lines and kept the clock OUT of the training variables: (a) reward-side - if the v8 saturation fix revives joint_pos_ref (the term that pins the gait to the clock), re-run hardware and see whether frequency returns to 2.5 Hz (hypothesis: v7's frozen actions meant NO reward was pinning the gait to the clock, and the real plant - with armature and friction making high frequencies expensive - slid down to the leg's pendulum natural frequency ~1.1 Hz); (b) plant-side - record suspended joint data (fit_actuator), fit armature/friction, load the fitted values into sim2sim and see whether the 1.25 Hz reproduces IN SIM. Decision rule fixed in advance: "谁先给出阳性结果谁定 v9 的方向 (奖励侧 vs plant 侧)" - whichever line goes positive first sets the next version's direction.
Change
Period-doubling excluded from the v8 change set; both diagnostic lines scheduled in parallel as non-blocking work; frequency reported factually in acceptance with no pass/fail attached ("倍周期是否消失 不设判定,它是 §9 的关键证据").
Outcome
The gap was routed into a decisive-experiment structure rather than a speculative retrain; the plant-side line pointed at exactly the unmodeled armature/friction that were later measured and installed as the plant baseline. Resolution (era-2c full-plant retest): the family had TWO causes - v8's low-speed period-doubling vanished once measured armature+friction were installed (1.30 -> 2.50 Hz, bifurcation-edge machine sensitivity), while v7's stood untouched at 1.20 Hz (saturation-freeze-driven policy property) - both evidence lines paid off, one per case.
Mechanism
A behavior appearing only on hardware has candidate causes on both sides of the sim2real boundary; changing training to fix it tests only one side per expensive cycle. Two cheap parallel probes - one intervening on the reward mechanism, one making sim reproduce the real behavior - localize the cause to a side before any training money is spent, and sim-reproduction of a real pathology is itself the strongest form of plant validation.
Applies when
- a gait pathology appears on hardware but never in any simulator
- deciding whether a sim2real gap is reward-side or plant-side
- tempted to change the gait clock/reward to chase a hardware symptom
“倍周期(真机 1.23~1.32 Hz ≈ 时钟一半,v6/v7 连续两代;sim 从不出现):两条证据线互为对照——(a)… 真机重跑看频率是否回 2.5 Hz(假说:v7 没有任何奖励把步态钉在时钟上,真机 plant 有 armature/摩擦、高频贵,自由滑落到复摆自然频率 ~1.1 Hz);(b)真机吊挂录 fit_actuator.py … 看能否在仿真里复现 1.25 Hz。谁先给出阳性结果谁定 v9 的方向。”
train/WALK_V8_SPEC.md § 9. 平行线 (倍周期) With no term on foot attitude the robot stood on the edges of its feet (ankle roll -26 deg) from R0.5 to V2.5b; pricing it fixed that and turned stance width into the next free variable - the user saw both on video before any metric flagged them
unpriced-foot-attitude-is-a-free-variableList every posture quantity the hardware cares about (foot attitude, stance width, yaw, knee flexion) and make sure each is priced by a term with real gradient at the observed error; pricing one exposes the next, so re-inspect the feet-level video after every change.
Symptom
Watching the v2_5 video the user said the ankle roll after standing looked very strange - the feet were not flat. The feet view showed the right foot standing on its outer edge; median ankle roll at t = 8 s was -/+26 deg (limit +/-35), mirrored. This was the physical form of the ankle_roll saturation criterion that had failed since R0.5.
Context
No term priced foot attitude: feet_contact_upright counts an edge contact as contact, and stand_pose's wide exp kernel (sigma^2 = 9) gives almost no gradient at 0.45 rad. HoST carries an "ankle parallel" term (+20); this table never had one. Edge standing was made a blocking precondition for hardware (continuous ankle load, unstable contact, wear).
Change
V2.6 (single variable): flat_feet = (|q_l_ankle_roll| + |q_r_ankle_roll|) x upright gate x height gate, a linear hinge with a 5 deg margin, weight -2, continued from V2.5b.
Outcome
Ankle roll -/+26 -> 5.3/5.2 deg, ankle_roll saturation 6.6% (< 10%), feet flat on the feet-view video; success 99.6%, re-falls 0-1%. Then the user watched v2_6: hip yaw constantly tense and the legs very close together. The numbers: hip roll -/+4.9/4.8 deg against a nominal 25 - feet flat and hips open 25 deg cannot coexist without ankle compensation, the new term taxed that compensation, and nothing priced stance width. "Foot attitude as a free variable" was fixed and "stance width became the new free variable" - which the real robot then exposed as splits.
Mechanism
An optimizer spends every posture degree of freedom no term prices; closing one reallocates the slack to the next unpriced one.
Applies when
- a standing or landing posture looks wrong on video while gates pass
- a saturation criterion keeps failing on one joint
- a new posture term was just added
“用户看 v2_5 视频:"起身之后 ankle_roll 非常奇怪,脚根本不是平着站立"。 … 机理:奖励表**无任何脚掌姿态项** —— feet_contact_upright 边缘接触也算触地, stand_pose 的 exp 核(σ²=9)对 26°=0.45 rad 梯度≈0。脚掌姿态是自由变量。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §41 预注册 V2.6:flat_feet —— ⑥ 老账的物理形态被用户目视锁定 Armature must be N^2 x rotor inertia, never 0 - measure it no-load
armature-n2-rotor-inertiaEvery geared actuator carries N^2 * I_rotor of reflected inertia at the joint; set armature from a no-load measurement, never leave it 0 and never guess it.
Symptom
Sim joints accelerate more easily than real joints; old MJCF had armature = 0 (rotor reflected inertia entirely unmodeled), a systematic sim2real gap on every joint.
Context
Original hand-written MJCF plant used armature 0. The team derived and then measured the correct value: torque needed at the rotor is I_rotor * N * alpha; after the N:1 gearbox the output shaft "feels" an extra N^2 * I_rotor of inertia. With a 9:1 reduction that is an 81x amplification of the rotor inertia, far too large to ignore.
Change
Set per-motor armature from no-load (motor out of the robot) measurement instead of 0: RS06 = 0.0070 kg*m^2, RS02 = 0.0032 kg*m^2, RS00 = 0.0015 kg*m^2. Landed together with measured friction as the first fully-measured plant parameter set.
Outcome
Plant parameters "第一次全部来自实测" (first time all from measurement); became the frozen plant baseline for all subsequent training generations.
Mechanism
Reflected inertia scales with the square of the gear ratio: the rotor spins N times faster than the joint, so its kinetic energy (and the torque needed to accelerate it) appears N^2 larger at the output. Omitting it makes simulated joints unrealistically fast/light, so policies learn action rates the real actuator cannot deliver.
Applies when
- building or auditing a simulation plant model for a geared/QDD actuator
- sim policy moves joints faster or snappier than the real robot can
- MJCF/URDF review shows armature or rotor inertia set to 0 or a default
“armature 转子反射惯量有问题 在sim里面一定要处理 不能是0,空机测试。转子处需要的力矩 = I_rotor × N × α 经减速箱放大 N 倍后 = N² × I_rotor × α 所以输出轴"感觉到"多了一个 N² × I_rotor 的惯量。 这就是 armature。… 关键是那个平方。减速比 9:1 就放大 81 倍。”
Experience.md § # armature 转子反射惯量有问题 (line 9) Slowing the gait clock at deployment is out-of-distribution and backfires - lower the commanded speed instead, or train the knob
cycle-time-override-is-oodAny deployment override must correspond to a dimension the policy was trained to handle; to make a parameter field-adjustable, randomize it in training and observe it - otherwise use the levers inside the trained envelope (commands) and leave the knob alone.
Symptom
Real-robot feedback "walks very fast and unstable" suggested slowing the gait; a deploy-side --cycle-time override existed, making "just slow the clock" a one-flag temptation.
Context
A sim sweep of the override on walk_v6 @cmd 0.3 showed monotone degradation away from the trained 0.40 s cycle: at 0.50 s tilt jumped 7.9 -> 13.2 deg and landing force 1.52x -> 2.24x; at 0.80 s (half speed) clearance collapsed to 3 mm - dragging again - with 20 deg tilt. Meanwhile the legitimate lever, lowering the commanded speed with the clock untouched, improved everything monotonically: cmd 0.1 gave 104% tracking, 6.7 deg tilt, minimum slip - the most stable operating point. The file distinguishes the two "slows" explicitly: lower command = smaller steps at the same 2.5 Hz rhythm; a slower rhythm itself requires retraining - randomize cycle_time (e.g. 0.40-0.65 s) during training and expose it as an observation, and only then does --cycle-time become a field-adjustable knob.
Change
Deployment guidance: never ship a cycle-time override the policy was not trained under; respond to "too fast/unstable" with lower commands; schedule clock variability as a training-time (contract-level) change if a field knob is wanted.
Outcome
The sweep quantified the trap before hardware paid for it (dragging and 2.2x landing force at slowed clocks); cmd 0.1 documented as the stable demo point.
Mechanism
The policy is a function fitted around the training distribution; a deploy-side override moves an input (phase rate) to values never seen, so behavior degrades unpredictably - the knob LOOKS like a capability because it exists in the code, but capability lives in the training distribution, not the interface.
Applies when
- a deploy tool exposes overrides (clock, scale, gains) beyond the training distribution
- hardware feels "too fast/aggressive" and a quick knob exists
- deciding between a deploy-side tweak and a retrain
“0.80s | 1.25Hz | 0.165 | 3mm(拖地) | 20.0° … 慢一半直接崩 … 策略按 0.40 训练, 别的周期属分布外。… 降指令速度才是有效杠杆 … cmd 0.1 是最稳的工作点。… 要节奏本身变慢必须重训 —— 训练期把 cycle_time 随机化(如 0.40~0.65s)并作为观测的一维, 部署时 --cycle-time 就成了现场可调的旋钮。”
train/WALK_DIAGNOSIS.md § 2026-08-01 追加: 调慢步态时钟(--cycle-time)在仿真里是反效果 The walking lines' safety setting, power-scale 0.8, broke the recovery policy's full-range contract - it cut the ends of the joint travel (4/50 could not get up) and left the torque spikes untouched; a kp x 0.9 gain profile inside the trained kp band did the job
power-derating-cuts-full-range-contractA deployment derating knob means something only relative to the action contract: before reusing a line's "safe setting" on a new skill, check what it does to that skill's reachable range and to the term that makes the spikes, prefer a gain change inside the band the policy was randomized over, verify it in simulation, and re-decide when the contract changes.
Symptom
After the violent first real get-up (2026-08-09), the recovery policy needed a gentler setting for its next hardware test, and the walking and omni lines' standard derating - deploying at power-scale 0.8 - was the obvious candidate.
Context
The V0 recovery contract maps actions to absolute targets over the full joint range: a = +/-1 lands exactly on the URDF limits, and standing puts the knee at the clip. Candidates were compared on R3.1 in MuJoCo (5 categories x 10 seeds) on 2026-08-10 before any hardware time was spent.
Change
A new gain profile, rl_kp090 (kp x 0.9, kd unchanged), recorded in robot.yaml as the recovery hardware-test setting, with power-scale 0.8 explicitly banned for recovery.
Outcome
kp x 0.9: 48/50 got up; median torque demand on hip_pitch/knee fell from 120-125% to 100-104% of the deployment limit; leg-leg contact frames 2,152 -> 1,095; the change sits inside the +/-10% kp randomization the policy trained with. power-scale 0.8: 4/50 could not get up, because under the full-range contract it removes the ends of the travel (the deep squat's tucked legs, the straight standing knee), and the torque spikes (kp x error) did not fall at all. When the line moved to the beta-anchored contract, rl_kp090 was declared a V0-era choice that does not fit (beta is calibrated at kp 30) and deployment returned to rl_default; the deploy switch applies power scaling to the walking side only.
Mechanism
A power scale multiplies the action, which under an absolute full-range mapping shrinks the reachable workspace instead of softening the actuator; the spikes come from the proportional term on large errors, which only a gain change reduces - and a gain change inside the trained randomization band stays in distribution.
Conflicts
The undated operator runbook still carries an R3.1 "B comparison" command at power-scale 0.8 beside the rl_default baseline; the sources do not say whether it was written before the ban or was ever run.
Applies when
- reusing a power, torque or action scale from one skill on another
- a policy whose actions map to absolute targets over the full joint range
- choosing a gentler setting for a first or second hardware trial
“kp×0.9 / kd 不动 —— recovery_r3_1 成功 48/50, τ 需求中位 hip_pitch/knee 120~125% -> 100~104% 部署限, 腿-腿接触 2152 -> 1095 帧; ±10% 在训练 kp DR 带内. ⚠️ power-scale 0.8 对 recovery **禁用**: 全 ROM 契约下 0.8 砍的是行程 端点 (深蹲收腿/站直够不到), 实测 4/50 起不来, 且尖峰 (kp·err) 一点不降 —— 它是 walk/omni 的安全档, 不是 recovery 的.”
git:Lucen-recovery@origin/recovery:robot.yaml § gain_profiles 注释: recovery 真机测试安全档 (2026-08-10) / rl_kp090 Removing a hand trim re-exposed the plant offset it had been silently compensating - and a slope scan told bias from sensitivity
hand-trims-hide-plant-offsetsTreat hand-tuned trims as undocumented plant measurements: before deleting one, find what it compensates and re-house that knowledge in the model or the reward budget; diagnose posture errors with a sensitivity sweep to distinguish constant bias from gain problems.
Symptom
After switching from the old hand-trimmed default to the clean geometric zero, the retrained stand policy's only regression was torso lean: 1.8 deg -> 4.1 deg backward.
Context
The old default's ankle-pitch trim (-0.0489/+0.0628) had been pre-compensating a fore-aft COM mismatch; removing the trim removed the hidden compensation, and the posture reward alone was too weak to win it back. A COM sensitivity scan settled what kind of problem this was: sweeping base COM offset -50 to +50 mm gave nearly identical slopes for old and new policies (~0.026 deg/mm) - "不是质心敏感度问题, 是恒定偏置" (not a sensitivity problem, a constant bias). Fix landed in stand_v1b: posture corrected to +0.24 deg while keeping symmetry (<=0.1 deg) and low effort (0.259), disturbance rejection better than both predecessors. Model credibility was checked the honest way: v0's sim prediction at the real COM position (-22 mm) was -2.31 deg lean vs real measured 2.2-3.1 deg - "预测精准命中" - which is what licensed trusting v1b's -0.52 deg prediction. (Side flag from the same file: a sign convention had been documented wrongly in early comments - gravity_base[0] > 0 is forward lean.)
Change
Trims retired in favor of explicit modeling: symmetric geometric default plus a posture-reward budget sized to carry the real COM offset; the offset itself known (real COM ~22 mm behind model).
Outcome
stand_v1b passed acceptance as the standing lineage's final version; the walk-line requirement "加大躯干姿态惩罚权重" was upgraded from suggestion to mandatory, since walking amplifies what standing tolerates (real walk_v1 hit 26 deg lean vs sim 7.4).
Mechanism
Hand trims are plant knowledge stored in the wrong place - invisible, asymmetric, and stale after recalibration; removing them re-exposes the raw plant error. A sensitivity sweep separates the two possible diagnoses (slope change = control problem; parallel offset = constant plant bias), each with a different fix.
Applies when
- cleaning up hand-tuned offsets/trims in defaults or calibration
- a posture bias appears after a default or calibration change
- deciding whether a lean is a COM-sensitivity or constant-offset issue
“两者斜率几乎相同(≈0.026°/mm),v1 只是整体多后仰约 2.4° —— 不是质心敏感度问题,是恒定偏置。成因:旧 default 的踝俯仰 trim(−0.0489/+0.0628)本就预补偿了前后质心偏差,换成零位 default 后这份补偿没了 … v0 在真机质心处(−22 mm)的 sim 预测为 −2.31° 后仰,真机实测 2.2~3.1° 后仰 —— 预测精准命中。”
train/RETRAIN_v2.md § 4b. stand_v1 独立验证结果 / 4c. stand_v1b 验收结果 FK-verify a borrowed reference's structure, then size its amplitude by the reference's job - it pins phase, the policy adds lift
reference-structure-fk-amplitude-divisionWhen borrowing a reference trajectory: verify its structural claim against your own kinematics (an invariant like flat-foot), assign it the phase-pinning job, and size amplitude low enough that the policy contributes the lift - moving toward a proven foreign value in halves, not jumps.
Symptom
walk_v4 had big knee swing (40-46 deg) but only 18-24 mm foot lift - amplitude without hip/knee/ankle phase coordination; later, walk_v5's real-robot swing ballooned to 73.6 deg (sim 55.7) with violent footfalls - amplitude over-driven by the reference.
Context
Structure first: Humanoid-Gym's 1:2:1 hip:knee:ankle reference was verified on the local model before adoption - the ratio exactly satisfies the locally derived flat-foot constraint hip - knee + ankle = 0, FK-tested at multiple amplitudes with sole pitch 0.00 deg throughout. Amplitude second, and here the first reasoning failed honestly: FK said shorter legs need LARGER reference scale (0.30 for 30 mm lift), and the FK was correct - but the premise was wrong ("FK 没错, 但前提错了"): it assumed foot lift must come from the reference. HighTorque Pi, same scale, uses 0.08 with a 0.02 m foot-height target - proof that lift is added by the policy ON TOP of the reference, whose actual job is pinning the phase relationship. Scale 0.30 made the reference the entire gait: over-constrained and over-driven. The correction went to 0.15, deliberately not Pi's 0.08: "一次只走一半, 留退路" (walk half the distance, keep a retreat).
Change
target_joint_pos_scale 0.30 -> 0.15 as one of v6-minimal's three changes, treating both the footfall force and the lateral kicking (yaw momentum scales with leg swing amplitude).
Outcome
v6 improved landing force 1.72x -> 1.55x, suspended tilt 45.9 -> 23.0 deg, turn-gain asymmetry 70% -> 19%; the later v6-halved-shaping experiment (35 mm -> 4 mm collapse) confirmed the reference still carries the gait's existence on this machine - the division of labor is real but machine-specific.
Mechanism
A joint-space reference plays two separable roles: encoding structure (phase relations that keep the foot flat) and injecting amplitude (energy). Structure transfers across robots and is checkable by FK against an invariant; amplitude is a negotiation with the policy, and over-assigning it to the reference removes the policy's freedom to modulate lift with state.
Applies when
- importing a reference gait / imitation target from another codebase
- reference amplitude reasoning based on leg length alone
- real swing amplitude far exceeds sim's under a strong reference
“FK 没错, 但前提错了。我默认抬脚必须由参考轨迹产生。HighTorque Pi 同尺度机器人 … 用 0.08, 而它 target_feet_height = 0.02 m —— 说明抬脚是策略在参考之上加出来的, 参考只负责钉住髋/膝/踝的相位配合。我们取 0.30 等于让参考本身就是整个步态, 过约束 + 过驱动”
train/WALK_V6_MINIMAL.md § ① target_joint_pos_scale 0.30 → 0.15 Three same-shaped judging errors - task metrics (survival, tracking, displacement) cannot stand in for posture metrics
task-metrics-vs-posture-metricsKeep validated posture-class rows (tilt, per-joint L/R asymmetry, temperature) in every acceptance battery alongside task rows; when operator feel contradicts the gates, suspect the metric class before the operator - and never build a new skill on what is actually an asymmetry defect.
Symptom
The C2 product judged "full pass" on task metrics (A800: turn-gap 7 pp, vx+0.30 19/20) felt WORSE in the operator's hands than the half-pass 700: A800 tilted up to 12.90 deg (700: 6.64), drifted left while standing, showed larger per-joint asymmetries, and ran its hip_rolls 5 degC hotter.
Context
The re-judgment catalogued three same-type metric errors in one campaign: (1) stand judged by SURVIVAL - missed 0.5-1.4 m wandering; (2) stand ranked by DISPLACEMENT - ordering was opposite to real feel (tilt ordering matched); (3) chirality judged by wz-tracking GAP - measured turning symmetry while the robot's actual disease was postural left/right asymmetry, "两个不同的东西,且结论相反". Common pattern named: "我一直用「任务指标」当判据,而真机手感对应的是「姿态 指标」… 任务类指标不能替代它". The fix was already in the data: the per-joint left/right asymmetry table (printed identically by sim2sim and deploy) agreed with hardware in direction on every row - "判据可用、有预测力,我只是没把它写进 PASS 条件". Shipping decision followed the posture read: product reverted to 700 ("又一次「买到 精度、卖掉别的」"), and the C4 root moved to 700 as well, with the sharpest line of the episode: A800's left-drift "like sidewalking" is probably its frontal-plane asymmetry defect, not a capability - "在缺陷上建能力是危险的".
Change
Two posture quantities with demonstrated real-robot predictive power promoted into every PASS battery: tilt-max median and per-joint left/right asymmetry (both sim-computable, deploy-homologous); motor-temperature readout added to session close-out.
Outcome
Deployment flipped to the posture-better checkpoint; the hip_roll temperature table (43-48 degC vs 25-28) confirmed the earlier 90%-of-heat account; the run-line acceptance battery inherited the posture rows from birth ("任务类替代不了姿态类").
Mechanism
Task metrics measure goal attainment under the evaluator's episode definition; posture metrics measure the body state trajectory that operators, motors, and downstream skills actually experience. The two can rank candidates oppositely because task success tolerates postural pathology - so a battery without posture rows is blind to exactly what hardware feel reports first.
Applies when
- hardware feel disagrees with a green acceptance table
- choosing between checkpoints that split task vs posture metrics
- selecting the root for a skill that resembles an existing defect
“共同模式:我一直用「任务指标」(存活 / 跟踪率 / 位移)当判据,而真机手感对应的是「姿态指标」(倾角、逐关节左右不对称)。→ 验收判据里必须有姿态类指标,任务类指标不能替代它。… A800 的「左飘像 side walk」很可能 … 是它更大的额平面不对称的表现 —— 在缺陷上建能力是危险的。”
train/README.md § C2 选点改判 (2026-08-09): 手性判据第三次选错指标 Add a single-point-suspension test to acceptance - the ground is a free stabilizer that hides divergence
suspension-probe-removes-free-stabilizerInclude at least one acceptance condition that strips the environment's free stabilization (suspension, or equivalent) - the sim-passing policy that fails on hardware is often failing a condition the battery never posed.
Symptom
walk_v5 looked healthy in every on-ground sim test yet diverged on the real robot - the acceptance battery had never measured a condition that would have revealed it.
Context
The battery gained a single-point-suspension probe (robot hung, feet free): measure torso tilt while the policy runs without ground contact. v5 scored 45.9 deg mean tilt suspended - wildly unstable - which the file calls "最灵敏的失稳探针(拿掉地面这个免费稳定器)": ground reaction forces passively stabilize a marginal policy, so on-ground metrics saturate long before the policy's internal balance is actually sound. v6 halved it (23.0 deg, target <10 deg) - progress visible on a scale where on-ground numbers showed nothing.
Change
Suspended-tilt added as a standing acceptance row; run under the honest contact parameters battery (accept_v2 with measured condim 4 / torsional friction 0.035), under which v5 correctly FAILS in agreement with the real robot.
Outcome
The sim battery's verdict on v5 flipped from pass to fail, matching hardware; suspended tilt became the discriminating metric between v5 and v6 (45.9 vs 23.0 deg) when ground metrics differed little.
Mechanism
Contact with the ground closes a stabilizing feedback loop the policy gets for free; removing it exposes the policy's own attitude control authority. A metric measured only in the assisted condition cannot rank policies by the unassisted quantity that hardware will actually demand during perturbations and flight phases.
Applies when
- sim acceptance passes but hardware diverges
- designing an acceptance battery for a legged robot
- two candidates tie on ground metrics
“单点吊那条是最灵敏的失稳探针(拿掉地面这个"免费稳定器"), v5 在地上一切正常却在真机发散, 就是因为验收从没测过这个工况。”
train/WALK_V6_MINIMAL.md § 5. 验收 Tightening the bridge's rate limiter under an unchanged policy cut torque peaks 30-50% and made other things worse - the policy cannot see the limiter, keeps commanding and winds up; a deploy-side limiter is a safety net, not a cure
deploy-rate-limiter-windupA rate or torque limiter added at deployment lowers peaks but the policy still commands as if unconstrained (saturation, windup, new contacts); use it as a safety net mirrored in evaluation, and put the constraint where the policy can learn around it.
Symptom
After the violent first real-robot get-up, the cheapest candidate fix was to tighten the bridge's slew (rate) limit for the recovery policy without retraining.
Context
Probe on R3.1 in MuJoCo (5 categories x 3 seeds, mu 1.0), monkeypatching the limiter with no repository change: TIGHT = RS06 4.0 / RS02 3.0 / RS00 2.0 rad/s (about 0.08/0.06/0.04 rad per policy step) against the current vel_limit setting.
Change
The probe decided the role of the limiter rather than a deployment.
Outcome
Success 14/15 -> 12/15; get-up median 2.35 -> 3.53 s (max 9.30); torque demand peak median hip_pitch 164% -> 111%, knee 166% -> 86%; action saturation still 100%; leg-leg contact 558 -> 860 frames. The limiter was kept only as a real-robot safety net (mirrored into sim2sim evaluation); the cure moved into training - where the next lesson was that a limiter anchored on the last command is itself an integrator (slew-anchor-is-an-integrator).
Mechanism
A policy that never trained with the limiter keeps issuing the targets it learned; the limiter clips them, the target window runs ahead (windup), and the robot follows a trajectory the policy never evaluated.
Applies when
- a trained policy is too violent on hardware and a quick deploy-side fix is tempting
- adding slew, torque or velocity limits in a bridge or firmware
- evaluation and deployment use different limiter settings
“判读:**链路侧收紧立等可取地把 τ 峰值砍 30~50%,但成功率掉、饱和率仍 100%、 腿-腿接触反升** —— 策略感知不到限速器,目标窗口继续狂奔。⇒ 收紧 slew 只配当 **真机侧安全网**(必须同步进 sim2sim 口径,基础设施现成),**不配当治法; 治法必须进训练**。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 探针:收紧桥层 slew,r3_1 不重训直接测 Diagnose a behavior failure by enumerating hypotheses and auditing each against the actual config, cheapest first
hypothesis-table-code-auditBefore changing anything, write the full hypothesis list for the symptom and audit each against the resolved config and measured magnitudes, cheapest check first; train only on the survivors.
Symptom
Real robot leaned forward "wanting to walk" but dragged its feet instead of lifting them - a symptom with many plausible causes and no obvious single fix.
Context
Seven hypotheses were listed and each checked against the actual training config files (velocity_env_cfg.py, isaac_values.py), ordered by check cost: missing foot clearance term (CONFIRMED, primary - feet_air_time existed but no swing-height term at all); energy penalties dominating (REJECTED - energy terms total -0.19 vs tracking +1.2, 16%); command range too narrow (CONFIRMED - (0.15,0.35)); nominal pose too crouched / action scale too small (HALF - knee 0.5 rad = 28.6 deg deep, scale fine); mixed PD across motor types (REJECTED - already grouped); missing base-height reward (REJECTED - present at -5.0); height-drop termination (REJECTED - none exists, which itself became finding #4 of the fix list).
Change
The audit produced a ranked fix list (add clearance penalty; widen speed range; reduce nominal crouch) with each rejected hypothesis documented so it would not be re-litigated.
Outcome
Three confirmed causes fixed over v5/v6: swing height went 22-23 mm -> 34 mm, tracking 81% -> 87%; the rejected hypotheses stayed rejected (no wasted rungs on energy weights or PD grouping).
Mechanism
Multi-cause symptoms invite guess-and-train loops; a written hypothesis table forces each candidate to be confirmed or rejected against actual values (not impressions), and cost-ordering the checks means most hypotheses die for the price of reading a config.
Applies when
- a real or sim behavior failure has multiple plausible causes
- the team is about to "try a fix" without an audit
- post-mortems keep re-proposing already-rejected causes
“真机现象:躯干前倾像要走,脚抬不起来(拖着蹭)。按成本从低到高逐条核查 … | 1 | 缺 foot clearance | ✅ 成立,首要 | 有 feet_air_time,无任何摆动足高度项 | | 2 | 能量惩罚压过跟踪 | ❌ 不成立 | 能量类合计 −0.19,跟踪 +1.2,只占 16% |”
train/WALK_DIAGNOSIS.md § walk 拖地问题 — 七条假设的代码核查结果 Ideal PD is not enough - add a delay buffer and fit armature/friction/delay per joint
actuator-delay-buffer-fittingNever ship ideal PD to hardware: add a measured delay (in control steps) and per-joint armature/friction fitted from step and sine responses, and treat remaining actuator mismatch as your standing largest sim2real residual.
Symptom
Standard ideal PD actuator model transfers poorly; sim assumes targets take effect instantly and joints reach arbitrary acceleration.
Context
A developer with a successful on-hardware Isaac Lab biped modified the actuator model in two ways and calibrated it against the real robot: step-response plus sine-sweep tests (positive step, negative step, sine tracking), overlaying sim curves on measured curves and hand-tuning.
Change
(1) Delay buffer: action targets take effect after a uniform 6 time-step delay on all joints; (2) acceleration limiting so the actuator cannot reach arbitrary acceleration; (3) per-joint fit of armature / friction / delay - different joints genuinely needed different values.
Outcome
Hip joints fit worst, knee best; the developer rated the result "not perfect, the best I could do" and still listed actuator-model improvement as next work - i.e. even the fitted model remained the dominant residual.
Mechanism
Real actuation is a lagged, bandwidth-limited system; a delay buffer and acceleration cap are the two cheapest structures that reproduce its phase and magnitude response. Per-joint differences come from differing load, wiring, and friction states, so a single global constant underfits.
Applies when
- actuator model in sim is ideal PD with no delay
- step-response of real joint visibly lags or overshoots the sim's
- budgeting which sim2real gap to attack first
“标准 ideal PD actuator 不够用,他改了两处:延迟缓冲:目标不是立即生效,全部关节统一 6 个 time step 延迟 / 加速度曲线:执行器不能瞬间达到任意加速度 … 用 armature / friction / delay 三个参数逐关节拟合,标定方法是阶跃响应 + 正弦扫描 … 髋部关节偏差最大,膝关节最好。”
Experience.md § 执行器建模 —— 最值得抄的一条 (lines 50-59) An outer heading P-loop at deploy cut drift 10x because its output stays inside the trained command band - then training was aligned to it
deploy-heading-loop-and-align-trainingFix drift-class problems first with an outer loop whose output provably stays inside the trained command band; when adopting it permanently, align the training command generator to the deployment's actual command mixture (feedback-driven AND constant), matching law, gain, and clip exactly.
Symptom
Persistent heading drift on straight-line walking (v6 net yaw 60.3 deg over 15 s) that reward-side fixes had only partially tamed.
Context
The deploy stack added --heading: an external P loop wz = clip(0.5 * wrap_to_pi(theta0 - theta), +/-0.6), recomputed each frame and fed into the policy's ordinary wz command slot. Measured: net yaw walk_v6 60.3 -> 5.8 deg, walk_v5 17.4 -> 4.2 deg. A run-level audit later corrected the mechanism story: training had heading_command=False since v1 - the policy had NEVER seen heading-error feedback, so the loop works purely because its output lands inside the trained command distribution wz ~ U(+/-0.6): "收益真实,当时的机理解释写错了" (the benefit is real; the mechanism explanation had been wrong). v8 then closed the loop properly: training-side heading command enabled with rel_heading_envs=0.5 - half the envs get heading-error-driven wz, half get explicit constant wz, because deployment feeds wz BOTH ways (straight-line = heading feedback, turning = constant command) and rel=1.0 would have made constant-wz turning out-of-distribution. The law, gain, and clip were aligned item-by-item between trainer and deploy tool.
Change
Deploy-side outer loop first (no retrain needed); then v8-D enabled the matching training-side heading command at rel=0.5 with identical gain (0.5) and clip (+/-0.6), contract unchanged (wz slot carries the computed value).
Outcome
Drift handled at deploy (5.8 deg) generations before training caught up; the alignment removed the residual train/deploy distribution mismatch, with the accepted cost booked (open-loop straight walking becomes more OOD for heading-envs - irrelevant since acceptance and deployment always run the loop).
Mechanism
A learned velocity-tracking policy is a valid inner loop for any outer controller whose commands stay within the trained command distribution - the policy needs no knowledge of the outer objective. Full alignment then requires training on the same mixture of command sources the deployment actually uses, in the observed proportions.
Applies when
- heading/position drift on a velocity-tracking policy
- designing outer loops over learned locomotion controllers
- training command distribution differs from how deployment feeds commands
“审计更正(2026-08-02,run 级 env.yaml):训练侧自 v1 复盘起就是 heading_command=False … 策略从未见过航向误差反馈。--heading 是评估/部署侧外加的航向 P 环(wz=clip(0.5·err,±0.6), 落在训练分布 wz~U(±0.6) 内)。实测净偏航 walk_v6 60.3° → 5.8° … 收益真实,当时的机理解释写错了”
train/WALK_V7_SPEC.md § 0. 本轮之前已经改掉 (航向闭环, 含审计更正) No parameter tuning on the floor - a failing config retries once, then it is out; anomalies go back to sim
no-field-tuning-protocolHardware time is for executing and measuring the pre-registered matrix, never for tuning: failing configs get one retry then elimination, anomalies get recorded and reproduced in sim, and contract-check bypass flags stay unused.
Symptom
Hardware sessions create pressure to fix problems live - nudge a gain, tweak a scale - which destroys attribution and risks the robot.
Context
The anomaly-handling section of the acceptance sheet is three fixed plays: (1) falls at start -> retry once at the same settings; falls again -> that configuration is eliminated, "不现场调参" (no on-site parameter tuning); (2) limit cycle or motor screech -> stop immediately, record the gain level and the joint, reproduce in sim before any discussion; (3) systematic disagreement with sim -> record it as a finding (hardware outranks sim) rather than adjusting anything to force agreement. Related guardrails elsewhere in the sheet: never pass --allow-unstamped / --allow-plant-drift to bypass manifest checks - if it errors, something real is wrong, stop and look.
Change
Field sessions restricted to executing the pre-written matrix; every fix path routed through sim reproduction and the normal config/rung process.
Outcome
Sessions stayed interpretable (each run matched a documented config) and safety overrides never became habit; anomalies arrived back in sim as reproducible cases instead of half-remembered floor stories.
Mechanism
Field-tuned values are measured under adrenaline on one floor with no logging or baselines - they contaminate the config lineage and are unattributable afterwards; and every bypass flag that skips a contract check converts a designed safety property into an operator promise.
Applies when
- a config fails or oscillates during a hardware session
- someone reaches for a live gain tweak or a bypass flag
- writing the anomaly-handling section of a deployment runbook
“起步即摔 → 换档重试一次, 仍摔则该档出局, 不现场调参。出现极限环/啸叫 → 立刻停, 记录档位与关节, 回 sim 复现再议。… 不要给 --allow-unstamped / --allow-plant-drift —— 三枚 ONNX 都已盖章 … 真要报错说明有别的问题, 停下来看。”
train/REAL_RUN_S2.md § 4. 异常处置 Deployment power derating damages non-forward axes far more than forward - sweep it in sim before deploying
power-scale-hurts-nonforward-axesTreat deployment power/torque scaling as a plant parameter: evaluate the policy in sim at the exact deployment scale, expect non-dominant axes to degrade first under derating, and either deploy at the training power or train with power randomization.
Symptom
Policies deployed at power-scale 0.8 (a safety derating of commanded torque) looked fine walking forward but were weak at backward and turning, inviting the wrong diagnosis "the skill was not trained well".
Context
Measured repeatedly: on s1e, going 1.0 -> 0.8 cost forward 18% but backward 58%; on C4-ff800, turn tracking was +25%/+40% at pw0.8 vs +75%/+58% at pw1.0, backward 51-52% vs 97-103%, while forward stayed 96-98% at both. Sim evaluation numbers in the plan were all pw1.0, but the robot was being run at 0.8.
Change
Pre-deploy protocol added: sweep the exported policy across power in sim (for PW in 0.8 0.9 1.0: eval_c_matrix --power $PW --seeds 20) and deploy at the first level where both turn directions reach >=50%. For C4 the recommendation was raise the robot to pw1.0 - the sweep showed it nearly free (saturation 47%->33%, left foot-clipping danger zone 25%->6%, cost only tilt 6.7->8.3 deg).
Outcome
Turning "weakness" resolved without any retraining; the sim sweep correctly predicted the real-robot signature at both power levels.
Mechanism
Forward walking is the reward-dominant, torque-cheapest skill with the most margin; backward/turn/sidewalk live closer to the torque envelope, so a uniform torque derating consumes their margin first. Training ran at power 1.0 (the trainer does no power scaling), so deploying at 0.8 is a systematic underactuation the policy never experienced.
Applies when
- deploying with any torque/power derating or safety scale
- secondary skills (backward, turn, lateral) underperform on hardware while forward walking looks fine
- choosing the deployment power level for a new policy
“power 衰减对非前进轴的伤害远大于前进轴(s1e:前进 1.0→0.8 掉 18%,后退掉 58%)。转向是非前进轴,0.8 下很可能明显跟不动。”
train/C_LADDER_RUN.md § 3c. A-2 上机前先定部署力度档 / 3p. 二 Freeze the deployment contract, stamp every export, and let an automated checker catch wiring bugs
contract-freeze-and-checkerFreeze and fingerprint the policy I/O contract; ship contract changes as new versioned profiles that leave old artifacts bit-identical; and extend the automated contract checker with every pipeline change, forcing the new path to execute in the check.
Symptom
Contract-level changes (observation layout, action pipeline) are where silent sim/real divergence is born; two real wiring bugs appeared the one time the action pipeline was extended.
Context
The 215-dim observation contract was frozen ("纪元 3,三机 digest" - an era number plus a digest agreed across three machines); proposals that would break it (e.g. a GRU memory) were rejected on contract grounds. Every exported ONNX is stamped and verified with a manifest (onnx_manifest --stamp / --verify), and deployment refuses mismatched combinations. When C4 added the lateral feed-forward, it went in as a NEW profile (omni_ff) leaving the existing omni profile's behavior bit-identical; the checker (check_contract) was extended to force the feed-forward path to actually execute (cmd_vy=0.13) and promptly caught two genuine bugs: (1) re-clamping with soft_joint_pos_limits after the feed-forward (0.23 rad deviation) instead of reusing the parent's clip; (2) indexing processed actions by asset.joint_names instead of the action term's own contract-ordered _joint_names, which landed the feed-forward on the wrong joints (l_hip_yaw / r_ankle_pitch).
Change
Contract discipline as implemented: frozen dims + digest; manifest stamping and refusal; contract changes only via new versioned profiles; checker updated in the same commit as any pipeline change, with inputs chosen so new code paths are exercised.
Outcome
Both wiring bugs caught before any training or deployment ("两个都是 check_contract 当场抓出来的 —— 这次它值回票价"); old deployments provably unaffected by the new profile.
Mechanism
The contract is the only interface the policy and robot share; freezing plus fingerprinting makes divergence detectable, and an executable checker turns "the contract holds" from a belief into a test - but only if its inputs actually drive the new code path.
Applies when
- modifying the action or observation pipeline of a deployed policy
- exporting policies for hardware
- proposals that would change observation dims or history structure
“契约校验抓到的两个真错误(记账,别再犯):1. 前馈后误用 soft_joint_pos_limits(URDF 限位 ×0.9)重钳 → 0.23 rad 偏差 … 2. 用 asset.joint_names 索引 _processed_actions → 前馈落到 l_hip_yaw/r_ankle_pitch 上 … 两个都是 check_contract 当场抓出来的 —— 这次它值回票价。”
train/C_LADDER_RUN.md § 3j. 契约级改动 / 契约校验抓到的两个真错误