Training Coach
Doctrine
A report may cite any of these as doctrine-N.
doctrine-1Contract freeze and fingerprint disciplineThe policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.
Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).
Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.
doctrine-2Attribution by resolved training params - never eval-override knobsCapability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.
Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).
Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.
doctrine-3PASS gates become constraints; FAIL gates become objectivesOnce a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.
Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.
Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".
doctrine-4One variable per ladder rung - counted against what the checkpoint sawA rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.
Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.
Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.
doctrine-5Pre-register risks, readings, and stop criteria before the ladderBefore a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.
Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).
Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.
doctrine-6Plant parameters are measured, never inventedEvery plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.
Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).
Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.
doctrine-7Sim2sim gate before sim2real - under deployment conditionsEvery checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).
Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).
Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.
doctrine-8Observation honesty - the actor's inputs are a hardware contractThe actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.
Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).
Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.
doctrine-9Reward economics are audited in realized currencyReward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.
Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).
Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.
doctrine-10The zero-cost option must be the desired behaviorFor every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.
Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).
Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.
doctrine-11Measurement discipline: independent referees, signs, distributionsA disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.
Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).
Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.
doctrine-12The deployment pipeline is plantIrreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.
Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).
Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.
doctrine-13DR budget is finite; its distribution is the measured supportRobustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.
Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).
Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.
doctrine-14Gates measure what hardware feels: posture, margins, stripped assistsAcceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.
Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).
Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.
doctrine-15Fork and root selection: recoverability, maturity, frozen rewardsChoose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.
Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).
Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.
doctrine-16Curricula: verified engagement, lineage counters, disease-phase gatingAutomatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.
Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).
Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.
doctrine-17Probe before training: feasibility first, hypotheses in tablesAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.
Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).
Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).
doctrine-18External advice is recomputed locally; values transfer as ratiosEvery external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.
Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).
Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.
doctrine-19Hardware sessions are scripted experiments, not tuning sessionsReal-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.
Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).
Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.
doctrine-20Close questions in writing; restart when the debt is structuralAudited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.
Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).
Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.
doctrine-21Name the quantity in the space it lives inA goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.
Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).
Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.
doctrine-22Continuation needs a live gradient; a release is chosen by a scanContinue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.
Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).
Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.
Experience cards
70 cards matching “dr-tail-plant-continuation”.
Set torque limits per joint from measured gait peaks - a uniform percentage is the wrong shape, and training must use the deployed numbers
torque-limit-shape-by-measured-peaksMeasure per-joint torque peaks in the actual gait and set each limit as measured-peak x margin capped at rating; then propagate the same numbers into training and add an automated deploy-time consistency check - never derate by a uniform percentage, never let training assume torque deployment will not grant.
Symptom
A uniform 50% torque derating (18/8.5/7) had piled safety margin on the joints that never use it while cutting the busiest joint below half its measured demand.
Context
Per-joint gait peaks were measured (walk_v5 at cmd 0.3/0.6): RS06 (hip_pitch/knee) uses 5.5-5.9 N*m = 15-16% of its 36 N*m rating - cutting it to 12 is a free safety win; RS02's ankle_pitch runs at 16.2 N*m = 95% of its 17 N*m rating - "它是速度的硬件瓶颈", no room to cut; RS00 measured 36-44%, capped at 11. The resulting shape 12/17/11 replaced the uniform percentage. Sweeps across several limit sets (rated / 50% / 14-17-11 / 12-17-11) produced identical speed, lift, and landing force - within this range the limits do not shape the gait; what matters is consistency: "关键是训练和硬件必须是同一个数", because the exporter fills effort_limit from tau_limit, and a policy trained at rated 36/17/14 "会假设有三倍力矩可用" while deployed at 12/17/11 (exactly the v5 cross-generation inconsistency later suspected in its wild kicking).
Change
robot.yaml tau_limit set to the measured-shape 12/17/11, firmware written to match, and train/isaac_values.py regenerated so training sees the same limits; the deploy tool self-checks limits against robot.yaml on every run.
Outcome
Free safety margin captured where demand is low, the real bottleneck joint left at rating, and the train/deploy torque worlds unified with an automated consistency check.
Mechanism
Torque demand is grossly unequal across joints in a gait (15% vs 95% of rating here); a uniform percentage misallocates the safety budget by construction. And since the trainer treats effort_limit as a plant truth, any train/deploy mismatch is an invisible plant gap of exactly the mismatch ratio.
Applies when
- choosing safety torque limits for a legged platform
- training-vs-deployment actuator limit audit
- one joint runs near rating while others idle
“曾用统一 50%(18/8.5/7)是错的形状: 把余量堆在用不到的 RS06 上, 却把 ankle_pitch 砍到需求的 52%。… RS02 在 0.6 m/s 已用到额定 95%, 它是速度的硬件瓶颈 … 实测多组限幅 … 完全一致 —— 限幅在这个范围对步态零影响, 关键是训练和硬件必须是同一个数。… 若训练仍按额定 36/17/14, 学出的策略会假设有三倍力矩可用。”
train/WALK_V6_MINIMAL.md § 3. 训练侧必须同步的一件事 A time gate that had cured one lineage's rushing made the from-scratch lineage trade away its stance width twice (0.364 -> 0.235 m, 0.355 -> 0.251 m) - its disease was absent there, so the fix was retired and the pre-gate checkpoint shipped
time-gate-vs-wide-stance-retire-the-fixCarry a fix into a new lineage only if its disease is present there; a mechanism that cured one lineage can be net negative in another, and when doubling a term's weight recovers almost nothing, treat the two objectives as structurally in conflict and remove the one whose purpose is gone.
Symptom
V3.1's phase 2 (the 3 s zero gate on standing income, continued from P1b) kept 100% success on every friction level and slowed the get-up, but the lateral stance drifted 0.364 -> 0.235 m and hip yaw crept to 57 deg against its 60 deg limit. With the width band's weight doubled (P2c, after P1c) it drifted again, 0.355 -> 0.251 m, below the pre-registered 0.30 m failure line.
Context
The zero gate had been introduced in V2.5/V2.5b to slow the old lineage's get-up. In V3.1 the rushing was already absent: P1c got up in 0.90-1.06 s with a worst torque ratio of 73.3%, better than the stamped v2_6c, because the full beta curriculum, second-difference smoothing and pull curriculum had cured the violence inside training.
Change
Recorded as a candidate law with two data points - the zero gate and a wide stance are mutually exclusive here - and the zero gate was removed from the V3.1 recipe. P1c (the pre-gate checkpoint) went through the full stamp-level acceptance instead.
Outcome
P1c passed everything: all six criteria, lateral stance 0.355 m, foot tilt P75 2.0 deg, mu {1.0, 0.8, 0.6, 0.4} x 10 seeds all 100%. recovery_v3_1p1c.onnx was stamped and pushed to the robot channel.
Mechanism
The zero gate moves the income toward "stay stable until the end", and under low-friction DR a wide stance has a slip tail, so survival outbids the width band; doubling the band's price bought back only 0.016 m - an auction that does not converge signals structural conflict, not an under-priced term.
Applies when
- porting reward mechanisms from an old lineage into a fresh recipe
- a width, margin or posture metric erodes during a late training phase
- a weight increase produces a negligible change in its target
“**定律候选(二实证):归零门 × 宽站互斥**。 … 加价翻倍只挽回 0.016,竞拍不收敛)。 … **归零门是 v2_5 血统的历史包袱,对 V3.1 配方是净负资产,P2 阶段除名 —— P1c 即终点形态**。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §49 终验(2026-08-14) When hardware underperforms, audit deployment knobs before prescribing retraining
deploy-knob-attribution-before-retrainingBefore any "retrain it" decision, reproduce the symptom in sim under the exact deployment configuration; if the symptom follows the deployment knob rather than the checkpoint, fix the knob or randomize it in training - never top-up-train the skill.
Symptom
Real-robot feedback after the C4 deployment - "turning is weak" - with two retraining options on the table: top up turn training, or restart from the s1e root.
Context
The sim account showed the policy turned well (75-81% at pw1.0); the robot was deployed at power-scale 0.8. The 3-6 pp difference between C2 and C4 policies at the same power was noise; the 40-50 pp difference between power levels was the entire effect. Both proposed retraining paths would have burned budget on a non-existent training gap, and restarting from s1e would additionally have discarded the sidewalk skill that took four rungs and a coordinate-bug hunt to obtain.
Change
Decision: retrain nothing. (1) Try pw1.0 on hardware first - sim says net gain; (2) only if 1.0 is unacceptable (heat/feel), the correct training fix is power/torque randomization in the S2 plant line (one variable, fixes turn and backward together) - not skill top-up; (3) restart-from-root explicitly ranked worst.
Outcome
The "weakness" was fully explained by the deployment knob; the sim/real signatures matched the earlier power-derating law verbatim ("与 C2 时代 power 衰减主要伤非前进轴 逐字吻合").
Mechanism
The policy's competence is defined under its training plant; deployment knobs (power scale, teleop mapping, command bands) silently define a different plant. Attributing a deploy-plant effect to a training gap produces exactly the wrong fix - more training on the wrong variable.
Applies when
- real robot underperforms a skill that sim says is fine
- proposals on the table include retraining or re-rooting
- deployment uses any override the trainer never saw (power scale, remapped commands, different control rate)
“正确的训练修法不是补训转向,而是训练时加 power/力矩随机化让策略在 0.8 下自己补偿 —— 单变量,属 S2 plant 线,一次同时修好转向与后退;从 s1e 重训是最差选项:丢掉四轮 + 一个指标 bug 才换来的侧走,而 C2 的转向本来就没问题。”
train/C_LADDER_RUN.md § 3p. 三 处置顺序(回答「补训转向 还是 回 s1e 重训」:都不该) A curriculum ramp keyed to the process step counter re-fires on every resume - and shipped policies never saw the penalty
curriculum-counter-lineage-stepsKey every curriculum/ramp schedule to lineage-cumulative progress, not per-process counters; and audit where your shipped checkpoints sit relative to every ramp - a penalty that no product ever experienced is not part of your training.
Symptom
vx+0.30 died at a fixed relative time in every resumed run: resume at 500 -> slide at 1100, zero at 1300; resume at 700 -> slide at 1300, zero at 1500 - absolute depths offset by exactly the resume offset, relative timetable identical.
Context
ramp_reward_weight (the saturation penalty ramp) read env.common_step_counter, which restarts at 0 for every run including --resume. So start_step=600 meant "600 iters after THIS resume", not "lineage iteration 600". The A/B arm pair was the clean proof: their env.yaml differed only in log_dir, only resume point distinguished them, and the omni CurriculumManager had exactly one active term - nothing else could produce that timetable. Second consequence: every shipped checkpoint (s1e-500 at +500, C2-700 at +200, A800 at +100) was selected before its run's +600, so the saturation penalty weight was 0.000 for every product ever shipped - explaining saturation 33% and raw |action| 1.9 against clip 1.0 (hip_roll in bang-bang), i.e. half the heat budget.
Change
Two independent recommendations recorded: (1) make the ramp count lineage steps (add the checkpoint's iteration offset on resume) or pin terminal weights in downstream rungs instead of ramping; (2) give the saturation penalty its own rung - never mixed into a skill-learning rung (that would be two variables again).
Outcome
Explained the recurring +600 death of the highest-amplitude command and the persistent actuator saturation of all shipped products with one root cause; honest caveat booked (at +800 the weight is only -0.086, small, but vx+0.30 is the command demanding the largest action amplitude, so it is squeezed first).
Mechanism
Resumable training splits "the lineage" from "the process"; any schedule keyed to process-local counters silently re-applies its transient to every descendant run, and any product-selection habit that picks checkpoints early systematically samples the pre-ramp regime - the curriculum exists in the config but never in any shipped policy.
Applies when
- resumed/forked training with any scheduled reward or DR ramp
- a metric dies at a fixed offset after each resume
- shipped policies show behavior a late-schedule penalty should prevent
“ramp_reward_weight 读的是 env.common_step_counter,它每个 run 从 0 开始,--resume 也不例外。… 原始 run / 臂B | 500 | 1100 = +600 | 1300 = +800;臂A | 700 | 1300 = +600 | 1500 = +800 … 所有出品其实从没见过饱和罚。… 这解释了 sat_max_pct 33%、raw |a| 最大 1.9(clip 是 1.0)—— hip_roll 一直在 bang-bang,而罚它的那一项权重恒 0。热账的一半在这里。”
train/C_LADDER_RUN.md § 3g. 系统性问题:saturation_ramp 每次 resume 归零 Choose the fork root by which candidate's shortfalls are recoverable, not by headline score
fork-root-recoverable-shortfallWhen picking a checkpoint to fork from, rank candidates by whether their weaknesses are trainable-back, not by current headline metrics; prefer the candidate whose deficits the upcoming training directly pays for.
Symptom
Multiple candidate checkpoints for the omni-command ladder root, each best at something different: fric-3000 had the best tracking precision (vx 88-91%) and hardened plant robustness; s1e-500 had lower precision (vx 82%) but was the only candidate that could still walk backward.
Context
Root selection ran as a data probe, not a preference vote: 6 candidates x 8 out-of-distribution omni commands x 20 seeds = 960 cells (probe_omni_0808.json). s1e-500 @pw1.0 survived 20/20 in all eight conditions including backward at 67% tracking; the deep-trained fric lineage scored backward 0-3/20 despite better forward precision.
Change
Decision criterion made explicit: list what each candidate exclusively wins at, then ask which of those wins the loser could train back. fric-3000's exclusive wins (precision, plant robustness) are both retrainable - precision is directly optimized by the reward, plant hardening is a planned later pass. s1e-500's exclusive wins (backward plasticity 20/20 vs 2/20, disturbance margin 159/160 vs 125/160, push chirality symmetry 40/40 vs 17/40) had all been shown unrecoverable - push-level rungs failed twice, chirality never recovered even with mirror augmentation on. Root = s1e-500.
Outcome
s1e-500 carried the whole C ladder; its backward skill was preserved through C2/C4 gates (regress budget <=2/20 enforced), and the final C4 product passed a 260-cell battery at 20/20 everywhere.
Mechanism
Training can re-earn anything the objective directly pays for, but capabilities that earlier training destroyed and never restored (plasticity, symmetry, robustness margins) are empirically one-way doors. The information-bearing comparison is therefore recoverability of each candidate's deficit, which the team stated as "独占项的可恢复性正好相反 —— 这就是判据" (the exclusive items' recoverability is exactly opposite - that is the criterion).
Applies when
- selecting a resume/fork root among several checkpoints
- one candidate is more precise but another retains a skill the rest lost
- planning a task-extension ladder from an existing lineage
“fric-3000 赢在精度(vx 88~91%…)与 plant 鲁棒性 → 两样都训得回来…;s1e-500 赢在可塑性(C1 20/20 vs 2/20)、抗扰余量(159/160 vs 125/160)、手性对称(推 ±6 N·s 40/40 vs 17/40)→ 三样都训不回来 … 独占项的可恢复性正好相反 —— 这就是判据。”
train/C_LADDER_RUN.md § 0. 为什么根是 s1e-500(数据,不是偏好) Prove an armless get-up exists before training it - a connected static domain, 25% torque on the cheapest path, an 8 mm hand-over gap, a static roll-over - and write down what each scan cannot represent
get-up-feasibility-accounts-before-trainingBefore training a get-up or any multi-contact skill, compute the quasi-static accounts - connectivity of the static domain, torque along the cheapest path, hand-over gaps, COM shift available for rolling - and state which configurations each scan cannot represent; when a policy gets stuck in one of those, extend the scan before blaming the reward.
Symptom
A torso-and-legs robot has no arms to push off the ground; whether it can get up from the floor at all was unknown when the line opened.
Context
recovery_feasibility.py ran three accounts before any training (the run line's "hard accounts first" discipline): a sagittal quasi-static scan (0.05 rad grid, 44,520 configurations, MuJoCo FK, flat-foot assumption). (1) The static standing domain (COM over the feet, torques in limit) has 29,586 cells, flood-fill connected with no islands, from a 0.097 m deepest squat to the 0.384 m stand. (2) The minimum-torque path peaks at 25% of the limits (knee 2.9/12, ankle 3.7/17 N*m) - a 4x margin. (3) All 508 ground-contact configurations have the contact behind the COM; the smallest gap to pure foot support is 8 mm. A roll-over account: swinging both straight legs to one side shifts the COM 96 mm against a 62 mm torso half-width - 1.6x, so rolling needs no momentum. Three conclusions were written down for later attribution: the legs are 80% of the mass (swinging them moves the COM), prone has no flat-foot hand-over face (merge into a supine/side sit first), and supine needs no sit-up (hip flexion is limited to 75 deg).
Change
The accounts gated opening the line and were cited in every later argument about what the robot can physically do.
Outcome
They held where they applied: in V1.0 every fall category was righted under a hard rate limit, which the spec records as the quasi-static roll-over account verified by training, and the 25% torque path was the basis for pursuing a slow get-up. They also misled once: account (3) is sagittal, and on 08-09 the spec corrected its scope - it cannot represent the splayed W-sit where the policy actually stalled. A follow-up prone hip-ROM scan (471,625 cells) found 3,912 two-foot-contact cells and none with both soles within 25 deg of level (best 40.2 deg): a flat-footed push-up from prone is infeasible on this robot, so the fix became where the feet go after sitting up.
Mechanism
A get-up needs a connected path through statically feasible configurations and enough torque along it; quasi-static accounts bound both cheaply, and momentum can only make the real problem easier. A reduced-dimensional scan, though, only speaks for the configurations it can express.
Conflicts
In R0.1-R0.2 the spec read account (3)'s "prone has no front hand-over" as "prone lacks the roll-over skill"; R0.3's confusion matrix showed prone had righted its torso 159/159, and the spec then restricted account (3) to the sagittal configurations it models.
Applies when
- opening a get-up, recovery or climbing skill on a new robot
- a robot lacks arms or other obvious contact options
- a policy stalls in a configuration a feasibility scan never modelled
“本机 **torso + legs、无手臂可撑地**,开训前先证明存在不依赖手臂的物理解 … 双腿同侧直腿摆最大横移 **96 mm = 1.6×** … **静态摆腿即可翻身,无需动量**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §1 机械可行性判决(2026-08-09,recovery_feasibility.py 三笔账) Train with self-collisions ON (filtering nested-link ghost pairs) - the reward wall prevents, the physics makes cheating impossible
self-collision-physics-plus-reward-wallNever train a contact-risk behavior with self-collisions disabled; enable them with an audited filter list for nested/overlapping pairs (zero contacts across a pose sweep), record the fps cost, and keep a calibrated distance penalty as the preventive layer on top.
Symptom
walk_v8 logged 107 frames of leg-on-leg contact while still earning 0.751 tracking score - because training-side self-collisions were OFF, leg clipping was literally imperceptible to the policy ("碰腿在训练里 根本感知不到").
Context
Enabling self-collisions naively is its own trap: an Isaac audit had shown PhysX auto-filters adjacent bodies (base-hip clean for free) but nested links generate ghost forces - calf and ankle_roll overlap 65 mm at the zero pose, producing 12x body-weight phantom forces. The v10 recipe: enable self-collisions, explicitly filter only the two nested pairs (l/r calf-ankle_roll), then run a zero-contact audit at three poses (nominal stand, walk crouch, swing-extreme) requiring contact count = 0, adding any residual pair to the filter and re-auditing; a 500-iter sanity run for NaN and an fps-cost record (measured -8.8%). Redundancy with the reward-side foot-distance wall was argued, not assumed: "N2 离得远(奖励侧预防),SC 碰了疼(物理侧兜底)" - the reward keeps distance at range, the physics makes contact hurt - so the v8-style "clip legs and still score" outcome becomes physically impossible.
Change
enabled_self_collisions=True + 2-pair filter + three-pose zero-contact audit (re-verified at 0.00 N after the later mass update) + fps budget recorded.
Outcome
Leg contact entered the training signal; the audit protocol caught the nested-pair ghost-force hazard before it corrupted training; combined with the calibrated distance wall, later versions held contact = 0 on hardware and in sim.
Mechanism
A hazard absent from the training physics cannot be learned about, no matter the reward; but collision meshes that interpenetrate at rest inject large fictitious forces if enabled blindly. Filtered enabling plus a pose-swept zero-contact audit gives true contact physics with no phantom energy - and layering prevention (reward) with consequence (physics) covers both learning and enforcement.
Applies when
- real robot self-contacts while training scored it healthy
- enabling self-collisions on a model with nested collision meshes
- deciding between reward-side and physics-side fixes for clipping
“PhysX 自动过滤相邻体(base↔hip_pitch 免费干净),幽灵力只在 calf↔ankle_roll(零位嵌套 65mm,12 倍体重)。… 与 N2 互补不冗余:N2 离得远(奖励侧预防),SC 碰了疼(物理侧兜底)—— v8 那种 107 帧互碰拿 0.751 跟踪分的事从此物理上不可能。”
train/WALK_V10_SPEC.md § 4. SC —— 训练侧自碰撞(范围已探明,比想象便宜) Before training a one-leg stand, the accounts and a probe showed the default gains could not hold it at all - kp 20 needs 0.39 rad of error to carry the static roll moment, more than the whole adduction range - so per-joint gains came first, and thermal limits set the session length
single-support-gain-authority-probeBefore training a posture that loads one joint statically, compute the steady tracking error load/kp and the series stiffness against m*g*h, and prove with a simple hand-written controller that the posture can be held under the deployment gains - change the gains first if it cannot; then size session length from the thermal account.
Symptom
The one-leg line (standing on one foot, the other folded back, no hopping) had to decide whether the existing gain profile could hold single support before any reward was designed.
Context
Hardware accounts (9.792 kg, COM 0.234 m high, 170 x 80 mm feet, legs 80% of the mass): moving the COM over one foot needs 107 mm of shift and the 20 deg hip-roll adduction range gives 131 mm - geometrically enough. The static frontal moment is 7.8-9 N*m, within RS02's 17 N*m - torque is enough. But at kp 20 carrying 7.8 N*m needs 0.39 rad of tracking error, more than the entire adduction range, and the real robot had already shown it: commanded +0.17, actual -0.04 (0.21 rad droop) under load, 0.0008 rad hanging - load, not the motor. A probe (probe_oneleg.py) then showed open-loop PD cannot hold single support on physics grounds, so the criterion became "an equilibrium exists and a hand-written 4-gain COM feedback can hold it": single-support roll stiffness is hip and ankle in series and must exceed m*g*h_com = 22.5 N*m/rad; ankle kp 12 in series with hip kp 80 gives only 10.4 (open loop 16/16 fell), ankle 60 with hip 80 gives 34.3 (52% margin).
Change
A per-joint gain profile (rl_oneleg: hip_roll kp 80, ankle_roll kp 60, the rest as rl_default) - which needed per-joint gain support in robot.yaml, the bridge, deploy and the trainer's actuator groups - decided before training. Thermal account: single support makes hip_roll the dominant heat load (about 7.8 N*m against a 7 N*m continuous rating), so acceptance and demos run in segments of at most 60 s with a temperature check.
Outcome
Under rl_oneleg the hand-written feedback held six cells cleanly for 6 s (hip_roll steady torque 2.1-3.4 N*m, half the thermal budget); under rl_default the same feedback on the same cells fell 0/4. The trained V0 policy then passed its 40-cell acceptance.
Mechanism
With PD position control, the steady error needed to carry a static load is load/kp; when that error exceeds the joint's range the posture is unreachable whatever the policy does, and series compliance between joints lowers the effective stiffness below the gravity stiffness that single support demands.
Applies when
- single-support, crouched or one-arm-load postures on PD actuators
- a joint "droops" under load on hardware but tracks well when hanging
- deciding whether a new skill needs its own gain profile
“但 kp=20 时撑住 7.8 N·m 需要 **0.39 rad 跟踪误差 > 整个内收行程**。真机已实测: 命令 +0.17 实际 −0.04(droop 0.21 rad),悬挂时 0.0008 rad——是负载不是电机。 … 单支撑滚转是 hip/ankle **串联**刚度,必须 > m·g·h_com = 22.5 N·m/rad;ankle kp12 串 hip80 只有 10.4(开环 16/16 全摔),60 串 80 = 34.3(裕 52%) … **rl_default 同反馈同格 0/4 全摔**(增益档必要性对照)”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §1-1 单脚站: 几何可行,卡点是 hip_roll 增益权限 / §2 A 线增益 / §5 probe 定谳 The first real-robot get-up was "very violent, kicking on the floor, dangerous" - a sim-perfect policy with no reason to be slow, unbounded absolute targets, no domain randomization and a rate limiter that filtered nothing; the task was restated as "safe, slow, transferable"
first-real-get-up-violent-stage-one-policyDo not put a get-up policy on hardware until its action is bounded (hard bound or state-anchored targets), smoothed, randomized and tested at the real pipeline's latency, and say explicitly that the task is "safe, slow and transferable" - a simulation-perfect policy optimizes only "gets up".
Symptom
On 2026-08-09 the user ran a V0-lineage recovery policy on the real robot and stopped it: very violent, kicking on the floor, dangerous. The planned next rung (a heavier torque_headroom) was never started.
Context
The spec had pre-registered that R0/R1 products stay in simulation and that the real-robot precondition was the R3 smoothing rungs plus a bridge-slew check plus a hanging protocol; the robustness (DR) rungs had not run. In simulation the policy passed 100% with a get-up of about a second. Which ONNX, which gain profile and whether a torque/joint log existed were left "to be recorded later" and never were.
Change
The V0 ladder was stopped at its best product (R3.1, sim only) and a re-rooting proposal was put to the user. The spec's four-layer account: style (the reward pays for standing early and nothing pays for slowness - HumanUP's "Stage I" get-up, "fast but unsafe ... infeasible for real-world deployment"); impact (full-range absolute targets with no hard bound, raw |a| up to 4.77, action saturation 100%, a single-step change of 0.306 saturating hip_pitch); transfer (zero DR, friction pinned at 1.0, the learned leg bracing); link (the bridge's RL slew equals vel_limit, 0.2-0.66 rad per step, while the real pipeline has 1-2 steps of time-varying latency and acceptance ran at delay 0).
Outcome
The line was re-rooted twice (training-side rate limit, then the beta-anchored action space) and gained a hang protocol before the next real attempt; on 08-11 a beta-anchored policy produced the line's first real get-up.
Mechanism
A task reward that pays for standing early selects the fastest feasible get-up; with absolute full-range targets every large target jump is a torque impulse bounded only by the clip; zero DR and braced-leg solutions do not transfer; and a limiter set at the velocity limit does nothing at 50 Hz.
Conflicts
The four layers are the spec's reconstruction from simulation probes and the literature; the real run's policy file, gain profile and log were never recorded, so no layer was confirmed against hardware data.
Applies when
- a first hardware trial of a high-effort skill is being scheduled
- sim success is high but the policy saturates actions or torques
- pre-registered hardware preconditions are not all met
“用户真机反馈:**非常猛、地上乱踢、危险**,叫停(R3.3 torque_headroom 加档已选型 weight −0.5→−1.5,未启动)。真机细节(哪个 onnx、什么档、有无 τ/q log)**待补记** … 任务从"能起来"变成 **"安全、慢、可迁移"** … **链路层**:桥层 slew RL 档 = vel_limit(10/20/33 rad/s ≈ 每拍 0.2~0.66 rad), 对 recovery 形同虚设;真机 1~2 拍时变延迟,验收默认 delay 0。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 真机叫停与换根判决(2026-08-09) When training fails repeatedly, inject the target behavior open-loop - stop tuning rewards for an unverified behavior
open-loop-probe-before-reward-tuningAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop on the real plant/sim first, and only resume training once you hold a measured, safe, sign-verified target trajectory.
Symptom
Three sidewalk training rounds failed; hypotheses multiplied (exploration failure / wrong reference waveform / insufficient authority) with no way to pick between them by running more training.
Context
Instead of a fourth reward guess, the team wrote probe_side_ref.py: the candidate reference is injected open-loop on top of a frozen policy's output (bypassing PPO entirely), directly measuring "what happens if the robot literally does this waveform" - separating all three hypotheses in one experiment (5 seeds x 8 s per condition, several waveform families and gains). The probe immediately eliminated the authority hypothesis (full-amplitude execution, 5/5 survival) and localized the problem to the waveform/measurement side. The closing principle was written down after the saga: without a verified target behavior, tuning rewards is "在黑暗里试钥匙" (trying keys in the dark).
Change
Standing method: before opening another training rung on a failing skill, build an open-loop (or task-space) generator of the intended behavior, measure whether the physical system can express it and what it looks like - then train toward a verified, quantified target.
Outcome
The probe chain produced the verified waveform (reversed-sign triangle, half gain), quantified safe amplitude (tilt 8.2 deg at band top, foot distance clear of the wall), exposed the metric bug when probe and training disagreed, and ultimately supplied the feed-forward that made C4 pass in +100 iters.
Mechanism
Training couples exploration, reward design, and feasibility into one opaque outcome; open-loop injection cuts the loop and tests feasibility and waveform alone. A behavior demonstrated open-loop converts the remaining failure into a pure credit-assignment/reward question - and its measured trajectory becomes the reference itself.
Applies when
- repeated training failures on one skill with multiple live hypotheses
- uncertainty whether the platform can physically express the behavior
- a reference trajectory's shape/sign/amplitude is guessed, not measured
“三轮 FAIL 之后不再猜,写 train/probe_side_ref.py 把参考开环注入到策略输出之上(绕过 PPO),直接量「照这个波形做会怎样」,一次分开三个假说:甲 探索 / 乙 波形 / 丙 权限。”
train/C_LADDER_RUN.md § 3f. C4 真因定谳(开环探针) / 3k. 建议的下一步 Deployment power derating damages non-forward axes far more than forward - sweep it in sim before deploying
power-scale-hurts-nonforward-axesTreat deployment power/torque scaling as a plant parameter: evaluate the policy in sim at the exact deployment scale, expect non-dominant axes to degrade first under derating, and either deploy at the training power or train with power randomization.
Symptom
Policies deployed at power-scale 0.8 (a safety derating of commanded torque) looked fine walking forward but were weak at backward and turning, inviting the wrong diagnosis "the skill was not trained well".
Context
Measured repeatedly: on s1e, going 1.0 -> 0.8 cost forward 18% but backward 58%; on C4-ff800, turn tracking was +25%/+40% at pw0.8 vs +75%/+58% at pw1.0, backward 51-52% vs 97-103%, while forward stayed 96-98% at both. Sim evaluation numbers in the plan were all pw1.0, but the robot was being run at 0.8.
Change
Pre-deploy protocol added: sweep the exported policy across power in sim (for PW in 0.8 0.9 1.0: eval_c_matrix --power $PW --seeds 20) and deploy at the first level where both turn directions reach >=50%. For C4 the recommendation was raise the robot to pw1.0 - the sweep showed it nearly free (saturation 47%->33%, left foot-clipping danger zone 25%->6%, cost only tilt 6.7->8.3 deg).
Outcome
Turning "weakness" resolved without any retraining; the sim sweep correctly predicted the real-robot signature at both power levels.
Mechanism
Forward walking is the reward-dominant, torque-cheapest skill with the most margin; backward/turn/sidewalk live closer to the torque envelope, so a uniform torque derating consumes their margin first. Training ran at power 1.0 (the trainer does no power scaling), so deploying at 0.8 is a systematic underactuation the policy never experienced.
Applies when
- deploying with any torque/power derating or safety scale
- secondary skills (backward, turn, lateral) underperform on hardware while forward walking looks fine
- choosing the deployment power level for a new policy
“power 衰减对非前进轴的伤害远大于前进轴(s1e:前进 1.0→0.8 掉 18%,后退掉 58%)。转向是非前进轴,0.8 下很可能明显跟不动。”
train/C_LADDER_RUN.md § 3c. A-2 上机前先定部署力度档 / 3p. 二 Two simulators disagreed 1.8x on one policy's torque demand - four explanations were eliminated with numbers, the surviving suspect was never tested, and neither reading was allowed to cancel the other
torque-disagreement-between-simulators-unresolvedWhen two simulators disagree on a safety-relevant quantity, eliminate the measurement explanations one at a time with numbers, name the survivor as a hypothesis, keep the pessimistic reading binding, and run the direct test (replay one action sequence open-loop through both plants) before the quantity is used to pass a gate for hardware.
Symptom
For R3.0, Isaac read hip_pitch torque demand at 75-80% of the limit (PASS); MuJoCo read the same policy at 148-152% (1.5x over the limit).
Context
Candidates were eliminated one by one: PD gains (identical, RS06 kp 30 / kd 1.5), the settle window (both skip the first 0.5 s), the statistic (worse-of-two-legs vs per-joint - explains ~15%), self-collision (the no-contact subset still read 152%), sampling (500 Hz peak vs 50 Hz sample - ~9%). About 1.8x remained. The prime suspect - Isaac's implicit actuator (PD solved inside the PhysX integration, chosen to mimic kHz firmware PD) against MuJoCo's 500 Hz explicit PD - was named but untested.
Change
The Isaac PASS was recorded as valid for the Isaac plant only; the prescribed decider was an open-loop replay of one action sequence through both plants, compared step by step, needing no training. It was upgraded to a precondition of R3.1.
Outcome
R3.1 ran in parallel with, not after, the replay, and the replay never appears as done. By §40 it was "the oldest open account" on the line, to be fed by real-robot logs - which the spec also never records. Meanwhile the 50 Hz sample under-read the peak by 34% as smoothing narrowed the spikes, so Isaac-side readings grew more optimistic exactly when the comparison mattered more.
Mechanism
Each measurement artifact explained a slice of the gap; what remained is a plant difference, and an untested plant hypothesis cannot license discarding the pessimistic simulator on a safety-relevant quantity.
Conflicts
The spec prescribes the open-loop replay (§23), upgrades it to an R3.1 precondition, then records that R3.1 ran in parallel without it (§24), and lists it as the oldest open account before hardware (§40); nothing through §50 (2026-08-14) records a result. Every later "torque gate PASS" on this line is an Isaac-plant reading plus a MuJoCo check, never a reconciled one.
Applies when
- Isaac and MuJoCo (or sim and hardware) report different torques, contacts or slips
- implicit vs explicit actuator models are in play
- a gate passes on one simulator only
“**同一个策略,一侧判 PASS 一侧判超限 1.5 倍。** … 口径项全部扣掉后仍剩 **~1.8×** 没有解释。 … 但**这条没有验证,不能拿它当结论去抵消 MuJoCo 的读数**。 … (做法:拿同一条 动作序列在两边开环回放,逐拍比 τ —— 不需要重训)”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §23 MuJoCo 复核门:成功率更稳,但 τ 与 Isaac 差 ~1.8× 没收敛 Ideal PD is not enough - add a delay buffer and fit armature/friction/delay per joint
actuator-delay-buffer-fittingNever ship ideal PD to hardware: add a measured delay (in control steps) and per-joint armature/friction fitted from step and sine responses, and treat remaining actuator mismatch as your standing largest sim2real residual.
Symptom
Standard ideal PD actuator model transfers poorly; sim assumes targets take effect instantly and joints reach arbitrary acceleration.
Context
A developer with a successful on-hardware Isaac Lab biped modified the actuator model in two ways and calibrated it against the real robot: step-response plus sine-sweep tests (positive step, negative step, sine tracking), overlaying sim curves on measured curves and hand-tuning.
Change
(1) Delay buffer: action targets take effect after a uniform 6 time-step delay on all joints; (2) acceleration limiting so the actuator cannot reach arbitrary acceleration; (3) per-joint fit of armature / friction / delay - different joints genuinely needed different values.
Outcome
Hip joints fit worst, knee best; the developer rated the result "not perfect, the best I could do" and still listed actuator-model improvement as next work - i.e. even the fitted model remained the dominant residual.
Mechanism
Real actuation is a lagged, bandwidth-limited system; a delay buffer and acceleration cap are the two cheapest structures that reproduce its phase and magnitude response. Per-joint differences come from differing load, wiring, and friction states, so a single global constant underfits.
Applies when
- actuator model in sim is ideal PD with no delay
- step-response of real joint visibly lags or overshoots the sim's
- budgeting which sim2real gap to attack first
“标准 ideal PD actuator 不够用,他改了两处:延迟缓冲:目标不是立即生效,全部关节统一 6 个 time step 延迟 / 加速度曲线:执行器不能瞬间达到任意加速度 … 用 armature / friction / delay 三个参数逐关节拟合,标定方法是阶跃响应 + 正弦扫描 … 髋部关节偏差最大,膝关节最好。”
Experience.md § 执行器建模 —— 最值得抄的一条 (lines 50-59) A policy's gain profile is part of its contract - the one-leg policy needs per-joint gains the default profile lacks, and the manifest refused an evaluation under the default once; the recovery contract's beta was never stamped, a known gap not to repeat
gain-profile-belongs-in-the-stampStamp everything that defines the closed loop a policy was trained in - gains included - into its manifest, and make every consumer refuse a mismatch; a profile field that is not in the stamp is a silent misconfiguration waiting for an operator to forget a flag.
Symptom
A policy trained with hip_roll kp 80 and ankle_roll kp 60 behaves differently, or falls, under the default kp 20/12 profile - and the gain profile is a command-line flag an operator can forget.
Context
The one-leg line added a gain_profile field to the contract so the stamped manifest carries it; the spec's deployment note says the manifest guard blocks rl_default and that it had already bitten once in simulation (an evaluation run without the one-leg profile). The same spec states the general rule - any new profile field must be synced into the manifest builder - and names the counter-example: the recovery line's beta was never put into the manifest. The recovery line itself had decided that its anchored authority is computed from the base rl gains and written into the contract so it cannot drift with the gain flag, and that the older kp x 0.9 profile chosen in the V0 era does not match the beta contract and must not be used.
Change
Gain profile as a contract field checked at load; per-contract gain choices written into the run sheets.
Outcome
Evaluations and hardware runs of the one-leg policy run under rl_oneleg or are refused; the recovery beta gap stayed recorded as known.
Mechanism
A policy is trained against a closed loop whose gains are part of the plant; running it under other gains is an out-of-distribution plant, exactly like a wrong observation scale.
Applies when
- a skill introduces per-joint or skill-specific gains
- deployment gains are chosen by a command-line flag
- adding any new field to a policy profile
“增益档 `--profile rl_oneleg` 必须给 —— manifest 防线会拦 `rl_default`(sim 已咬合一次) … (recovery 的 β 未进 manifest 是已知缺口,不再复制)”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §9b AGX 真机手顺 要点 / §3 契约 Prove a new penalty actually fires - two ways a clearance term silently did nothing
inert-reward-term-auditBefore training with a new reward term, log its realized per-step value under the current policy and confirm it is nonzero where intended - check coordinate zero-points against FK and check who occupies the term's gate; and never weaken the term that creates the states your new term needs.
Symptom
A newly designed swing-height clearance penalty could have trained as a no-op twice over, and the companion advice to lower feet_air_time actively backfired when tried.
Context
Instance 1 (zero-point offset): the proposed code used body_pos_w of the foot link, but that is the ankle_roll_link frame origin, which sits 0.0585 m above the ground even with the foot flat on it - so (0.03 - 0.0585) is always negative and the penalty is永远 0; the 0.0585 offset must be subtracted (verified identical in MuJoCo FK and Isaac). Instance 2 (gate occupancy): the clearance penalty fires only in swing phase; a dragging policy keeps both feet in contact, so the penalty is constantly 0 for exactly the policy it was meant to fix - and worse, any slight lift immediately incurs it, a reverse threshold. Lowering feet_air_time to 0.5 on that advice measurably collapsed air time to 0.0002 (below v3). Corrected understanding: "clearance 是把已有的摆动相抬高, 造出摆动相仍要靠 air_time" - air_time creates the swing phase, clearance raises it.
Change
Fixed the height zero-point; kept feet_air_time as the swing-phase creator with clearance layered on top; both errors documented as corrections to the team's own earlier advice.
Outcome
With both fixed, swing height rose from 22-23 mm (v2) to 29 mm (v5) to 34 mm (v6); the inert-term failure class entered the standing checklist.
Mechanism
A penalty's gradient exists only where its gate is occupied and its argument crosses its threshold; frame offsets shift the threshold out of reach, and phase gates can have zero occupancy under exactly the policy being treated. Terms interact as an ecology - one term must create the states in which another can act.
Applies when
- adding any gated or thresholded penalty (clearance, impact, slip)
- a new term produces no behavioral change at any weight
- body-frame positions are used in reward code
“body_pos_w 是 ankle_roll_link 坐标系原点,平放触地时仍高出地面 0.0585 m。… (0.03 − 0.0585) 恒为负 → 惩罚永远是 0 … clearance 惩罚只在摆动相生效,拖地时两脚始终触地 → 惩罚恒 0;而一旦轻微抬脚就立刻扣分,对正在拖地的策略是反向门槛。… 正确认识:clearance 是"把已有的摆动相抬高",造出摆动相仍要靠 air_time。”
train/WALK_DIAGNOSIS.md § walk_v4 独立验收 — 本文档给的两处代码/建议是错的 Fix a too-deep nominal pose before adding any penalties - the default stance defines the basin training starts in
nominal-posture-before-penaltiesBefore tuning penalties on a degenerate gait, audit the nominal pose and height targets against morphology and published ratios; if the default stance encodes the degenerate behavior, fix it first - and recompute dependent quantities (init height) by FK, not by hand.
Symptom
Policy lived in a crouched shuffle; nominal knee angle was 0.5 rad (28.6 deg) - deeper than published configs (Unitree G1 0.3 rad / 17.2 deg, Booster T1 0.4 rad) - so the policy's starting point and its action-space center both sat inside the crouch basin.
Context
Initially ranked "secondary" in the local diagnosis, this was promoted to co-first priority by the cross-check against published reward tables, which states that with nominal knee flexion above ~0.4 rad, fixing the posture must precede adding any penalties ("改这个之前别加任何 惩罚都是白费"). Companion base-height items: walk profile had weakened base_height_l2 to -5.0 (base class -10, field standard -10 to -20, "the second most common cause of death"), and the height target must be the STANDING height (0.384), not the crouch height.
Change
Nominal knee 0.5 -> 0.3 rad with init_base_height recomputed by MuJoCo FK (0.3739 -> 0.3802); base_height_l2 restored to -10 with standing height target; both bundled as first-priority alongside the clearance term.
Outcome
Part of the v5/v6 package that lifted swing height to 34 mm and tracking to 87%; the crouch basin stopped being the default answer.
Mechanism
The nominal pose is the fixed point every regularizer pulls toward and the point where action=0 lands; if that point is itself the degenerate posture, every penalty fights the geometry. Correcting the attractor is prior to shaping the gradient field around it.
Applies when
- policy converges to a crouched or collapsed posture
- nominal joint angles were chosen for stability rather than gait
- base-height reward targets or weights were locally weakened
“研究明确说"nominal 膝屈超过 ~0.4 rad 必须先改,改这个之前别加任何惩罚"。我们是 0.50,超标。… base_height_l2 在 walk profile 里被减到 −5.0(基类是 −10)。研究说这是"第二常见死因"且应 −10 ~ −20。改回 −10。目标高度用站立高 0.384 是对的(研究要求 target 必须是*站立*高度而非蹲姿)。”
train/WALK_DIAGNOSIS.md § 修正 ①(升级优先级) / 修正 ④ The standing-pose reward had been pulling toward the narrow stance the whole line was fighting - a zero-training kinematic audit of the target vector found it, after first auditing the wrong nominal
pose-target-geometric-auditBefore training on a pose target, audit it with forward kinematics - is it geometrically consistent (feet flat, intended stance, intended width) and is it the frame you think it is (action nominal vs standing default)? A posture term's target may itself be the attractor you are fighting.
Symptom
Several rungs aimed at widening the stance failed; the stance stayed narrow as if something kept pulling it back.
Context
The audit (MuJoCo forward kinematics, no training): every 5 deg of hip roll widens the stance ~5.5 cm (0.271 m at 5 deg, 0.383 m at 15 deg); at 47 deg of hip yaw a wide stance cannot be flat-footed (residual foot tilt ~0.7 x hip roll), which explained the stalled rungs. The first report also said the stand_pose nominal (hip roll 25, knee 60) has a 63 deg residual foot tilt - but that was the action frame's nominal (the limit-midpoint squat), which stand_pose never used, despite a docstring warning not to mix them. stand_pose's real target was DEFAULT_JOINT_POS: the contract's all-zero pose, legs parallel, ~0.22 m apart.
Change
The disease statement was corrected in writing: the narrow stance was not an accidental by-product of proxy traps but the target stand_pose had been actively rewarding. A stored "narrow the stance" knife was marked toxic. V2.8 moved the target to a flat 15-deg stance (sigma 3 -> 1.5, flat_feet margin 5 -> 20 deg).
Outcome
V2.8 still failed in-lineage (stance unchanged, feet nearly overlapping, mu 0.4 transfer 2%) - see stance-decided-by-get-up-path - and the from-scratch V3.1 removed both roll joints from stand_pose and put width into a task-space term, which is what finally produced a 0.355 m flat stance. The 63 deg finding was kept as a warning: an action nominal used as a standing target would be a ready-made pit.
Mechanism
A posture term with a sharp kernel around the wrong target is an active attractor; every other term fighting it pays twice.
Applies when
- a posture keeps returning despite penalties against it
- a reward uses a default or nominal pose as its target
- the contract has more than one "nominal" (action frame vs standing pose)
“上文"stand_pose 的 nominal (hip25/knee60) 残倾 63°"**审计错了对象**:那是 **动作参考系 nominal**(限位中点蹲),stand_pose 从未指向它(函数 docstring 原文即警告"两者别混",还是混了 —— 记档)。 … **修正后的病根陈述:窄站距不是代理陷阱的意外副产物,而是 stand_pose 一直在主动奖励的目标本身**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 勘误(实现时抓到):审计混了两个 nominal —— 真病根比误诊的更直白 FK-verify a borrowed reference's structure, then size its amplitude by the reference's job - it pins phase, the policy adds lift
reference-structure-fk-amplitude-divisionWhen borrowing a reference trajectory: verify its structural claim against your own kinematics (an invariant like flat-foot), assign it the phase-pinning job, and size amplitude low enough that the policy contributes the lift - moving toward a proven foreign value in halves, not jumps.
Symptom
walk_v4 had big knee swing (40-46 deg) but only 18-24 mm foot lift - amplitude without hip/knee/ankle phase coordination; later, walk_v5's real-robot swing ballooned to 73.6 deg (sim 55.7) with violent footfalls - amplitude over-driven by the reference.
Context
Structure first: Humanoid-Gym's 1:2:1 hip:knee:ankle reference was verified on the local model before adoption - the ratio exactly satisfies the locally derived flat-foot constraint hip - knee + ankle = 0, FK-tested at multiple amplitudes with sole pitch 0.00 deg throughout. Amplitude second, and here the first reasoning failed honestly: FK said shorter legs need LARGER reference scale (0.30 for 30 mm lift), and the FK was correct - but the premise was wrong ("FK 没错, 但前提错了"): it assumed foot lift must come from the reference. HighTorque Pi, same scale, uses 0.08 with a 0.02 m foot-height target - proof that lift is added by the policy ON TOP of the reference, whose actual job is pinning the phase relationship. Scale 0.30 made the reference the entire gait: over-constrained and over-driven. The correction went to 0.15, deliberately not Pi's 0.08: "一次只走一半, 留退路" (walk half the distance, keep a retreat).
Change
target_joint_pos_scale 0.30 -> 0.15 as one of v6-minimal's three changes, treating both the footfall force and the lateral kicking (yaw momentum scales with leg swing amplitude).
Outcome
v6 improved landing force 1.72x -> 1.55x, suspended tilt 45.9 -> 23.0 deg, turn-gain asymmetry 70% -> 19%; the later v6-halved-shaping experiment (35 mm -> 4 mm collapse) confirmed the reference still carries the gait's existence on this machine - the division of labor is real but machine-specific.
Mechanism
A joint-space reference plays two separable roles: encoding structure (phase relations that keep the foot flat) and injecting amplitude (energy). Structure transfers across robots and is checkable by FK against an invariant; amplitude is a negotiation with the policy, and over-assigning it to the reference removes the policy's freedom to modulate lift with state.
Applies when
- importing a reference gait / imitation target from another codebase
- reference amplitude reasoning based on leg length alone
- real swing amplitude far exceeds sim's under a strong reference
“FK 没错, 但前提错了。我默认抬脚必须由参考轨迹产生。HighTorque Pi 同尺度机器人 … 用 0.08, 而它 target_feet_height = 0.02 m —— 说明抬脚是策略在参考之上加出来的, 参考只负责钉住髋/膝/踝的相位配合。我们取 0.30 等于让参考本身就是整个步态, 过约束 + 过驱动”
train/WALK_V6_MINIMAL.md § ① target_joint_pos_scale 0.30 → 0.15 A soft joint-limit penalty charged the standing pose itself - the geometric-zero knee sat on its hard limit, so stand_v1 bent its knees to dodge 0.419 per step and leaned 4.1 deg forward; excluding the knee gave 0.24 deg
soft-limit-penalty-charges-nominal-poseBefore training, evaluate every penalty at the nominal pose; if a joint's soft limit sits inside the pose the task requires (a straight knee on its hard stop), exclude that joint from the soft-limit penalty and let the action clip enforce the hard limit.
Symptom
stand_v1 (retrained after the default pose moved to the CAD geometric zero and mirror augmentation was added) fixed left/right asymmetry (6.8 -> 0.0 deg) but settled at a 4.1 deg forward lean, where pure PD at the same default settled at 0.1 deg - the policy was actively pushing itself forward, which is exactly the real robot's failure direction.
Context
soft_joint_pos_limit_factor = 0.9 shrank the knee's soft limit to -/+0.1047 rad, while the geometric-zero default has the knee at q = 0, exactly on the hard limit. Standing in the nominal pose therefore paid 0.2094 x 2.0 = 0.419 per step in dof_pos_limits (alive earned only 0.5). The policy's way out was to bend the knees to -/+0.1013 rad, and the cost was the forward lean.
Change
stand_v1b: the knees excluded from dof_pos_limits (a straight knee IS the standing pose; the hard limit is still enforced by the action clip). No other change.
Outcome
stand_v1b: max tilt 0.3 deg, steady tilt 0.24 deg, asymmetry 0.1 deg, height 0.384 m - exactly nominal - with knees at -0.0007 / +0.0005 rad. It became the standing release used on the real robot, and later the standing side of the recovery switch. The same exclusion was carried into the recovery contract (knee at the clip in the standing pose) and the one-leg reward table.
Mechanism
A limit penalty whose soft boundary lies inside the nominal pose turns the nominal into a taxed state, and the policy buys its way out with whatever posture change is cheapest - here a knee bend paid for with lean.
Applies when
- a standing or default pose has a joint at or near its hard limit
- a policy settles in a small steady tilt that pure PD does not show
- soft-limit factors shrink limits uniformly across joints
“`soft_joint_pos_limit_factor=0.9` 把膝软限位内缩到 ∓0.1047,而几何零位 default **膝盖 q=0 正好压在硬限位上** ⇒ 站在标称姿态每步白扣 `0.2094 × 2.0 = 0.419` (alive 才 +0.5)。策略只能屈膝到 ∓0.1013 躲罚,代价是躯干前倾 —— 恰好是真机的 失效方向。 … 修掉"软限位罚标称姿态"后重训(`dof_pos_limits` 排除膝盖)。**前倾问题彻底消失** … 不再屈膝躲惩罚,高度正好落回标称 0.3840。”
train/README.md § 三期: 镜像对称增强 + 站立 v1 (2026-07-28) / stand_v1b (2026-07-28): 站立定版 A 2-degree joint-zero calibration fix moved the whole runnable envelope - re-test old "cannot run" verdicts after recalibration
zero-offset-calibration-shifts-envelopeDate every hardware verdict with the calibration state; after any zero/mount recalibration, re-test previously condemned policy-power combinations and previously "unexplainable" posture offsets before attributing either to training or model.
Symptom
s1d was on record as "only runs at power 0.7" (kicked wildly at 0.8); after a calibration pass, the same policy ran 12 s at 0.8 with no kicking at all.
Context
The calibration had fixed a 2.08 deg zero offset on r_hip_roll - exactly the constant error source on the dominant joint of the kicking oscillation loop ("恰是乱踢振荡环主导关节的常值误差源"). The three-generation post-calibration hardware sweep also closed a second case: the robot's mysterious "backward lean" disappeared after calibration, and the sim-real posture difference collapsed from opposite-sign 5+ deg to same-sign ~2 deg ("后仰案实质了结") - the lean had been a sensing/zero artifact, not a mass-model error. Booked consequence: if the s1d recovery re-verifies, "真机可跑档整体 上移" - every policy's runnable power envelope shifts up, and downstream lineages' hardware expectations get revised.
Change
Joint-zero and mount calibration promoted from setup chore to a variable that dates hardware verdicts: verdicts about which power/scale levels a policy can run are conditioned on the calibration state they were measured under.
Outcome
One policy rehabilitated at a higher power level; one standing sim-real posture discrepancy closed without touching model or training; a pending re-verification booked rather than asserted.
Mechanism
A constant joint-zero error acts as a persistent disturbance injected at the feedback loop's most-loaded joint; near an oscillation threshold, removing a 2-degree bias is the difference between a stable and an unstable loop. Since the error is additive and machine-side, it shifts every policy's stability envelope simultaneously - which is why verdicts must carry their calibration date.
Conflicts
The s1d rehabilitation awaited one confirming re-run at the time of writing ("待复核一跑坐实") - the offset-as-cause reading is the head suspect, not a closed verdict.
Applies when
- a policy oscillates at a power level others tolerate
- sim and real disagree on a constant posture offset
- deciding whether to re-test old hardware verdicts after maintenance/calibration
“发现①:s1d@0.8 能跑了(旧账「只有 0.7 能跑」)——12s 无乱踢。头号嫌疑 = 标定修正:r_hip_roll offset 修 2.08°,恰是乱踢振荡环主导关节的常值误差源。… 发现②:「后仰」标定后消失 … sim-real 姿态差从反号 5°+ 收敛到同号 2°,后仰案实质了结。”
train/README.md § 真机 @0.8 三代横评(2026-08-07 标定后) The shipped checkpoint was chosen by scanning checkpoints on the full gate - neighbours 100 iterations apart failed 1 and 38 cells, late checkpoints degraded - never by taking the last one, and training stopped on signals, not on a schedule
checkpoint-choice-is-a-full-gate-scanChoose a release checkpoint by running the full acceptance battery over a band of checkpoints (including the transfer axis), stop training on measured signals rather than a fixed iteration count, and expect adjacent checkpoints to differ sharply.
Symptom
Gate results moved sharply and non-monotonically between checkpoints of the same run, and the last checkpoint was often not the best.
Context
One-leg V0r1: the 2,000 neighbourhood was best; from 2,500 on the nominal gates degraded (late overtraining); 2,000 itself had one real micro-hop (17.7 mm over 5 frames); 2,300 was all green and shipped. V0r2: failed cells per checkpoint 2,000:19, 2,100:38, 2,200:1, 2,300:3, 2,400:27, 2,500:12, 3,000:18 - 2,200 shipped. The recovery line learned the same from the other side: stopping v2_6 early at a scheduled point left a policy whose re-fall rate had spiked to 9-22% before consolidation healed it ("stop on signals, not on the schedule"), and a continuation's transfer decayed checkpoint by checkpoint while Isaac stayed perfect.
Change
The acceptance rule "scan checkpoints, do not look only at the last one" is written into the one-leg gates (called the S1 discipline); release candidates are chosen from the scan.
Outcome
Both one-leg releases were mid-run checkpoints (2,300 and 2,200) chosen by the full 40-cell battery.
Mechanism
PPO keeps changing the policy after the gates saturate; with no gradient toward the gate's conditions, later checkpoints wander, so gate quality is a noisy function of iteration.
Applies when
- picking which checkpoint of a run to export and stamp
- a run is stopped at a fixed iteration budget
- final-checkpoint results are worse than mid-run smoke tests
“Isaac 侧 S1 纪律: 验收扫 checkpoint,不是只看最后一个。 … 扫描判决: 2000 邻域最优——2500+ 标称面退化(⑤③② 散挂, 晚期过训), 2000 有一例真微跳(L s100 μ1.2, 17.7mm/5帧), 2300 全绿。”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §6 验收门 / §8 核查单 5 与 7 The restart reward table lists a reason for every term AND a lesson for every exclusion - absent terms are removed, not zero-weighted
minimal-reward-table-with-provenanceMaintain the reward table as an evidence ledger: every term cites the episode that justifies it, every excluded term cites the episode that convicted it (including the development stage it is valid at), and retired terms are deleted from the config, never left at weight zero.
Symptom
Seven walk generations had accumulated an entangled reward table where nobody could say which term earned its place; the restart needed a table that could be audited line by line.
Context
The minimal table v2 was built under three written principles: "一项管 一件事、结构性反抬脚的项一个不留、塑形只留一套相位逻辑;不在表里的一律 不加" (one term per job; zero structurally-anti-lift terms; exactly one phase-shaping logic; nothing outside the table gets added). Every row carries its provenance (e.g. base_height target = standing height cites the v4 crouch lesson; split x/y tracking cites the merged-exp gradient hole; world-frame yaw cites the v1/v3 body-frame lesson). Every EXCLUSION carries its same-type precedent: feet_landing_vel out because it is poison while the gait is unformed (drag pays 0, lifting pays - the v4-clearance / v8a-B "reverse threshold" family) though it was fine in v7 when the gait already existed - term validity depends on developmental stage; feet_air_time out with its measured non-lever evidence (v6 had it at 2.0 and still lifted 4 mm); and replaced terms are REMOVED from the config ("置 None,不是权重 0 挂着") so audits see truth, not dormant weight.
Change
Reward table rebuilt as ~20 rows each with weight + provenance column; exclusion list maintained alongside with the falsifying episode for each; dormant terms deleted rather than zeroed.
Outcome
Later revisions (S1.2/S1.3) modified the table by citing and updating specific rows' evidence rather than re-arguing the whole design; the table doubled as the lineage's reward-lesson index.
Mechanism
A reward table is a set of standing hypotheses; attaching each row's evidence makes revisions targeted and reversible, and recording why a term is absent prevents the cycle of re-adding known poisons. Deleting vs zero-weighting matters because config audits and DR interactions see the term either way - a zero-weight term is dormant complexity waiting to be flipped on wrongly.
Applies when
- designing a reward table for a restart or new task
- someone proposes re-adding a previously removed term
- auditing which reward rows still earn their place
“原则:一项管一件事、结构性反抬脚的项一个不留、塑形只留一套相位逻辑;不在表里的一律不加 … feet_landing_vel(评审 #4):拖地时代价恒 0、抬脚才收费——与 v4-clearance/v8a-B 同属「反抬脚门槛」家族,步态未成形时是毒;v7④ 加它时步态已存在。… 已从 cfg 移除(置 None),不是权重 0 挂着。”
train/OMNI_V0_SPEC.md § 3. 最小奖励表 v2 / 明确不带 Training at scale 0.4 is NOT the twin of deploying 0.5 at power 0.8 - the algebra matches, the learned policy does not
deploy-scaling-not-training-equivalentNever assume deploy-side scalings can be folded into training-time constants ("burning the crutch into training"): the learned optimum depends on the training-time authority, so treat such conversions as full experiments with pre-registered expectations and a sim2sim gate before any hardware.
Symptom
s1g (S1.6) trained from zero at action_scale 0.4 - meant as the "training twin" of the hardware-proven s1c-at-power-0.8 (0.8 x 0.5 = 0.4) - was all green in Isaac (zero falls, reward 117) yet scored 0/3 across all eight checkpoints and 0/20 at 20 seeds in the MuJoCo gate, falling forward at median 1.57 s with a 2.9x speed overshoot.
Context
The pre-registered expectation (survival gate should pass, since the conviction matrix showed s1c@0.8+delay2 all-survive) was cleanly falsified, and the harness was acquitted by controls: --delay 0 fell identically (not a delay fragility), check_contract all green, and s1c through the same harness survived 2/3. The verdict: "「s1c@0.8 = 0.4 训练孪生」的代数等价不成立" - a policy deployed with a derated output still LIVES in the 0.5 internal model it trained under (its value function, its expectations of its own authority), while a policy that starts training with reduced authority learns a different, clip-hugging gait with zero margin for plant differences ("部署端打折的策略活在 0.5 的内模里,训练起点收权限学出的是贴 clip 的 另一套步态,对 plant 差异零余量"). Result: the policy was withdrawn before hardware ("撤回——不上真机"), the lineage root moved back to the 0.5-contract s1c-5500, and this became the C ladder's cited fact-check ("s1g 是 0/20 证伪出局的那一代").
Change
The amplitude-surgery route abandoned; contract kept at scale 0.5; the deploy-side 0.8 crutch later retired on its own merits when the delay-complete s1e generation ran at full power.
Outcome
One training run bought a clean falsification of a plausible algebraic identity; no hardware time was spent on it because the sim2sim gate caught it.
Mechanism
Output scaling commutes with the network arithmetic but not with learning: the training-time scale shapes which gait solutions are reachable and how much clip headroom the optimum keeps. A derated mature policy retains the wide-authority solution executed softly; a from-zero narrow-authority policy finds a different optimum that saturates its smaller envelope - the two are not the same controller in different units.
Applies when
- proposing to move a deployment derating into a training constant
- a scaled-down contract policy hugs the action clip
- Isaac-green / cross-sim-zero results on a re-scaled lineage
“预注册 a) 证伪——Isaac 全绿(零摔/reward 117)但 MuJoCo --delay 2 八档 checkpoint 扫描全数 0/3、iter6500 20-seed 0/20 … 「s1c@0.8 = 0.4 训练孪生」的代数等价不成立: 部署端打折的策略活在 0.5 的内模里,训练起点收权限学出的是贴 clip 的另一套步态,对 plant 差异零余量。”
train/OMNI_V0_SPEC.md § 3. S1.6 判决(2026-08-07 验收) The advisor's "runtime five-stage state machine + per-stage reference poses + RL residual" appeared in none of the three papers it cited - reading the originals changed the plan and downgraded two widely repeated industry claims
advisor-paraphrase-vs-paperRead the primary source behind any piece of advice before adopting its architecture; record where the paraphrase and the original differ, downgrade claims the originals do not support to speculation, and adopt what the verified sources actually share.
Symptom
After the violent first real run, an advisor proposed re-architecting recovery as a runtime staged state machine with reference poses and an RL residual, citing HoST, HumanUP and StableMimic.
Context
The three papers were read in full on 2026-08-09 and tabulated (deployment form, what the "stages" really are, hard constraints, references). HoST: one end-to-end policy, height-gated rewards in training, action anchored as q + beta*a with a beta curriculum. HumanUP: two training stages with the same observation/action; the vendor's three-stage state machine is the baseline it beats (41.7% vs 78.3%); Stage II tracks an 8x slowed Stage I trajectory (4x too violent, 10x does not converge). StableMimic: a learned soft gate. On 08-10 more sources were checked the same way: the Agility page does not say Digit's self-righting was learned in simulation (only step recovery is stated as RL), so the claim was downgraded to speculation; HoST's support for the 12-DoF armless Mini Pi exists in its code repository, not in the paper text.
Change
Adopted only what the originals share: a hard action bound or anchored action space, strong smoothing including a second-difference term, a slowed hidden reference, heavy DR and real fallen states. The staged runtime state machine was not adopted.
Outcome
The next re-rooting candidates came straight from the verified material, and the beta-anchored action space (HoST, with the Mini Pi configuration as the nearest real-robot precedent) became V2, the lineage that later stood up on hardware.
Mechanism
A paraphrase compresses a paper into the advisor's own architecture; only the original shows what was actually deployed, what was a baseline, and which numbers came with which ablation.
Applies when
- an advisor, agent or summary proposes an architecture with citations
- an industry claim ("X learned it in sim") is about to justify a design
- several papers are cited for one combined recipe
“⇒ **顾问的核心形态"runtime 五阶段状态机 + 每阶段参考姿态 + RL residual"在三篇引文 里均不存在**,其中 HumanUP 还点名 state machine 是局限。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 三篇引文精读判决(2026-08-09 全文核对;顾问转述与原文有出入) After seven patch-generations, freeze the lineage as a regression baseline, fix the structural debts, and retrain from zero
freeze-lineage-fix-structure-restartWhen successive rungs keep trading one symptom for another, ask whether the remaining problems are structural (contracts, latency, sampling, reward-table architecture); if so, freeze the lineage as regression baselines, pay the structural debts, and restart minimal - carrying forward laws and instruments, not weights and weights' patches.
Symptom
The v5-v11 walk lineage had accumulated interacting patches (reward terms, gates, clamps, per-joint scales) faster than it converged on the user's goal; v12's spec itself was superseded before training by an external review's verdict that the remaining problems were structural, not parametric.
Context
The 2026-08-05 status banner records the pivot: the walk profile was rolled back wholesale to v10b parameters, the v5-v11 lineage frozen "只作回归对照" (kept only as regression baselines), and four structural debts were named as prerequisites for a from-zero straight-walk baseline: the action-latency FIFO (fixed with its own test), the ONNX manifest contract, discrete command sampling, and a minimal reward table. The v12 spec - fully designed, partially implemented - was suspended: "本规格挂起,不再按此开训".
Change
Strategy switched from "one more patch generation" to freeze-fix-restart: lineage checkpoints retained as comparison anchors, infrastructure hardened first, then a clean retrain with a minimal reward table (this restart produced the s* generation that later became the real-robot SOTA line).
Outcome
A designed-and-ready training generation was deliberately not run - the review's structural findings outranked sunk design cost; the restart line inherited seven generations of laws (calibrations, gate batteries, falsified fixes) without inheriting their entangled reward table.
Mechanism
Patch lineages accumulate coupled terms whose interactions eventually cost more to reason about than a restart costs to train; the knowledge worth keeping is the laws and instruments (measured plant values, calibrated gates, falsified directions), not the entangled weights. A restart on hardened structure converts the lineage's lessons into a clean initial design instead of another delta.
Applies when
- repeated rungs shuffle symptoms without net progress
- an external review flags infrastructure/contract debts
- deciding between another patch generation and a clean retrain
“同日外部评审定调换路线:冻结 v5~v11 血统(只作回归对照),修结构性问题(latency FIFO 已修 tests/test_action_latency.py、ONNX manifest 契约、离散命令采样、最小奖励表)后从零训直行基线。本规格挂起,不再按此开训。”
train/WALK_V12_SPEC.md § ⚠️ 状态 (2026-08-05) Every power cycle starts with the same read-only pre-flight - read the buses, check the torque limits against 12/17/11, verify the IMU axes, check the ports after any new USB device - and any reassembly re-measures the joint zeros
power-cycle-preflightStart every powered session with a fixed, read-only pre-flight - bus responses, torque limits equal to the simulated ones, IMU axes, device identities - and re-measure joint zeros after any mechanical reassembly before running a policy.
Symptom
Hardware state drifts between sessions in ways no policy can see: a motor that stops answering after a power cycle, a torque limit that differs from the one simulated, an IMU axis flipped, two USB devices swapping identities, a joint zero moved by reassembly.
Context
The runbook's session order before any policy runs: read every motor on both CAN buses without enabling them (the first command after every power cycle); set_torque --check, all twelve motors must read 12/17/11 N*m, and any difference is written back; imu_reader --verify-axes, where the operator tilts the robot forward and to the right and every check must pass before continuing; check_ports after plugging in any new USB device (the IMU and a CAN adapter once collided on USB identity). After re-mounting motors: read the buses, then re-measure the calibration offsets (three repeats, written back) - "skipping it means running everything on the wrong zero". Hanging checklists repeat the torque-limit check (the deploy script also self-checks at start).
Change
A fixed, read-only pre-flight run in the same order every session.
Outcome
The runbook records one earlier hardware check in the same spirit: all 12 motors' implied kp fell within 18.4-22.0 for a commanded 20, inside the kp randomization range used in training.
Mechanism
A policy transfers only if the plant matches the one it was evaluated on; the pre-flight turns silent hardware drift into a failed check before the robot moves.
Applies when
- the first command after powering a robot on
- after swapping adapters, cables or motors
- a policy that worked last session suddenly behaves differently
“python tools/set_torque.py --check # 12 颗应全对 12/17/11, 有 diff 就 --write … 插任何新 USB 设备后都先跑一次 check_ports.py(IMU 和 CANable 的 USB 身份撞过车) … python tools/calib_stance.py --repeat 3 --write # 重标 offset —— 8/9/10 重新装, 机械零位变了”
RL系统/FOLLOW THIS copy 2.md § WALK / STAND 每次开始前 / 换CAN / 装回后必做两件 No parameter tuning on the floor - a failing config retries once, then it is out; anomalies go back to sim
no-field-tuning-protocolHardware time is for executing and measuring the pre-registered matrix, never for tuning: failing configs get one retry then elimination, anomalies get recorded and reproduced in sim, and contract-check bypass flags stay unused.
Symptom
Hardware sessions create pressure to fix problems live - nudge a gain, tweak a scale - which destroys attribution and risks the robot.
Context
The anomaly-handling section of the acceptance sheet is three fixed plays: (1) falls at start -> retry once at the same settings; falls again -> that configuration is eliminated, "不现场调参" (no on-site parameter tuning); (2) limit cycle or motor screech -> stop immediately, record the gain level and the joint, reproduce in sim before any discussion; (3) systematic disagreement with sim -> record it as a finding (hardware outranks sim) rather than adjusting anything to force agreement. Related guardrails elsewhere in the sheet: never pass --allow-unstamped / --allow-plant-drift to bypass manifest checks - if it errors, something real is wrong, stop and look.
Change
Field sessions restricted to executing the pre-written matrix; every fix path routed through sim reproduction and the normal config/rung process.
Outcome
Sessions stayed interpretable (each run matched a documented config) and safety overrides never became habit; anomalies arrived back in sim as reproducible cases instead of half-remembered floor stories.
Mechanism
Field-tuned values are measured under adrenaline on one floor with no logging or baselines - they contaminate the config lineage and are unattributable afterwards; and every bypass flag that skips a contract check converts a designed safety property into an operator promise.
Applies when
- a config fails or oscillates during a hardware session
- someone reaches for a live gain tweak or a bypass flag
- writing the anomaly-handling section of a deployment runbook
“起步即摔 → 换档重试一次, 仍摔则该档出局, 不现场调参。出现极限环/啸叫 → 立刻停, 记录档位与关节, 回 sim 复现再议。… 不要给 --allow-unstamped / --allow-plant-drift —— 三枚 ONNX 都已盖章 … 真要报错说明有别的问题, 停下来看。”
train/REAL_RUN_S2.md § 4. 异常处置 Pre-register the ladder's risks and how each future result will be read - before training
preregister-risks-and-fork-readingsBefore a training ladder or risky rung, write the risks, the stop rules, and how every plausible outcome will be interpreted - then do not edit them after seeing results.
Symptom
Without pre-registration, ladder results get rationalized after the fact; the team had already seen post-hoc reads go wrong and adopted written pre-commitment.
Context
The C-ladder execution sheet opens with three numbered pre-registered risks: (1) S2 plant robustness will not carry into omni - a full S2 redo is budgeted from the start; (2) the s1e recipe has a collapse valley at iter 1500+ (recorded twice), so every rung runs a watch_ckpt --every 100 smoke loop with stop-on-degradation; (3) the root's OOD survival edge may partly be "it is slower / commits less" - with the reading fixed in advance: if C1 training raises tracking while survival drops, the edge was bought with slowness, and root selection reopens. Later rungs went further, pre-registering a full result-to-conclusion table for the A/B arms ("预注册读法(事后不改)") and even pre-registering the author's own doubt that a level would fail and what its failure would prove.
Change
Standing practice: before each rung, write down (a) known risks with their mitigations, (b) the interpretation of each possible outcome, (c) stop criteria - all frozen before the run starts ("开训前写死,事后不许改").
Outcome
When arm-A/arm-B and redo results arrived, conclusions were read off the pre-registered table instead of argued; a predicted-likely-FAIL level (C4-redo3) was still run because its pre-registered value was eliminating the regularization hypothesis - which it did.
Mechanism
Pre-commitment converts each training run into a decisive experiment: outcomes falsify or confirm named hypotheses instead of being absorbed into a story; it also makes negative results valuable (a FAIL that eliminates a hypothesis advances the search).
Applies when
- launching a multi-rung training ladder
- running an A/B fork whose outcome will drive a fork/root decision
- a rung is expected to fail but is run for its diagnostic value
“三条预注册风险 … 若 C1 训后跟踪提上去而存活掉下来,说明这条优势是速度买的、不是通用性 → 那时重开选根。”
train/C_LADDER_RUN.md § 0. 三条预注册风险 A real-robot verdict is (policy x deployment stack) - when the stack changes materially, old verdicts expire
stale-verdicts-under-old-stackDate every hardware verdict with the deployment-stack version it was measured under; after any material stack change, re-test before trusting old condemnations or old praises - with the interpretation of each possible result written down first.
Symptom
walk_v5 stood condemned as "kicks wildly" and walk_v6 as "cannot walk unassisted" - but those verdicts were issued under an earlier deployment stack (--heading did not exist yet, several fixes had just landed); only v7 had ever run under the current unified stack.
Context
New sim evidence sharpened the doubt: a same-harness four-version sweep showed v6 was the HEALTHIEST archive at the real operating point (cmd 0.15: dominant frequency locked at 2.50, steps 35:35 perfectly symmetric, foot distance 203/190 mm best of four, mean tilt 4.2 deg, saturation 15%). A full re-test under the unified stack was scheduled with per-version questions and a pre-filled interpretation table ("判读表(预填假设,回来对号)"): e.g. v6 walks + frequency ~2.5 -> old verdict was the stack's fault, v6 becomes the comparison champion; v6 walks but at ~1.25 -> period-doubling on hardware = confirmed plant gap, actuator fitting promoted to mainline; v5 no longer kicks -> the kicking was an old-stack artifact.
Change
All four versions re-queued on hardware under one stack (same torque limits, slew profile, heading loop, logging), with the version-specific legacy profile pinned; verdicts held provisional until re-issued.
Outcome
The re-test design separated policy properties from stack artifacts before any policy was permanently written off - and turned each outcome into a specific conclusion via the pre-filled table.
Mechanism
A deployed behavior is produced by the policy plus everything between it and the motors (heading loop, slew limits, torque caps, clock); verdicts implicitly condition on that whole stack. Fixing the stack invalidates the conditioning, so old failures may be stack artifacts and old successes may not survive either.
Applies when
- deployment tooling (limits, filters, loops) changed since a policy was last judged
- deciding which historical policy is the rightful baseline
- a sim sweep contradicts an old hardware verdict
“只有 v7 在完整的今日部署栈下上过真机 … v5"左右乱踢"、v6"未能自主"的判决全部来自更早的栈(--heading 尚不存在, 部分修复刚落地)——判决已过期。且 2026-08-02 四代同机仿真横测翻出了新证据:v6 在真机工况(cmd 0.15)下是四代里最健康的仿真档案”
train/REAL_SWEEP_V5_V8.md § 0. 为什么重测 Every acceptance run records video of the very rollout that produced the numbers - the seated basin, edge-standing feet, tangled legs and the narrow stance were all seen on video before, or instead of, a metric catching them
video-as-acceptance-recordMake video a default output of every acceptance run, rendered from the same rollout the metrics come from (fixed views including the feet), and watch it - posture failures are visible before any gate row exists for them; never let video replace or override the numeric gate.
Symptom
Posture failures in the recovery line kept arriving as things the numbers had no row for: a kneeling W-sit, feet standing on their outer edges, crossed legs, a stance too narrow to hold on hardware.
Context
From 2026-08-09 (user decision) accept_recovery renders offscreen by default, following one env for the whole episode and archiving the clip. The MuJoCo gate's --video renders three views (side, front, feet) of the same rollout the metrics come from, with all plant modelling (delay, push); the older replay-based renderer produced an independent trajectory without delay and was not used for acceptance. Video never gates: if rendering breaks, --no-video keeps the numeric gate running.
Change
Video as a default acceptance artifact, named per policy, category and view, reviewed by the user.
Outcome
The R0.1 prone clip showed the same kneel-sit as R0; the R3.1 failure clip showed the crossed legs; the v2_5 feet view showed edge standing and led to V2.6; the v2_6 videos led the user to order a real-robot A/B between v2_5b and v2_6. Once, the recorder did not start (P1c final acceptance) and the visual material had to be produced separately.
Mechanism
Metrics exist only for failure modes someone anticipated; video shows the unanticipated ones, and rendering the metric rollout itself guarantees the picture and the numbers describe the same episode.
Applies when
- setting up an acceptance pipeline for posture-sensitive skills
- numbers pass but a human reviewer is uneasy
- sim videos are rendered by a separate replay tool
“**验收存视频(用户定 2026-08-09)**:`accept_recovery.py` 默认开 Isaac 离屏 渲染,跟拍一个 env 的整局并归档 … 跪坐这类盆地在数字表出现前肉眼先看见,视频是验收的 定性存档,数字门不受它影响”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §5 R0 验收门(预注册)验收存视频(用户定 2026-08-09)