Training Coach
Doctrine
A report may cite any of these as doctrine-N.
doctrine-1Contract freeze and fingerprint disciplineThe policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.
Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).
Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.
doctrine-2Attribution by resolved training params - never eval-override knobsCapability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.
Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).
Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.
doctrine-3PASS gates become constraints; FAIL gates become objectivesOnce a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.
Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.
Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".
doctrine-4One variable per ladder rung - counted against what the checkpoint sawA rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.
Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.
Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.
doctrine-5Pre-register risks, readings, and stop criteria before the ladderBefore a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.
Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).
Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.
doctrine-6Plant parameters are measured, never inventedEvery plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.
Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).
Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.
doctrine-7Sim2sim gate before sim2real - under deployment conditionsEvery checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).
Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).
Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.
doctrine-8Observation honesty - the actor's inputs are a hardware contractThe actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.
Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).
Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.
doctrine-9Reward economics are audited in realized currencyReward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.
Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).
Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.
doctrine-10The zero-cost option must be the desired behaviorFor every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.
Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).
Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.
doctrine-11Measurement discipline: independent referees, signs, distributionsA disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.
Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).
Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.
doctrine-12The deployment pipeline is plantIrreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.
Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).
Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.
doctrine-13DR budget is finite; its distribution is the measured supportRobustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.
Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).
Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.
doctrine-14Gates measure what hardware feels: posture, margins, stripped assistsAcceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.
Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).
Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.
doctrine-15Fork and root selection: recoverability, maturity, frozen rewardsChoose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.
Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).
Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.
doctrine-16Curricula: verified engagement, lineage counters, disease-phase gatingAutomatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.
Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).
Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.
doctrine-17Probe before training: feasibility first, hypotheses in tablesAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.
Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).
Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).
doctrine-18External advice is recomputed locally; values transfer as ratiosEvery external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.
Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).
Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.
doctrine-19Hardware sessions are scripted experiments, not tuning sessionsReal-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.
Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).
Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.
doctrine-20Close questions in writing; restart when the debt is structuralAudited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.
Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).
Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.
doctrine-21Name the quantity in the space it lives inA goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.
Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).
Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.
doctrine-22Continuation needs a live gradient; a release is chosen by a scanContinue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.
Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).
Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.
Experience cards
164 cards matching “same-sign-response-is-yaw-bias”.
Teleop fed the sidewalk axis a command beyond its training band - feet clipped; give each axis its own speed setting
teleop-command-band-per-axisGive every command axis its own teleop scale, clamped to that axis's training band, and reproduce any hardware incident in sim with the exact deployed command values before touching training.
Symptom
Robot stepped on its own foot when sidewalking left under teleop - and only when going left.
Context
The teleop tool used one speed setting for all axes: --teleop-speed 0.20 applied to A/D sent cmd_vy = 0.20, above the training band's top (0.08-0.18) where foot-spacing margin is thinnest. Sim reproduction of the incident (product policy, pw0.8, 5 seeds x 20 s, true collision threshold = single foot width 104 mm): at vy 0.20 the minimum foot distance was 111-115 mm - 7-11 mm from self-collision - vs 147 mm at vy 0.10. Left was 4x more dangerous than right (25% vs 6% of time inside the 160 mm soft wall at vy 0.10), matching the left-only symptom; the margin did not degrade over time (pressing more just lengthened exposure).
Change
deploy_policy gained --teleop-side (default 0.10), separating the lateral speed from the forward speed so each axis's teleop command sits inside its own trained band.
Outcome
Command now inside the band with 43 mm margin at default; the incident became a quantified, reproduced, closed account rather than a mystery.
Mechanism
The policy's competence envelope is the training command distribution per axis; teleop mappings that share one scalar across axes silently command out-of-band inputs on the weakest axis. Asymmetric risk (left vs right) came from the policy's own chirality bias, so a symmetric command produced an asymmetric hazard.
Applies when
- wiring a joystick/teleop layer over a learned policy
- a hardware incident occurs on one command direction only
- training bands differ across command axes
“A/D 一直与 W/S 共用速度档,所以按 A 下发的是 vy = 0.20 —— 既超训练带(0.08~0.18)上沿 … 0.20(遥控实际值)| 111~115 mm | 7~11 mm … 且左比右危险 4 倍 … 处置:deploy_policy 新增 --teleop-side(默认 0.10),侧移与前进档分开。”
train/C_LADDER_RUN.md § 3p. 一 向左走踩到自己 → --teleop-speed 0.20 同时喂给了 vy Run acceptance under measured contact parameters - honest condim/torsional-friction flipped a false PASS into a real-matching FAIL
eval-plant-honesty-contact-paramsPin the evaluation plant's contact model to measured values (contact dimension, torsional/rolling friction, mu) before trusting any gate that involves slip, impact, or drift - a gate can only fail a policy for physics its simulator contains.
Symptom
walk_v5 passed the old acceptance battery yet failed on the real robot (footfall force, drift, kicking) - the evaluation plant was flattering the policy.
Context
The battery was re-run under "honest contact parameters" - condim 4 (adding torsional contact) with measured torsional friction 0.035 - and v5 then FAILED exactly the rows corresponding to its real problems: heading 185 deg (limit 30), support-foot yaw slip 284 deg (limit 80), landing force 1.72x (limit 1.5x), suspended tilt 45.9 deg (limit 10). The slip physics depends on torsional friction, which the default contact model (condim 3) does not even simulate - a slip problem is invisible to an evaluator that cannot represent yaw friction at the foot. Term-sizing measurements for the new rewards were likewise taken under the same honest parameters (cmd 0.45, skipping the 5 s start transient).
Change
Acceptance harness pinned to condim 4 / torsion 0.035 (measured); verdicts issued under defaults declared non-citable for these rows.
Outcome
Sim acceptance verdicts began agreeing with hardware ("现在失败, 与真机一致"); the v6 fixes could be developed and validated against an evaluator that could actually see the disease.
Mechanism
An evaluator is a plant model too: contact dimensionality and friction values decide which failure modes exist in the simulation at all. Evaluating under default contact parameters tests the policy in a world where its real failure is physically impossible, producing structurally false PASSes.
Applies when
- sim acceptance passes policies that fail on hardware
- slip/drift/impact gates run under default simulator contact settings
- setting up a cross-simulator evaluation harness
“accept_v2.py 已加三条判据, walk_v5 在诚实的接触参数下(--condim 4 --torsion 0.035)现在失败, 与真机一致:直行 15s 航向累计 <30° | 185° ✗ … 落脚力峰值 <1.5× 体重 | 1.72× ✗”
train/WALK_V6_MINIMAL.md § 5. 验收 The 1.5 Hz step-frequency gate was retracted - the author had misread his own actuator data, and the limit fought pendulum dynamics
gate-threshold-retracted-frequencyEvery gate threshold must cite its measurement and survive a first-principles sanity check; when a gate keeps failing otherwise healthy behavior, re-derive the threshold from the raw data before enforcing it again - and retract wrong gates in writing.
Symptom
An acceptance criterion "step frequency <= 1.5 Hz" kept failing healthy policies (walk_v4 at 2.33 Hz), and an earlier attempt to force slower stepping (v3) had killed stepping altogether.
Context
The threshold had been derived from the author's own actuator frequency-response measurements - but re-reading the raw table showed the misread: amplitude ratio at 2.0 Hz is 0.88 (knee) / 0.83 (ankle), acceptable; the genuinely bad point was walk_v2's 3.8 Hz at 0.58. Mechanism check agreed: the leg as a compound pendulum (L ~ 0.30 m) has a natural frequency ~1.1 Hz, swing half-period 0.45 s - the observed 2.1-2.4 Hz sits near where the leg wants to swing, and forcing 1.5 Hz "是跟摆动动力学对着干" (fights the swing dynamics). The frequency definition itself was pinned by two independent methods (contact-event counting 2.39 Hz vs FFT 2.33 Hz, agreeing): reported numbers are cycle frequency = steps per leg per second.
Change
Gate retracted in writing: "步频 ≤1.5 Hz 应删除或放宽到 ≤2.5 Hz"; frequency definition standardized before entering any config.
Outcome
walk_v4/v6's 2.1-2.4 Hz reclassified from disease to normal; the v3 failure got its probable explanation (suppressing stepping to meet a wrong gate).
Mechanism
A gate is only as good as the measurement and the reading behind it; thresholds inherited from a misread plot become invisible design constraints that later training obeys at real cost. Cross-checking a threshold against first-principles dynamics (pendulum frequency) is a cheap way to catch such misreads.
Applies when
- an acceptance threshold repeatedly fails policies that look healthy
- thresholds were set from a single person's reading of raw data
- a forced compliance with a gate degrades the behavior it guards
“我当初依据自己测的执行器频响定的,但看错了区间。… 2.0~2.4 Hz 的幅值比 0.83~0.88 是可接受的;真正不行的是 walk_v2 的 3.8 Hz。… 腿按复摆算(L≈0.30 m)自然频率约 1.1 Hz … 把它压到 1.5 Hz 是跟摆动动力学对着干(walk_v3 把迈步压没了,可能正是这个原因)。”
train/WALK_DIAGNOSIS.md § ② 撤回"步频 ≤1.5 Hz"这条验收标准 —— 是我定错了 Two machines, one configuration - every gain, offset and torque limit changes only in robot.yaml, whoever edits pushes at once, both checkouts show the same commit before the robot moves, and a pulled policy file is size-checked and synced before power-off
two-machine-config-disciplineTreat the robot's configuration as a versioned artifact with one source of truth, push every change immediately, verify identical commits on every machine before a hardware session, keep hardware limits in the repo and the firmware in sync in both directions, and verify transferred model files (size, digest) before running them.
Symptom
The robot's onboard computer runs the bridge and deploy scripts from its own checkout while training and analysis happen on other machines; a hot fix left on one side, or a half-written file, silently makes the robot run something other than what was evaluated.
Context
The operator runbook's "wall version" of the two-machine discipline: configuration changes only in robot.yaml (calibration offset/sign, gains, torque limits), committed and pushed from the Mac, pulled on the robot, bridge restarted; code is not edited on the robot, and if it is, it is committed and pushed on the spot - nothing unpushed overnight; 30 seconds before every real-robot session both checkouts must show a clean status and the same last commit hash; changing tau_max requires writing the motor's limit_torque too (and the reverse); re-zeroed motors require re-measuring offsets. The recovery line added: after pulling on the robot, check the ONNX is not zero bytes (a lesson from a corruption incident on 08-12) and sync before cutting power; the recovery and main lines are separate worktrees, each pulled with --ff-only.
Change
Operating rules, pinned on the wall and repeated in the hanging checklists ("git pull, both machines on the same commit").
Outcome
The sources record the rules and the incident that produced the size check; they do not record a count of sessions the rules caught.
Mechanism
A policy is evaluated against one configuration; any divergence between the machines, or a truncated file, turns a hardware result into a result about an unknown configuration.
Applies when
- a robot's onboard computer and a workstation both hold the configuration
- someone hot-fixes code or gains on the robot
- model files are copied or pulled to the robot before a session
“改配置只改 robot.yaml(标定 offset/sign、增益、限扭全在里面)→ Mac git commit + push → NX git pull → 重启桥。 … 谁改完谁立刻推,永远不留未推送的改动过夜。 … 每次上真机前 30 秒检查:两边 git status 干净、git log -1 哈希一致。 … 铁律不变:改 tau_max 必须同步写电机 limit_torque(反之亦然);电机重新标零后 offset 必须重测回填 yaml。”
RL系统/FOLLOW THIS copy 2.md § ② 双机维护纪律(贴墙版) The deploy-side walk/recovery switch - into recovery at tilt > 65 deg held 0.3 s, back at tilt < 15 deg with angular rate < 1 rad/s and straight knees held 1 s, a 15 s timeout, last action cleared both ways - and the handoff steps the design required are only partly implemented
walk-recovery-fsm-handoffSpecify a deploy-time controller switch as hysteretic, time-filtered predicates the robot can measure (proxy what it cannot, e.g. straight knees for height), a timeout that ends in a safe stop, and a complete handoff (history, clock, last action, command ramp) - then test that the code performs every handoff step, because the design document is not the implementation.
Symptom
With a recovery policy and a locomotion policy as separate networks, the robot needs a switch: when is it "fallen", when is it "up", and what state must be reset so the next policy does not act on the previous one's history.
Context
The 08-09 design: enter recovery when fallen (tilt > 55 deg or height < 0.60 x 0.384 m) for 150 ms, leave for a stand-hold when upright (tilt < 12 deg, height > 0.85 x 0.384 m, feet steady, |omega| < 0.8) for 400 ms - wide entry, strict exit, hysteresis - then a mandatory handoff trio before walking resumes (reset the walking policy's observation history, restart its phase clock at 0, clear its previous action and latency buffer) and a command ramp instead of a jump. The 08-14 implementation in deploy_policy (--recovery-policy): each policy under its own manifest contract (walk: nominal + scale; recovery: beta-anchored), the same rl_default gains, the tilt cutoff disabled; RECOVERY at tilt > 65 deg for 0.3 s; LOCO again at tilt < 15 deg and |omega| < 1 and knees straight (< 0.35 rad) for 1 s - the robot's computer has no height estimate, and straight knees stand in for height so a V3.0-style upright kneel cannot pass as standing; RECOVERY longer than 15 s ends in a safe stop; last_action cleared on both switches; command forced to 0 during RECOVERY; power scaling applies to LOCO only.
Change
An open account was written down with it: the recovery end state (0.355 m stance, hip yaw -/+27 deg) is outside the walking policy's training start distribution, so the first test must use stand / zero command as LOCO, and the long-term fix is to widen the walking policy's initial states rather than bend recovery's stance to suit walking.
Outcome
The spec records only a successful compile (py_compile); the hang test was left to be done on site. The operator runbook carries the three-step procedure (hang with stand as LOCO, mat and push, then omni walk as LOCO) but no outcome.
Mechanism
Two policies trained separately each assume their own history, clock and last action; a switch that carries any of them across feeds the next policy a state it never saw - the design called this "the walking policy seeing a ghost history".
Conflicts
The 08-09 design requires resetting history, clock and previous action plus a command ramp; the 08-14 implementation records last_action clearing and a zero command during recovery; the one-leg spec of 2026-09-14 lists the handoff hygiene as specified but not implemented - reset_history() is called by nothing (a 215-dim policy would carry four frames of pre-fall history), the phase clock is not zeroed, and there is no command ramp back to LOCO. No hardware run of the FSM is recorded in either source.
Applies when
- switching between separately trained policies on hardware
- a policy with history or phase observations is re-enabled mid-run
- the robot lacks a sensor the switching criterion was designed around
“**判据**:进 RECOVERY = 倾角 >65°(`--fall-tilt-deg`)持续 0.3 s;回 LOCO = §5 真机可测子集:倾角 <15° ∧ |ω|<1 ∧ **膝直 <0.35 rad(NX 无高度观测, 高度门用膝直代理 —— 防 V3.0 型"跪坐但直立"误判)** 持续 1 s (`--recover-hold`);RECOVERY 单次 >15 s(`--recovery-max-time`)安全停。 … **切换卫生**:两向切换 last_action 清零;RECOVERY 态 cmd 强制 0”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §50 FSM 双策略调度(2026-08-14,用户令):deploy_policy --recovery-policy Compare the achieved reward to the computed ignore-floor to tell "never learned" from "learned but unprofitable"
ignore-floor-diagnosisFor any skill that trains flat, compute the reward the null policy would earn on that term; achieved==floor means the behavior never paid out (find why: exploration, reward observability, or feasibility) - do not tune weights first.
Symptom
C4 sidewalk failed on both arms; the question was whether the policy had found sidewalk and rejected it as unprofitable, or never found it at all - two diagnoses with opposite fixes.
Context
The theoretical value of the sidewalk tracking term for a policy that completely ignores the command was computable from the command distribution: 0.189. Trained final values landed at 0.187 (arm A) and 0.204 (arm B) - sitting exactly on the ignore-floor - while the honest balance at the optimum actually favored sidewalking (net +0.88/step inside the side bucket). Later the same arithmetic closed the whole saga: for the true reward landscape, doing real sidewalk scored 0.153 vs 0.944 for ignoring - the policy's refusal "是理性最优,不是探索失败" (rational optimum, not exploration failure) under one hypothesis, and under the final measurement-corrected account the policy had "每一次都在 做理性选择" (made the rational choice every time).
Change
Diagnostic rule adopted: compute the ignore-floor for the new term; if the achieved value sits on it, the behavior was never expressed in useful volume (or the reward cannot distinguish it - check both); if the achieved value is above floor but the behavior is absent at deployment, the policy sampled it and priced it out - then the reward balance, not exploration, is the lever.
Outcome
Correctly identified that PPO had not merely under-valued sidewalk; each subsequent hypothesis (waveform sign, regularization cage, exploration form, reward kernel) was tested against this floor arithmetic, which kept the search honest through three reversals.
Mechanism
Every reward term has a computable value under the null behavior; the achieved-vs-floor gap is a one-number audit of whether the optimizer ever monetized the target behavior. It converts "training failed" into one of two mechanistically distinct states with different fixes.
Applies when
- a new skill's tracking reward plateaus early
- deciding between exploration fixes and reward-weight fixes
- post-mortem of a failed curriculum rung
“track_lin_vel_y_exp 训练终值恰好坐在「完全无视指令」的底分上(A 0.187 / B 0.204,理论值 0.189),而终点 balance 明明有利(side 桶内净 +0.88/步)。不是学会了不划算,是根本没学到。”
train/C_LADDER_RUN.md § 3e. 为什么首战 FAIL Knee swing collapsed because it directly trades against the slip penalty - price the conflict explicitly and clamp what reward cannot hold
knee-swing-vs-slip-pricingWhen a behavior collapses as another metric improves, look for the term pair trading them and set their price ratio deliberately (with escalation and reverse tripwires pre-registered); where the policy actively spends action budget to undo your target, stop paying more reward and clamp the target space structurally.
Symptom
Knee peak-to-peak swing collapsed across generations - v5 33 deg, v10 26-30, v10b 7-8, v11 6.5-8.6 - and rolling the clock back did not recover it, acquitting the clock; the collapse tracked the gated slip penalty instead: "屈膝与不打滑在当前奖励里直接对抗" - v10b's excellent 93 deg slip was purchased with knee amplitude.
Context
Reward-side flexion fixes had failed three times: raising reference amplitude backfired twice (v9/v11), and v11's deep-squat default was actively fought by the policy - it spent 0.68 of action budget pulling the squat straight ("被策略花 0.68 动作拉直反杀"). v12's design accepted the conflict as real and attacked on two tracks: (1) ECONOMICS - a direct knee_swing_amplitude reward (+0.3, target 0.55 rad, capped at 0.6/step = 55% of tracking), explicitly opposed to the slip penalty by design ("显式对立——这正是设计:v12 就是这场对抗的定价实验"), with an escalation ladder (K +0.3 -> +0.5, then slip -0.5 -> -0.3, one layer at a time) and a reverse tripwire (slip telemetry back at v10 levels -> slip weight to -0.8, accept ~20 deg knee compromise); (2) STRUCTURE - knee target bounds [0.2, 0.9] rad so full straightening is physically impossible (straightest 11.5 deg) and the 0.68 fighting budget is released. A bonus falsifiable prediction was attached: phase-lock strength tracks amplitude (v9_probe 48 deg locked 2.5 Hz; v11 low-amplitude 1.36 Hz unlocked), so if K works, hardware phase-lock should return - one change, two verdicts.
Change
knee_swing_amplitude reward + knee target clamp + pre-registered escalation/reverse levers; the failed reward-side-only approach retired.
Outcome
The lineage was frozen before v12 trained (strategic reset), but the diagnosis stands as the walk line's clearest example of two reward terms trading a behavior between them, with the pricing experiment and structural clamp fully designed and calibrated.
Mechanism
When two terms price opposite aspects of one motion (swing amplitude creates yaw momentum that becomes slip), the optimizer settles wherever the price ratio puts it - patching one side moves the equilibrium, not the conflict. Explicit pricing makes the trade a designed quantity; structural clamps remove the regions where the policy spends budget fighting the designer.
Conflicts
The pricing experiment (K vs slip) was designed and calibrated but never trained - the 2026-08-05 reset suspended v12; the collapse attribution table and the 0.68-action counterattack are measured, the remedy's效果 is untested.
Applies when
- one gait quality degrades in lockstep with another's improvement
- the policy visibly fights a default pose or reference
- repeated reward-side fixes for the same behavior have failed
“膝摆塌在 v10→v10b,头号嫌疑是门控滑移罚(四代实测膝 p2p:v5 33° / v10 26~30° / v10b 7~8° / v11 6.5~8.6°;退时钟没救回 → 非时钟)——"屈膝"与"不打滑"在当前奖励里直接对抗 … 奖励侧修屈膝已三败 … v11 深蹲 default 被策略花 0.68 动作拉直反杀”
train/WALK_V12_SPEC.md § 0. 定位 / 2. K —— 膝摆经济(与滑移罚的对偶) Write each config's expected hardware signature before the session - and if reality disagrees, change the books, not the conclusion
preregistered-real-expectationsBefore hardware runs, write per-config expected signatures and the disagreement rule (hardware outranks sim; discrepancies get recorded, not reconciled); validate the harness by checking it reproduces at least one known real behavior.
Symptom
Hardware impressions are easily narrated after the fact; without written expectations, any real-robot outcome can be made to "match" the sim story.
Context
The S2 acceptance sheet carried a section titled "sim 侧预注册预期 (事后核对, 不许事后改)" - per-configuration behavioral signatures written before the session: s1e@0.8 the disturbance king (push 159/160, zero chirality, all-mu 20/20) at the cost of speed gates 0/20 and zero-command wander ~0.98 m with -29.5 deg/20 s rotation; fric-3000 "walks accurately but is easier to push over"; fric-2400 neither. Credibility check included: the sim harness had reproduced the already-recorded real behavior (pace in place + right drift + net rotation -30 deg/20 s), which "提高本单全部预期的可信度". The anomaly clause fixed the epistemics in advance: if results systematically disagree with sim, "不改结论改账" - don't massage the conclusion, write the discrepancy into the books, and per the earlier zero-warning lesson, hardware wins.
Change
Every hardware session ships with a pre-registered expectation table (signature per config), a baseline-match credibility check, and a written precedence rule for disagreement.
Outcome
The A/B session became falsifiable: agreement confirms the proxy, disagreement is booked as a proxy-bias finding rather than argued away.
Mechanism
Pre-registration converts qualitative hardware sessions into tests of the sim-to-real mapping itself; a reproduced known behavior calibrates trust in the remaining predictions; and fixing "who wins on disagreement" beforehand prevents authority from drifting to whichever source flatters the plan.
Applies when
- planning any hardware acceptance or A/B session
- the sim harness's credibility in this regime is unestablished
- post-session write-ups tempt narrative fitting
“⚠️ sim 复现了真机已记录的「原地踏步 + 右漂 + 净旋 −30°/20s」—— harness 与真机行为对得上, 提高本单全部预期的可信度。… 结果与 sim 系统性不符 → 不改结论改账: 写进 README 该节, 按 「Isaac 指标三次零预警」的教训, 以真机为准。”
train/REAL_RUN_S2.md § 2. sim 侧预注册预期 (事后核对, 不许事后改) / 4. 异常处置 The standing-pose reward had been pulling toward the narrow stance the whole line was fighting - a zero-training kinematic audit of the target vector found it, after first auditing the wrong nominal
pose-target-geometric-auditBefore training on a pose target, audit it with forward kinematics - is it geometrically consistent (feet flat, intended stance, intended width) and is it the frame you think it is (action nominal vs standing default)? A posture term's target may itself be the attractor you are fighting.
Symptom
Several rungs aimed at widening the stance failed; the stance stayed narrow as if something kept pulling it back.
Context
The audit (MuJoCo forward kinematics, no training): every 5 deg of hip roll widens the stance ~5.5 cm (0.271 m at 5 deg, 0.383 m at 15 deg); at 47 deg of hip yaw a wide stance cannot be flat-footed (residual foot tilt ~0.7 x hip roll), which explained the stalled rungs. The first report also said the stand_pose nominal (hip roll 25, knee 60) has a 63 deg residual foot tilt - but that was the action frame's nominal (the limit-midpoint squat), which stand_pose never used, despite a docstring warning not to mix them. stand_pose's real target was DEFAULT_JOINT_POS: the contract's all-zero pose, legs parallel, ~0.22 m apart.
Change
The disease statement was corrected in writing: the narrow stance was not an accidental by-product of proxy traps but the target stand_pose had been actively rewarding. A stored "narrow the stance" knife was marked toxic. V2.8 moved the target to a flat 15-deg stance (sigma 3 -> 1.5, flat_feet margin 5 -> 20 deg).
Outcome
V2.8 still failed in-lineage (stance unchanged, feet nearly overlapping, mu 0.4 transfer 2%) - see stance-decided-by-get-up-path - and the from-scratch V3.1 removed both roll joints from stand_pose and put width into a task-space term, which is what finally produced a 0.355 m flat stance. The 63 deg finding was kept as a warning: an action nominal used as a standing target would be a ready-made pit.
Mechanism
A posture term with a sharp kernel around the wrong target is an active attractor; every other term fighting it pays twice.
Applies when
- a posture keeps returning despite penalties against it
- a reward uses a default or nominal pose as its target
- the contract has more than one "nominal" (action frame vs standing pose)
“上文"stand_pose 的 nominal (hip25/knee60) 残倾 63°"**审计错了对象**:那是 **动作参考系 nominal**(限位中点蹲),stand_pose 从未指向它(函数 docstring 原文即警告"两者别混",还是混了 —— 记档)。 … **修正后的病根陈述:窄站距不是代理陷阱的意外副产物,而是 stand_pose 一直在主动奖励的目标本身**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 勘误(实现时抓到):审计混了两个 nominal —— 真病根比误诊的更直白 FK-verify a borrowed reference's structure, then size its amplitude by the reference's job - it pins phase, the policy adds lift
reference-structure-fk-amplitude-divisionWhen borrowing a reference trajectory: verify its structural claim against your own kinematics (an invariant like flat-foot), assign it the phase-pinning job, and size amplitude low enough that the policy contributes the lift - moving toward a proven foreign value in halves, not jumps.
Symptom
walk_v4 had big knee swing (40-46 deg) but only 18-24 mm foot lift - amplitude without hip/knee/ankle phase coordination; later, walk_v5's real-robot swing ballooned to 73.6 deg (sim 55.7) with violent footfalls - amplitude over-driven by the reference.
Context
Structure first: Humanoid-Gym's 1:2:1 hip:knee:ankle reference was verified on the local model before adoption - the ratio exactly satisfies the locally derived flat-foot constraint hip - knee + ankle = 0, FK-tested at multiple amplitudes with sole pitch 0.00 deg throughout. Amplitude second, and here the first reasoning failed honestly: FK said shorter legs need LARGER reference scale (0.30 for 30 mm lift), and the FK was correct - but the premise was wrong ("FK 没错, 但前提错了"): it assumed foot lift must come from the reference. HighTorque Pi, same scale, uses 0.08 with a 0.02 m foot-height target - proof that lift is added by the policy ON TOP of the reference, whose actual job is pinning the phase relationship. Scale 0.30 made the reference the entire gait: over-constrained and over-driven. The correction went to 0.15, deliberately not Pi's 0.08: "一次只走一半, 留退路" (walk half the distance, keep a retreat).
Change
target_joint_pos_scale 0.30 -> 0.15 as one of v6-minimal's three changes, treating both the footfall force and the lateral kicking (yaw momentum scales with leg swing amplitude).
Outcome
v6 improved landing force 1.72x -> 1.55x, suspended tilt 45.9 -> 23.0 deg, turn-gain asymmetry 70% -> 19%; the later v6-halved-shaping experiment (35 mm -> 4 mm collapse) confirmed the reference still carries the gait's existence on this machine - the division of labor is real but machine-specific.
Mechanism
A joint-space reference plays two separable roles: encoding structure (phase relations that keep the foot flat) and injecting amplitude (energy). Structure transfers across robots and is checkable by FK against an invariant; amplitude is a negotiation with the policy, and over-assigning it to the reference removes the policy's freedom to modulate lift with state.
Applies when
- importing a reference gait / imitation target from another codebase
- reference amplitude reasoning based on leg length alone
- real swing amplitude far exceeds sim's under a strong reference
“FK 没错, 但前提错了。我默认抬脚必须由参考轨迹产生。HighTorque Pi 同尺度机器人 … 用 0.08, 而它 target_feet_height = 0.02 m —— 说明抬脚是策略在参考之上加出来的, 参考只负责钉住髋/膝/踝的相位配合。我们取 0.30 等于让参考本身就是整个步态, 过约束 + 过驱动”
train/WALK_V6_MINIMAL.md § ① target_joint_pos_scale 0.30 → 0.15 The real robot's right-leg kicking was over-trained-delay times loop gain - irreducible pipeline latency is plant, model it fully from day one
pipeline-latency-is-plant-not-drMeasure the end-to-end action pipeline delay and build it into the nominal plant and every acceptance gate from day one; treat power/scale deratings that "fix" oscillation as gain-reduction crutches flagging an unmodeled delay, and expect higher-feedback-gain policies to be MORE delay-fragile.
Symptom
On hardware, s1c/s1d at action scale 1.0 always kicked wildly with the right leg (s1c only ran as SOTA at power 0.8; s1d only at 0.7) - while sim showed nothing under default evaluation.
Context
Sim reproduced the incident item by item once the real pipeline delay was injected: s1d@1.0 with --delay 1 fell at 10.2 s, --delay 2 at 5.2 s; s1c@1.0 stressed (r_hip_roll saturation 5 -> 16%; "右脚" = the policy's chirality makes the right leg its high-gain leg); and the combos that worked on hardware (s1c@0.8+delay2, s1d@0.7+delay2) all survived in sim. Mechanism: the real pipeline is ~1-2 ticks (BusWorker next-cycle pickup + CAN round trip) but S1.1 trained only to 1 tick - "超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定" (delay beyond training x full loop gain = oscillation; the power derating had been buying stability by compressing loop gain). s1d was MORE fragile than s1c because its yaw 3-layer stack had learned higher feedback gain - higher gain, lower delay tolerance. Three changes: latency DR widened to cover reality; acceptance gates and smoke runs moved permanently to --delay 2 ("门必须在真机条件下预测 真机"); and the doctrine written twice-paid: "不可约的管线属性(延迟、 限速)不是'随机化选项',是 plant 本体,第一天就该全额建模" - S1's nominal-then-robust staging falsified by hardware for the second time. The later s1e hardware run at power 1.0 (no kicking, normal force) closed the loop: "0.8 = 旧代拐杖" - the derating had been a crutch for the under-modeled delay, not a real requirement.
Change
Latency modeled as plant from day one of any lineage (measured 1-2 ticks covered, bridge-layer rate limits likewise modeled by default); every gate and smoke evaluation issued under --delay 2.
Outcome
Kicking reproduced, explained, and eliminated in the s1e generation at full scale and full power; the deploy-side crutches (0.7/0.8) retired for the new lineage.
Mechanism
Feedback oscillation onset is a product of loop gain and phase lag; a policy trained below the real delay learns gains that sit past the real stability margin, and any output derating masks it by scaling gain down. Since pipeline delay is deterministic hardware property - not an uncertainty - it belongs in the nominal plant, and every evaluation must include it or the gate predicts a robot that does not exist.
Applies when
- hardware oscillation/kicking that sim only reproduces with added delay
- a policy only runs on hardware at reduced power/scale
- defining what belongs in the nominal plant vs the DR list
“真实链路延迟 ~1~2 拍 … S1.1 只训到 1 拍——超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定。s1d 比 s1c 更脆 = yaw 三层栈学出更高反馈增益,增益越高延迟容忍越低。… 教训入账:S1「先标称后鲁棒」第二次被真机证伪——不可约的管线属性(延迟、限速)不是"随机化选项",是 plant 本体,第一天就该全额建模。”
train/OMNI_V0_SPEC.md § 3. S1.4(真机右脚乱踢事故强制) The restart reward table lists a reason for every term AND a lesson for every exclusion - absent terms are removed, not zero-weighted
minimal-reward-table-with-provenanceMaintain the reward table as an evidence ledger: every term cites the episode that justifies it, every excluded term cites the episode that convicted it (including the development stage it is valid at), and retired terms are deleted from the config, never left at weight zero.
Symptom
Seven walk generations had accumulated an entangled reward table where nobody could say which term earned its place; the restart needed a table that could be audited line by line.
Context
The minimal table v2 was built under three written principles: "一项管 一件事、结构性反抬脚的项一个不留、塑形只留一套相位逻辑;不在表里的一律 不加" (one term per job; zero structurally-anti-lift terms; exactly one phase-shaping logic; nothing outside the table gets added). Every row carries its provenance (e.g. base_height target = standing height cites the v4 crouch lesson; split x/y tracking cites the merged-exp gradient hole; world-frame yaw cites the v1/v3 body-frame lesson). Every EXCLUSION carries its same-type precedent: feet_landing_vel out because it is poison while the gait is unformed (drag pays 0, lifting pays - the v4-clearance / v8a-B "reverse threshold" family) though it was fine in v7 when the gait already existed - term validity depends on developmental stage; feet_air_time out with its measured non-lever evidence (v6 had it at 2.0 and still lifted 4 mm); and replaced terms are REMOVED from the config ("置 None,不是权重 0 挂着") so audits see truth, not dormant weight.
Change
Reward table rebuilt as ~20 rows each with weight + provenance column; exclusion list maintained alongside with the falsifying episode for each; dormant terms deleted rather than zeroed.
Outcome
Later revisions (S1.2/S1.3) modified the table by citing and updating specific rows' evidence rather than re-arguing the whole design; the table doubled as the lineage's reward-lesson index.
Mechanism
A reward table is a set of standing hypotheses; attaching each row's evidence makes revisions targeted and reversible, and recording why a term is absent prevents the cycle of re-adding known poisons. Deleting vs zero-weighting matters because config audits and DR interactions see the term either way - a zero-weight term is dormant complexity waiting to be flipped on wrongly.
Applies when
- designing a reward table for a restart or new task
- someone proposes re-adding a previously removed term
- auditing which reward rows still earn their place
“原则:一项管一件事、结构性反抬脚的项一个不留、塑形只留一套相位逻辑;不在表里的一律不加 … feet_landing_vel(评审 #4):拖地时代价恒 0、抬脚才收费——与 v4-clearance/v8a-B 同属「反抬脚门槛」家族,步态未成形时是毒;v7④ 加它时步态已存在。… 已从 cfg 移除(置 None),不是权重 0 挂着。”
train/OMNI_V0_SPEC.md § 3. 最小奖励表 v2 / 明确不带 Anchoring the action on the measured joint angle (target = q + beta*a), with a beta curriculum down to tau_limit/kp, bounded torque by construction, removed the re-falls and later stood the robot up on hardware
beta-anchored-action-targetFor large-motion skills on position-controlled actuators, bound the action relative to the measured joint angle with a per-joint authority of tau_limit/kp, curriculum the authority down from full range, keep the curriculum state out of the observation and pin acceptance at the deployed authority - and make the deployment code refuse to run the anchored contract without a measured q.
Symptom
The V0 full-range absolute action produced violent targets; the V1 command-anchored rate limit made standing oscillate. Both failure modes came from how the action becomes a target.
Context
V2.0 (user approved, from scratch): BetaAnchorJointPositionAction, target = q_measured + beta_j(m)*a, memoryless per step. beta_j(m) = floor + m*(beta0 - floor), beta0 = the contract half-range (m = 1 reproduces V0 authority), floor = min(tau_limit/kp, beta0): hip_pitch 1.309 -> 0.40, knee 1.047 -> 0.40, hip_yaw -> 0.917, the other joints unchanged - the tightening lands exactly on the joints the V0 torque account convicted. m drops 0.1 per step when a standing-share EMA exceeds 0.35. beta is NOT in the observation, so the 45-dim contract is untouched; acceptance is pinned at m = 0 because the Python curriculum state is not saved in the checkpoint. The deployment chain got a new profile (recovery_v2: action_anchor current_q, explicit per-joint beta written into the contract, independent of the gain profile), and policy_io raises if q is missing rather than silently falling back to the absolute contract; the old profile's check reproduced its pre-change deviation bit for bit.
Change
New action term and beta curriculum; later the RS06 floor was lowered 0.40 -> 0.30 -> 0.25 (kp*beta 7.5 N*m) and the stamped deployment profile was synced to 0.25.
Outcome
First acceptance at m = 0 (v2_0b): re-falls 0% in every category, the torque gate passed for the first time on the line (worst 69.9%), knee jitter 0.004; supine 98.8 / side 88.8% with prone and mid still failing (fixed by the conditional pull curriculum). MuJoCo showed demand at or under the limits (hip_pitch 11.7/12 against V0's 26.8). Lowering beta cut impact (hip_pitch demand 9.7 -> 8.5 N*m) but barely slowed the get-up - it had become coordination-limited. Enabling the policy moves the target only +/-beta around the current pose, so there is no homing fling; the 08-11 real get-up and the later v3_1p1c both run on this contract.
Mechanism
kp*beta caps the proportional torque in a single step with no build-up delay and no memory, giving both a hard impact bound and full balance bandwidth.
Applies when
- a skill needs full joint range but hardware torque limits are low
- absolute position targets cause impacts or saturation
- changing the action semantics of a contract that deployed policies share
“**动作项** `BetaAnchorJointPositionAction`:`target = q_实测 + β_j(m)·a`, 逐步无记忆 … **Play/验收钉 m=0(= floor = 部署档)**:python 课程状态不进 checkpoint, Play cfg 显式 `beta_m_start=0` … 判读:**结构赌注兑现** —— 站姿零再摔 + 力矩账首过(kp·β 封顶按构造)”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §33 V2.0 预注册(2026-08-10,用户点头开工):β 锚定动作空间,从零训 Freeze the deployment contract, stamp every export, and let an automated checker catch wiring bugs
contract-freeze-and-checkerFreeze and fingerprint the policy I/O contract; ship contract changes as new versioned profiles that leave old artifacts bit-identical; and extend the automated contract checker with every pipeline change, forcing the new path to execute in the check.
Symptom
Contract-level changes (observation layout, action pipeline) are where silent sim/real divergence is born; two real wiring bugs appeared the one time the action pipeline was extended.
Context
The 215-dim observation contract was frozen ("纪元 3,三机 digest" - an era number plus a digest agreed across three machines); proposals that would break it (e.g. a GRU memory) were rejected on contract grounds. Every exported ONNX is stamped and verified with a manifest (onnx_manifest --stamp / --verify), and deployment refuses mismatched combinations. When C4 added the lateral feed-forward, it went in as a NEW profile (omni_ff) leaving the existing omni profile's behavior bit-identical; the checker (check_contract) was extended to force the feed-forward path to actually execute (cmd_vy=0.13) and promptly caught two genuine bugs: (1) re-clamping with soft_joint_pos_limits after the feed-forward (0.23 rad deviation) instead of reusing the parent's clip; (2) indexing processed actions by asset.joint_names instead of the action term's own contract-ordered _joint_names, which landed the feed-forward on the wrong joints (l_hip_yaw / r_ankle_pitch).
Change
Contract discipline as implemented: frozen dims + digest; manifest stamping and refusal; contract changes only via new versioned profiles; checker updated in the same commit as any pipeline change, with inputs chosen so new code paths are exercised.
Outcome
Both wiring bugs caught before any training or deployment ("两个都是 check_contract 当场抓出来的 —— 这次它值回票价"); old deployments provably unaffected by the new profile.
Mechanism
The contract is the only interface the policy and robot share; freezing plus fingerprinting makes divergence detectable, and an executable checker turns "the contract holds" from a belief into a test - but only if its inputs actually drive the new code path.
Applies when
- modifying the action or observation pipeline of a deployed policy
- exporting policies for hardware
- proposals that would change observation dims or history structure
“契约校验抓到的两个真错误(记账,别再犯):1. 前馈后误用 soft_joint_pos_limits(URDF 限位 ×0.9)重钳 → 0.23 rad 偏差 … 2. 用 asset.joint_names 索引 _processed_actions → 前馈落到 l_hip_yaw/r_ankle_pitch 上 … 两个都是 check_contract 当场抓出来的 —— 这次它值回票价。”
train/C_LADDER_RUN.md § 3j. 契约级改动 / 契约校验抓到的两个真错误 Calibrate a wall penalty by measuring healthy and sick policies - healthy pays ~0, the disease pays a wall
calibrate-threshold-between-healthy-and-sickCalibrate every threshold penalty by evaluating its exact formula on replays of at least one healthy and one sick policy: place the threshold between their distributions, size the weight so the sick policy pays a decisive fraction of tracking while the healthy one pays ~0, and pre-compute neighboring thresholds for cheap adjustment.
Symptom
Real walk_v7 occasionally clipped its own legs (stance narrowed to 133 mm mean vs nominal 214.5); a foot-distance penalty was needed, but an uncalibrated threshold/weight risked either doing nothing or becoming a reverse barrier.
Context
The term (relu(d_min - lateral foot distance), measured in the base yaw frame because world-frame y is meaningless after turning) was calibrated before training by replaying three known policies through the exact reward formula at the acceptance operating point: healthy v5 (183 mm) pays 0.6% of tracking - effectively free; narrowed v7 (133 mm) pays 22% - effective widening pressure; collapsed v8 (93 mm) pays 55% - a wall. d_min 0.16 was placed deliberately between healthy and sick, with alternative thresholds (0.14/0.18) pre-computed in the tool output for later adjustment. The shape self-check was named as a standing question: "先问'零代价的选项是什么'" - the zero-cost region must be exactly the desired behavior.
Change
feet_lateral_distance added at d_min 0.16 / weight -10, with telemetry expectation pre-registered (should decay toward 0 as stance learns >160 mm; if bow-legged over-widening >214 appears, only then discuss an upper bound).
Outcome
v9_probe on hardware: no leg contact ("没碰腿(N2 兑现)"), stance min 145/126 mm green; the term's zero-cost design left healthy gait untaxed.
Mechanism
A relu threshold penalty defines a free region and a priced region; its correctness is entirely in where the boundary sits relative to the healthy and pathological distributions. Replaying known-good and known-bad policies through the exact formula measures both distributions in the term's own currency, making the threshold and weight a placement decision instead of a guess.
Applies when
- adding any relu/threshold-style wall penalty
- a safety margin (foot distance, joint limit, clearance) needs enforcement without taxing normal behavior
- choosing between candidate thresholds for a new term
“形状自检 (v6 横向组/v8-B 的教训 —— 先问"零代价的选项是什么"): 标称站距付 0, 健康步态付 ~0, 收窄才付费 … v5(健康) 183 mm … 0.6% ≈ 免费 | v7(收窄) 133 mm … 22% —— 有效推宽 | v8(塌陷) 93 mm … 55% —— 墙 … d_min=0.16 恰在 v5(183)与 v7(133)之间”
train/WALK_V9_SPEC.md § 2. N2 —— 脚距惩罚(已完成权重预标定) Training-log reward values and fixed-command eval values live on different distributions - comparing them once claimed a 44% improvement that was really 6-10%
same-distribution-reward-comparisonQuote reward-term values only with their distribution attached (command range, DR on/off, environment), and compare across runs only when those match; re-measure in a common environment before claiming any improvement percentage.
Symptom
A v6-era analysis concluded slip had dropped 44% by comparing the training log's Episode_Reward against values calibrated in a fixed-command play environment; a same-condition re-measurement showed the true improvement was 6-10%.
Context
The training log's reward is an expectation over the training command distribution (vx 0.15-0.5, yaw +/-0.6, with pushes and domain randomization); play-environment calibrations are taken at a single fixed command with DR off. Subtracting one from the other compares apples to oranges - the warning was written into the v7 pre-flight: "奖励数值只能在同一指令分布下比较 … 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次)".
Change
Rule adopted: any before/after reward-term comparison must hold the command distribution, DR state, and evaluation environment fixed; training-log values compare only against training-log values of runs with identical command/DR configs.
Outcome
The phantom 44% improvement was retracted; later term-level accounting (e.g. the C4 ignore-floor work) consistently specified its distribution before quoting numbers.
Mechanism
A reward term's expectation depends on the visited-state distribution as much as on the policy; changing the command distribution or DR moves every term's baseline. Cross-distribution differences therefore measure the distributions, not the policy change.
Applies when
- comparing reward telemetry across training runs or vs play evals
- claiming improvement percentages from training logs
- term-level reward accounting for diagnosis
“奖励数值只能在同一指令分布下比较。训练日志的 Episode_Reward 是在训练指令分布上算的(vx 0.15~0.5 / 偏航 ±0.6 / 带推力与域随机化), 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次: 据此以为滑移降了 44%, 同条件对拍只有 6~10%)。”
train/WALK_V7_SPEC.md § 3. 开训自查 ⚠️ Training the final recipe from scratch in one run - every mechanism the lineage had accumulated - produced 0% and a seated robot; the order in which the lineage acquired those mechanisms was part of why it worked
curriculum-history-is-part-of-the-productA recipe that ends a lineage is not a recipe for a from-scratch run: consolidate it as ordered curriculum phases matching how the lineage acquired its mechanisms, check that every curriculum criterion is reachable from the starting policy, and read the run's raw term values, not the total reward, before calling it green.
Symptom
V3.0 trained the lineage's whole final recipe from scratch in one 9,000 iteration run - full beta curriculum, prone-conditioned pull assist, friction DR, the 3 s zero gate on standing income, flat_feet - testing the proposition "the product is defined by its configuration, not by its training history". Training looked all green (reward 26.33, episode length 500, 100% time-outs).
Context
Read in raw units against v2_6c at the same weights, the green was a seated equilibrium: base_height 0.377 vs 0.640, stand_pose 0.205 vs 0.516, flat_feet 0.0000 (zero because it sits outside its height gate, not because the feet were flat). Acceptance: 0.0% in Isaac at the deployed authority, 0.0% on every MuJoCo friction level, and still 0.0% at the training-time authority (100% seated at 0.222 m, upright and still).
Change
The full stdout (316k lines) was read: the beta curriculum's criterion (standing share over 0.35) was met zero times, so beta never left the wide setting and the policy had no experience at the deployed authority; the pull curriculum was stuck on the same criterion. The lineage had escaped the seated basin with immediate income (V2.0-V2.2) and only then added the zero gate to cure rushing (V2.5); from scratch, the zero gate removed the early "stand fast, earn more" gradient needed to escape. Verdict "the curriculum history is part of the product", limited to n = 1. v2_6c stayed the product.
Outcome
V3.1 kept the order as explicit phases: P1 from scratch with immediate income (zero gate off) until the curricula advance, P2 adding the zero gate. P1b/P1c escaped the seated basin and passed; P2 was later judged net negative and dropped (time-gate-vs-wide-stance-retire-the-fix).
Mechanism
Mechanisms that refine a competent policy (time gates, tight authority) can delete the gradient a naive policy needs, and a curriculum whose advancement criterion the naive policy never meets freezes at its first level.
Conflicts
The spec limits the falsification to "this recipe + this curriculum criterion" (n = 1, no seed sweep, no criterion tuning). V3.1's phased run succeeding is consistent with the ordering reading but changed other terms too.
Applies when
- consolidating a long lineage of continuation fixes into one clean recipe
- a from-scratch run with all mechanisms enabled plateaus early
- curriculum state is not logged or never advances
“命题:产物由配置定义,而非训练史定义。 … 训练 log(完整 stdout 316k 行)`[beta_anchor]` 仅初始 1 行,**达标 0 次** … V2.0~V2.2 靠**即时计酬**爬出坐姿盆地(§33),站立巩固后 V2.5 才装归零门 治"过快"(§39)。**课程史是产品的一部分。**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §42 V3.0 判决(2026-08-11):从零单 run 全机制 FAIL 于坐姿盆地 A pull-assist curriculum keyed to a global success share was satisfied by the categories that already worked and withdrew before prone learned anything - conditioning the criterion on prone took it from 2.5% to 98.7%
curriculum-criterion-conditioned-on-lagging-categoryMeasure a curriculum's advancement criterion on the population the scaffold is meant to help; a global success share is met by whatever already works, and the help is withdrawn before the lagging case learns.
Symptom
Under the beta-anchored action, supine and side stood reliably while prone still sat (1.3%). A pull-assist curriculum added to help it was withdrawn completely within ~790 iterations and prone moved only to 2.5% (noise).
Context
The pull assist follows HoST: an upward force on the base, active only when the torso is within 30 deg of vertical, scaled by body weight (HoST's 200 N on G1 = 0.583 BW -> 56 N here, steps of 5.6 N, ten levels to zero); the product must pass with no assist. It had already taught sit-to-stand in V1. In V2.1 its advancement criterion was the standing-time share over all envs (threshold raised to 0.55 because the share was already ~0.53).
Change
V2.2: PullAssistForce with gate_category = "prone" - only envs whose first step classifies them as prone count toward the criterion - and the threshold back at 0.35. A feasibility signal was pre-registered: if prone's share stayed near zero under the full 56 N, return to the roll-over path instead of adding force.
Outcome
The prone-conditioned curriculum withdrew level by level only as prone itself passed: prone 98.7%, and the four-category gate passed for the first time on the line (98.6% overall, re-falls 0%, torque gate PASS).
Mechanism
A pooled success share is filled by the categories that already succeed (supine/side ~53%), so the scaffold is removed on their account before the lagging category has used it.
Applies when
- an assist, guide force or easier setting is withdrawn by a success threshold
- one task category lags while the pooled metric looks healthy
- a curriculum ran to completion without changing the lagging category
“**教训:全局站立占比阈会被存量类别(supine/side ~53%)凑够,拉力在 prone 学会前就撤光了 —— metric 设计失误,不是拉力机制失效**(它在 V1 教会过 坐→站)。 … **V2.2(已启动)**:`PullAssistForce` 加 `gate_category="prone"` —— 达标判据 只统计 prone 类 env”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §34 V2.1 判决(2026-08-10) A policy's gain profile is part of its contract - the one-leg policy needs per-joint gains the default profile lacks, and the manifest refused an evaluation under the default once; the recovery contract's beta was never stamped, a known gap not to repeat
gain-profile-belongs-in-the-stampStamp everything that defines the closed loop a policy was trained in - gains included - into its manifest, and make every consumer refuse a mismatch; a profile field that is not in the stamp is a silent misconfiguration waiting for an operator to forget a flag.
Symptom
A policy trained with hip_roll kp 80 and ankle_roll kp 60 behaves differently, or falls, under the default kp 20/12 profile - and the gain profile is a command-line flag an operator can forget.
Context
The one-leg line added a gain_profile field to the contract so the stamped manifest carries it; the spec's deployment note says the manifest guard blocks rl_default and that it had already bitten once in simulation (an evaluation run without the one-leg profile). The same spec states the general rule - any new profile field must be synced into the manifest builder - and names the counter-example: the recovery line's beta was never put into the manifest. The recovery line itself had decided that its anchored authority is computed from the base rl gains and written into the contract so it cannot drift with the gain flag, and that the older kp x 0.9 profile chosen in the V0 era does not match the beta contract and must not be used.
Change
Gain profile as a contract field checked at load; per-contract gain choices written into the run sheets.
Outcome
Evaluations and hardware runs of the one-leg policy run under rl_oneleg or are refused; the recovery beta gap stayed recorded as known.
Mechanism
A policy is trained against a closed loop whose gains are part of the plant; running it under other gains is an out-of-distribution plant, exactly like a wrong observation scale.
Applies when
- a skill introduces per-joint or skill-specific gains
- deployment gains are chosen by a command-line flag
- adding any new field to a policy profile
“增益档 `--profile rl_oneleg` 必须给 —— manifest 防线会拦 `rl_default`(sim 已咬合一次) … (recovery 的 β 未进 manifest 是已知缺口,不再复制)”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §9b AGX 真机手顺 要点 / §3 契约 Raising a command bucket's share does not strengthen its per-state gradient - it only starves the other modes
bucket-share-is-not-a-gradient-leverWhen a skill is not learning, first prove its per-state signal is nonzero (ignore-floor and probe checks); only rebalance sampling shares to fix genuine sample starvation, and account the regression risk to the diluted modes before doing it.
Symptom
Sidewalk was not learning, and the reflex proposal was to give the side bucket a larger share of sampled commands.
Context
The C4-redo3 rung explicitly kept the 20/40/20/20 bucket (stand/forward/turn/side) with the reasoning written out: PPO computes advantages per state, so bucket proportion does not change the per-state gradient of side states; at 4096 envs x 20% x 24 steps the rollout already contained ~19.7k sidewalk states - sample count was not the bottleneck. And the cost side was already measured: cutting forward from 60% to 40% had made vx+0.30 die at +400 in an earlier run - more cuts would only collapse it sooner.
Change
Bucket proportions held constant across the entire C4 redo series; the actual bottlenecks (metric frame bug, reward variance penalty, exploration form) were pursued instead.
Outcome
Sidewalk was eventually fixed with zero bucket changes (feed-forward delivery, +100 iters); forward/turn skills never suffered starvation-induced regressions during the redo series.
Mechanism
Policy-gradient credit is assigned per visited state; oversampling a mode multiplies its states in the batch but not the informativeness of each, so if the per-state gradient is ~0 (behavior unreachable or reward indifferent), N times zero is still zero - while the displaced modes genuinely lose data and regress.
Applies when
- proposing to oversample a failing task/command mode
- a majority mode regresses after share rebalancing
- budgeting env count vs mode share for a multi-skill policy
“比例不动:PPO 逐状态算优势,桶占比不改变单状态梯度;4096 env × 20% × 24 = 每 rollout 已有 1.97 万个侧走状态,样本数不是瓶颈;而 forward 60%→40% 已实测让 f30 在 +400 处死掉,再加码只会更早塌。”
train/C_LADDER_RUN.md § 3i. 桶 20/40/20/20 不动(比例不动) The first real-robot get-up was "very violent, kicking on the floor, dangerous" - a sim-perfect policy with no reason to be slow, unbounded absolute targets, no domain randomization and a rate limiter that filtered nothing; the task was restated as "safe, slow, transferable"
first-real-get-up-violent-stage-one-policyDo not put a get-up policy on hardware until its action is bounded (hard bound or state-anchored targets), smoothed, randomized and tested at the real pipeline's latency, and say explicitly that the task is "safe, slow and transferable" - a simulation-perfect policy optimizes only "gets up".
Symptom
On 2026-08-09 the user ran a V0-lineage recovery policy on the real robot and stopped it: very violent, kicking on the floor, dangerous. The planned next rung (a heavier torque_headroom) was never started.
Context
The spec had pre-registered that R0/R1 products stay in simulation and that the real-robot precondition was the R3 smoothing rungs plus a bridge-slew check plus a hanging protocol; the robustness (DR) rungs had not run. In simulation the policy passed 100% with a get-up of about a second. Which ONNX, which gain profile and whether a torque/joint log existed were left "to be recorded later" and never were.
Change
The V0 ladder was stopped at its best product (R3.1, sim only) and a re-rooting proposal was put to the user. The spec's four-layer account: style (the reward pays for standing early and nothing pays for slowness - HumanUP's "Stage I" get-up, "fast but unsafe ... infeasible for real-world deployment"); impact (full-range absolute targets with no hard bound, raw |a| up to 4.77, action saturation 100%, a single-step change of 0.306 saturating hip_pitch); transfer (zero DR, friction pinned at 1.0, the learned leg bracing); link (the bridge's RL slew equals vel_limit, 0.2-0.66 rad per step, while the real pipeline has 1-2 steps of time-varying latency and acceptance ran at delay 0).
Outcome
The line was re-rooted twice (training-side rate limit, then the beta-anchored action space) and gained a hang protocol before the next real attempt; on 08-11 a beta-anchored policy produced the line's first real get-up.
Mechanism
A task reward that pays for standing early selects the fastest feasible get-up; with absolute full-range targets every large target jump is a torque impulse bounded only by the clip; zero DR and braced-leg solutions do not transfer; and a limiter set at the velocity limit does nothing at 50 Hz.
Conflicts
The four layers are the spec's reconstruction from simulation probes and the literature; the real run's policy file, gain profile and log were never recorded, so no layer was confirmed against hardware data.
Applies when
- a first hardware trial of a high-effort skill is being scheduled
- sim success is high but the policy saturates actions or torques
- pre-registered hardware preconditions are not all met
“用户真机反馈:**非常猛、地上乱踢、危险**,叫停(R3.3 torque_headroom 加档已选型 weight −0.5→−1.5,未启动)。真机细节(哪个 onnx、什么档、有无 τ/q log)**待补记** … 任务从"能起来"变成 **"安全、慢、可迁移"** … **链路层**:桥层 slew RL 档 = vel_limit(10/20/33 rad/s ≈ 每拍 0.2~0.66 rad), 对 recovery 形同虚设;真机 1~2 拍时变延迟,验收默认 delay 0。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 真机叫停与换根判决(2026-08-09) Foot dragging is an attractor, not a low amplitude - and joint damping is the mode switch, adjustable at deploy time
swing-bistability-damping-switchWhen a quality metric is bimodal, stop treating it as an amplitude to be trained up: map the modes against initial conditions and plant parameters, find the parameter that switches basins, apply it first as a deployment lever, and only then bake it into the training distribution (as a plant-family shift, never as an execution-mapping change).
Symptom
s2e_pd-1400's swing height "median 12.1 mm" hid a perfect bimodal distribution: 20 seeds split into a drag mode (2.6-4.9 mm) and a step mode (19.3-24.0 mm) with NOT ONE seed in between - the median sat in the empty gap, and "swing debt -11 mm" really meant "50% probability of falling into the drag attractor".
Context
Two designed experiments closed the mechanism. Test A (nominal plant, 40 seeds): step 42% / drag 58% / middle 0 - at nominal gains, initial conditions alone pick the mode, both modes 100% survivable. Test B (fixed init, kp x kd grid): kd is the mode SWITCH - at kd 1.3 all surviving cells step (13-22 mm), at kd 0.7 nearly all drag (2.7-4.3), only at kd 1.0 does init get a vote; kp >= 1.2 is dangerous (5/6 falls). Global verification at kd x1.3 (20-seed, delay 2): survival 20/20 at ZERO cost, step share 42 -> 80%, swing median 12.1 -> 18.4 mm, slip record low 334, thicker tilt margin - costs: vx 85 -> 78%, saturation +5 pp. A Pareto sweep then priced the knob: step share 42/72/75/88/82/90 across kd 1.00-1.30 with a linear vx tax of -2.3 pp per 0.1 kd - the basin gain is fully collected at kd 1.20 ("1.30 是 over-damping 纯多付税"). Mechanism: low damping leaves a landing micro-oscillation / ground-slide channel the policy can exploit to drag; damping plugs the channel.
Change
Deployment lever adopted: kd-scale 1.20 (conservative 1.15) as the legitimate successor to the power-0.8 crutch ("前者削幅度保稳,后者堵 拖地通道换步态,且不牺牲存活"); training-side prescription: move the DR band to nominal-1.2 x (0.9,1.1) = [1.08,1.32], deleting the [0.7,1.0) drag-teaching zone - a contract-level change requiring digest re-baselining, gated on measuring the real robot's actual kd dispersion first.
Outcome
The kd surgery rung (s2e_kd) delivered basin 8 -> 11/20, slip 405 -> 331, vx 81 -> 85% with no out-of-band fragility (below-band check 20/20) - "拐杖烧进分布的正确姿势", explicitly contrasted with the failed s1g amplitude version: this one changes the plant family the policy has seen, that one changed the execution mapping the policy would have to relearn.
Mechanism
The gait's swing behavior is a bistable dynamical system whose basin boundaries are set by plant parameters; a policy trained across a kd band that includes the drag basin has learned to inhabit it. Shifting the deployed (and then trained) damping moves the system into the step basin without touching the policy - a plant-side fix for what looked like a training deficiency.
Applies when
- a gait quality metric splits into distinct modes across seeds
- deciding between more training and a gain/damping change
- converting a deployment crutch into a training-distribution change
“20-seed 里拖地模式 2.6~4.9mm 与迈步模式 19.3~24.0mm 各半,中间一个不落 … kd 是模式开关——kd1.3 下 6/6 存活格全迈步 … kd0.7 下几乎全拖地 … swing 债的解(至少大半)在部署端阻尼档,不在训练端 … 机理:低阻尼下落脚微振荡/贴地滑给了策略顺势拖行的通道,加阻尼堵之。”
train/README.md § swing 双稳态定性 + kd 部署杠杆 (2026-08-07, 用户设计 Test A/B) The get-up kept getting faster because standing earlier paid more every step - lowering torque authority barely slowed it, and only zeroing the standing income for the first 3 s moved the pace into the design band
per-step-income-drives-speed-time-gateWhen a skill is too fast, find the term that pays for finishing early and gate that income by time; keep the "get into position" term ungated so the policy does not learn to wait, use a ramp instead of a cliff, and confirm with a paired same-level experiment that the drift is motivational before changing it.
Symptom
The user judged the get-up too fast (Isaac medians about 0.7-1.6 s) and suspected path dependence: the policy seemed to get faster the longer it trained.
Context
Lowering the beta authority 0.40 -> 0.30 cut impact but moved supine only 1.70 -> 2.00 s: coordination-limited, not torque-limited. A paired experiment inside one beta level (checkpoint 15,600 vs 18,499, +2,900 iterations, same ruler) measured the drift: get-up medians -7 to -10%. The spec concluded the motive, not the path, was the cause - per-step standing income pays for every early step, and any lineage (even one from scratch) races toward the fastest solution inside its constraints.
Change
V2.5: the standing income (base_height, stand_pose, still, feet_on_ground) multiplied by w(t) = clamp(t/3 s, 0, 1); upright deliberately NOT gated, so righting and sitting up early still pay and the policy is not taught to lie flat and wait; a ramp, not a step. V2.5b: zero before t0 = 3 s, then a 1 s ramp.
Outcome
V2.5: Isaac 100%, get-up +17-43% slower, MuJoCo 100/100/98/100% (the best cross-simulator reading yet), still short of the 3.5-4.5 s design band - a linear ramp only discounts early income. V2.5b: MuJoCo supine 2.04 -> 4.10 s and prone 3.18 -> 4.04 s, inside the band; the Isaac pace barely moved (a lineage habit on a gradient-free plateau). Later the zero gate proved harmful when trained from scratch (curriculum-history-is-part-of-the-product) and in the V3.1 lineage (time-gate-vs-wide-stance-retire-the-fix).
Mechanism
Constraints on authority or velocity change how the fastest solution looks; the time structure of the task income decides how fast the fastest solution is.
Applies when
- a policy is faster or more aggressive than wanted and constraints do not slow it
- progress-style rewards pay every step spent at the goal
- performance drifts faster with more training at fixed settings
“**V2.5 机制(唯一)**:站立收入(base_height/stand_pose/still/feet_on_ground) 乘时间斜坡 w(t)=clamp(t/T_gate,0,1),T_gate=3.0 s;**upright 刻意不门控** (翻正/坐直早期照常拿钱,防"躺平等门开" … **用户假设量化 证实:逐步计酬动机在 β 包络内持续压缩时间,约束挡不住动机 —— V2.5 动机层 修法为正解。**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §39 V2.5 预注册 / §39 补 配对实验 Fix the task first, harden the plant second - DR budget spent on a dying task is wasted
task-shaping-before-plant-hardeningFreeze the task/command distribution before spending DR budget on plant robustness; if the task will still change, schedule plant hardening as a final pass and book the interim robustness gap explicitly.
Symptom
Tempting default ordering was to keep the plant-hardened (S2) lineage and teach it new commands; but the S2 plant adaptation had been earned on the straight-walk task, and the new omni tasks (sidewalk, in-place turn) use completely different contact patterns.
Context
The team had direct evidence that DR robustness is a budget that gets reallocated when the data distribution changes ("push/μ 两轮已实证 DR 预算有限且会被重分配") - robustness trained under one task/command distribution does not persist when training continues under another.
Change
Ladder order set to: first C (task shaping - add command modes until the task family is final), then a second S2 pass (plant hardening) on the C product. The plant-robustness gap this creates mid-ladder is accepted and booked explicitly ("此处不欠账" - the debt is assigned to the second S2 pass, not denied).
Outcome
The first S2 pass was not wasted: its laws (kd bandwidth <-> low mu, push need not be trained, ground mu need not be trained, bistability) let the second pass drop from five rungs to three. The C ladder itself ran on the softer plant band without incident.
Mechanism
DR robustness is carried by the policy's visited-state distribution; changing the task changes that distribution, so robustness bought under the old task partially dissolves. Hardening before the task is final means paying for robustness on states that will no longer be visited - "给一个即将不存在的任务花预算" (spending budget on a soon-to-not-exist task).
Applies when
- deciding ordering between skill/command expansion and DR hardening
- a hardened lineage is proposed as the root for a task change
- robustness regressions appear after adding new command modes
“S2 的 plant 适应是为直行步态调的,C4 侧走/C3 原地转是完全不同的接触模式,先硬化再改任务 = 给一个即将不存在的任务花预算(push/μ 两轮已实证 DR 预算有限且会被重分配)。故顺序改为 先 C(任务定型)→ 再 S2(plant 硬化)。”
train/C_LADDER_RUN.md § 0. 决策逻辑 = 短板可不可恢复 (末段) Audit rewards by realized contribution (weight x achieved value) - a weight of 2.0 was really paying 0.04
realized-contribution-auditEvaluate a reward table by each term's realized per-step contribution under the current policy, never by its weight column; if a term's realized value is ~0, escalating its weight is a no-op - change the term's structure instead.
Symptom
Foot dragging persisted through repeated weight escalation: feet_air_time had been raised 0.25 -> 1.0 -> 2.0 across versions with no behavioral change, and the training-side comment even recorded the fact ("几乎没有 单支撑相, 是拖着脚蹭") without the fix changing form.
Context
Computing realized per-term contributions in the trained state exposed the economy: feet_air_time contributed weight 2.0 x achieved 0.019 = 0.038 per step, against tracking's +1.20 - lifting the leg earned 3% of what tracking earned, so dragging was the rational optimum no matter the weight escalation. The same table acquitted the energy penalties (total -0.30 negative vs +1.74 positive) that a naive read of weights (-5.0 orientation!) would have blamed.
Change
Fix redirected from "raise the weight again" to "add a term whose realized contribution changes the optimum": a clearance penalty sized so its realized magnitude (~0.018/foot when dragging) is comparable to feet_air_time's, enough to flip the optimum without drowning tracking.
Outcome
With the term economy corrected (plus posture/range fixes), swing height reached 34 mm and tracking 87% by v6; weight escalation of the old term was abandoned.
Mechanism
A reward weight is only a multiplier on whatever the policy currently achieves on that term; when the achieved value is near zero (behavior absent), escalating the weight multiplies near-zero. Optimizer behavior is governed by realized per-step magnitudes, so audits must be conducted in that currency.
Applies when
- a behavior persists despite repeated weight increases
- auditing whether penalties are "too strong" or rewards "too weak"
- sizing a new reward term against existing ones
“把 feet_air_time 权重从 0.25 → 1.0 → 2.0 一路加,但没有加高度项。量级算下来:feet_air_time 权重 2.0 × 实得 0.019 = 0.038,而跟踪奖励是 1.2。抬腿的边际收益只有跟踪的 3%,拖地当然是最优解。… 正项 +1.74,负项 −0.30。能量惩罚不是瓶颈,抬腿没收益才是。”
train/WALK_DIAGNOSIS.md § 决定性证据(#1) / 各项奖励的实际量级 Tightening the bridge's rate limiter under an unchanged policy cut torque peaks 30-50% and made other things worse - the policy cannot see the limiter, keeps commanding and winds up; a deploy-side limiter is a safety net, not a cure
deploy-rate-limiter-windupA rate or torque limiter added at deployment lowers peaks but the policy still commands as if unconstrained (saturation, windup, new contacts); use it as a safety net mirrored in evaluation, and put the constraint where the policy can learn around it.
Symptom
After the violent first real-robot get-up, the cheapest candidate fix was to tighten the bridge's slew (rate) limit for the recovery policy without retraining.
Context
Probe on R3.1 in MuJoCo (5 categories x 3 seeds, mu 1.0), monkeypatching the limiter with no repository change: TIGHT = RS06 4.0 / RS02 3.0 / RS00 2.0 rad/s (about 0.08/0.06/0.04 rad per policy step) against the current vel_limit setting.
Change
The probe decided the role of the limiter rather than a deployment.
Outcome
Success 14/15 -> 12/15; get-up median 2.35 -> 3.53 s (max 9.30); torque demand peak median hip_pitch 164% -> 111%, knee 166% -> 86%; action saturation still 100%; leg-leg contact 558 -> 860 frames. The limiter was kept only as a real-robot safety net (mirrored into sim2sim evaluation); the cure moved into training - where the next lesson was that a limiter anchored on the last command is itself an integrator (slew-anchor-is-an-integrator).
Mechanism
A policy that never trained with the limiter keeps issuing the targets it learned; the limiter clips them, the target window runs ahead (windup), and the robot follows a trajectory the policy never evaluated.
Applies when
- a trained policy is too violent on hardware and a quick deploy-side fix is tempting
- adding slew, torque or velocity limits in a bridge or firmware
- evaluation and deployment use different limiter settings
“判读:**链路侧收紧立等可取地把 τ 峰值砍 30~50%,但成功率掉、饱和率仍 100%、 腿-腿接触反升** —— 策略感知不到限速器,目标窗口继续狂奔。⇒ 收紧 slew 只配当 **真机侧安全网**(必须同步进 sim2sim 口径,基础设施现成),**不配当治法; 治法必须进训练**。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 探针:收紧桥层 slew,r3_1 不重训直接测 A walking policy's tilt cutoff is a legal state for a recovery policy - the default 45 deg fall guard had to be raised for recovery tests and is disabled once the switch owns falls, so the abort chain becomes the recovery timeout, the operator's cut, and the firmware torque limits
fall-guard-becomes-a-stateWhen a new skill makes a safety cutoff's trigger a legal state, replace the cutoff with a bound of the skill's own (a timeout ending in a safe stop) instead of just switching it off; keep the operator's cut and the firmware limits as independent layers, and write every flag change into the run sheet.
Symptom
deploy_policy's default protection stops the robot beyond 45 deg of tilt. A recovery policy starts lying at roughly 90-97 deg, so under the default it is refused on the spot - a flag the first hanging checklist forgot.
Context
The layers in the sources: deploy_policy's tilt cutoff (default 45 deg, a line in the safety chain); for standalone recovery tests the cutoff was raised (110 deg in the spec's A/B sheet; 181 deg, effectively off, in some runbook commands); with --recovery-policy the cutoff is disabled because a fall is now a state, not an exception, and RECOVERY lasting over 15 s ends in a safe stop (the runbook calls it the line where the spotter steps in). Independent of the policy: firmware torque limits checked at start (12/17/11 N*m, set_torque --check), the operator cutting enable at any kicking or oscillation, and in the one-leg teleop a space-bar stop that puts the foot down.
Change
The flag was added to the run sheets, and the FSM replaced the removed cutoff with its own bound (the timeout).
Outcome
The spec records the flag omission and its fix; it does not record the FSM's timeout being exercised on hardware.
Mechanism
A safety cutoff encodes one policy's notion of "abnormal"; a new skill whose normal operation lies beyond it either cannot run or runs with the cutoff off, and only a replacement bound keeps the chain closed.
Applies when
- deploying recovery, fall-damage or acrobatic skills behind existing safety checks
- a run sheet disables a protection flag
- listing the abort chain for a hardware session
“--max-tilt-deg(默认 45°,安全链第 13 行写的那个)。recovery 的合法状态覆盖整个倾角域,把它抬到 181 = 实效关闭 … RECOVERY 超时 15s 会自动安全停(看护介入线)”
RL系统/FOLLOW THIS copy 2.md § FSM 吊挂首测 ② 落地测 / #### Recovery Policy (operator runbook, undated) The best checkpoint to SHIP is not the best checkpoint to CONTINUE FROM - maturity is capital against adaptation shock
root-maturity-vs-product-qualityDecide shipping points and fork roots separately: gates rank products, but a root candidate must prove itself by surviving a continuation under the next rung's shift (dual-arm if in doubt) - and prefer the more-trained point as root when product metrics conflict with maturity.
Symptom
A band re-audit found s1e-300 beat the incumbent root s1e-500 on nearly every quality gate (stepping 19/20 vs 13/20 with historically-best 26.9 mm swing, speed gate 14/20 vs 2/20, heading 26 vs 54 deg/20 s) - suggesting the root had been mis-picked and the younger point should take over.
Context
The dual-arm control settled it the other way: continuing the S2 PD rung from s1e-500 adapted smoothly (3/3 smoke throughout), while the b300 control arm (same config, from s1e-300) fell into a survival valley under the PD shock (+100 iters: 1/3 -> 0/3), never climbed out within budget, and its 800-iter product scored 13/20 survival - eliminated. Verdict: "幼年点自身指标再好也扛不住新 DR 适应冲击, 成熟度是本钱,s1e-500 根被数据背书" - a young point's own metrics, however good, do not survive new-DR adaptation shock; maturity is capital. The audit still yielded value: the band scan (200-1000, per-100) mapped the lineage's arc (200 dragging -> 300 peak -> 400+ decay -> 900+ drift blowout), and 300 remains the better PRODUCT answer for shipping-as-is questions.
Change
Selection doctrine split into two questions with different answers: best-product point (quality gates at the point itself) vs best-root point (survives adaptation shocks; more training age = more capital), each decided by its own evidence - and root claims settled by a dual-arm continuation test, not by point metrics.
Outcome
s1e-500 kept the root role with data behind it; the S2e ladder built on it passed rung after rung, while the b300 line was closed at the cost of one control arm.
Mechanism
Early checkpoints sit near sharp optima with less accumulated robustness structure; their headline metrics reflect the narrow training distribution, not resilience to distribution shifts. A continuation rung is itself a distribution shift, so the root property being selected for is shock tolerance - observable only by actually continuing, never by static gates.
Applies when
- a younger checkpoint outscores the current root on quality gates
- choosing the base for a robustification or command ladder
- a continuation run stalls in an early survival valley
“b300 对照臂 … PD 冲击下存活谷(+100 起 1/3→0/3),预算尽未爬出,800 档 20-seed 存活 13/20 出局——幼年点自身指标再好也扛不住新 DR 适应冲击,成熟度是本钱,s1e-500 根被数据背书”
train/README.md § omni_s2e_pd (b300 对照臂) / s1e 选点重审 A plateau in a training curve was a population mix, not a half-learned skill - 60% standing at 0.372 m and 40% sitting at 0.19 m - and a category at hard zero stayed at zero through 3,000 more iterations
zero-partial-credit-is-not-an-iteration-problemBefore buying iterations for a plateau, split the metric by category and check whether it is bimodal; a category at hard zero with no partial credit is missing a capability or a reachable state, and more iterations under an unchanged config will only polish the categories that already work.
Symptom
After R0.1 the training-side base_height sat near 0.29 m and the curve was still climbing at the iteration cap, which read as "train it longer".
Context
Candidate A (user decision) was a child-run from R0.1's last checkpoint with zero config change - the logged env.yaml files differ only in log_dir - for 3,000 more iterations (R0.2). The per-category acceptance split was already available: prone had scored 0/156 with no partial credit.
Change
Continue training unchanged, then read the result by category rather than by the pooled curve.
Outcome
supine 91.5 -> 96.4%, side 82.2 -> 87.9%, get-up 0.96 -> 0.84 s, pose error 0.91 -> 0.54 - all improvements to categories that already stood. Prone stayed 0/153; mid 55.9 -> 41.2% was within noise (n=34). Height by category was binary - standing groups 0.372/0.373 m, seated groups 0.187/0.194 m, nothing between - so the pooled 0.29 was 0.61 x 0.372 + 0.39 x 0.19 = 0.30 (measured 0.307): six in ten standing, four in ten sitting. The rendered prone episode was still kneel-sitting at t = 8 s.
Mechanism
The pooled mean of a binary outcome only moves when the mix moves; PPO kept polishing the subpopulation that already succeeded while the failing one produced no advantage signal to follow.
Conflicts
R0.2 recorded the missing capability as "prone lacks rolling over"; R0.3's end-state confusion matrix retracted that - prone had righted its torso in 159/159 episodes and was failing to stand from the W-sit. The lesson that iterations could not fix it holds; the named cause was wrong.
Applies when
- a training curve plateaus while acceptance shows one category at zero
- deciding between "train longer" and "change something"
- pooled training metrics are read as the typical episode
“**分类别 h 中位把"平台 = 人口混合"钉死了**:数值是**二值**的 —— 站立组 0.372/0.373,坐姿组 0.187/0.194,**中间没有过渡态**。 … **这也是本仓此后读该指标的通用告诫:全体混合的期望会把 双峰分布平均成一个不存在的中间值,必须分类别看。** … **结论:A 不能过门,原因确定为 prone 缺"翻身"这一技能,不是迭代不够。**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §13 R0.2(recovery_r0_2,child-run 续训):A 走完了 —— 推不动 prone Three hardware accounts locked the run design point - and the knee's real speed ceiling is tau_limit/kd, not the firmware limit
feasibility-accounts-lock-design-pointBefore opening a dynamic-gait training line, compute the full account set - tau_limit/kd effective speed ceilings, joint ROM under the intended reference geometry, and thermal RMS at the duty cycle - and let the accounts lock the design point; move only to pre-registered in-table alternates, re-running the accounts first.
Symptom
The run line was believed to require a firmware raise of the RS06 speed limit (10 rad/s) as a hard precondition, and the feasibility script's motor-envelope scan had marked 80/100 mm foot-lift cells "physically feasible".
Context
Three added accounts re-decided everything. (1) Damping tax: in MIT mode tau = kp*(q_des-q) - kd*qd, so sustained rotation is capped at tau_limit/kd = 12/1.5 = 8 rad/s - below the firmware's 10; at peak speeds 6.7-7.9 rad/s the damping term alone eats 10.1-11.9 N*m (84-99% of the torque limit). "提固件 limit_spd 越不过这道税 —— 它是 kd 与限扭的比,不是固件旋钮." (2) Joint ROM: the feasibility script had checked motor envelopes but NOT joint range - the ankle-pitch ROM caps 1:2:1 leg-shortening lift at 62 mm (soft) / 77 mm (hard), so the 80/100 mm "feasible" cells were voided; also firmware-independent. (3) Ankle thermal: duty 0.40 puts ankle RMS at 87% of continuous rating (0.35 -> 93%); long-period big-stride cells hit both ankle torque peak and heat. Verdict: firmware raise DEQUEUED (50 mm design point needs knee 6.7-7.3 < the 8 rad/s effective ceiling < firmware 10); vel_limit stays 10 so sim == robot. The three accounts uniquely lock the design point - 50 mm lift / T 0.60 s / duty 0.40 - "三笔账 唯一锁定,不是调参空间", with pre-registered alternates allowed only inside the table and only after re-running the accounts.
Change
Design point frozen from accounts; hardware precondition reversed by arithmetic rather than by test; reference amplitude (0.84 rad = FK inverse of 50 mm) derived, per-joint action scales sized to the required travel (knee 0.9, hip_pitch 0.6, ankle deliberately NOT amplified - hard limit is adjacent).
Outcome
A firmware work item left the critical path; an infeasible region of the design space was closed before any training; the remaining risk (knee tracking lag from the damping tax) was pre-registered with its own criterion and in-table fallback (duty 0.35) - "这不是'奖励没调好', 是 plant 账".
Mechanism
PD actuators in MIT mode pay kd*velocity out of the same torque budget that tracks position, so the effective speed ceiling is a ratio of configuration constants, invisible to firmware settings; and feasibility is the intersection of ALL constraint families (torque envelope, joint ROM, thermal RMS) - a scan that omits one family certifies impossible cells.
Applies when
- planning running/jumping or any high-rate gait on PD actuators
- a firmware or hardware upgrade is assumed as a training precondition
- a feasibility scan covers motor limits but not ROM or heat
“膝的有效速度顶 = τ_limit/kd = 12/1.5 = 8 rad/s,不是固件的 10。… 提固件 limit_spd 越不过这道税 —— 它是 kd 与限扭的比,不是固件旋钮。… 可行性脚本只查了电机包络没查关节 ROM —— 其 80/100mm 的"物理可行"格作废。… 判决:RS06 提固件对 run v0 不是前置,出队”
train/RUN_V0_SPEC.md § 1. 硬件账判决 / 2. 步态设计点 Single-impulse push recovery is a binary chaotic quantity - cross-machine floating-point divergence can flip the outcome
single-impulse-recovery-is-chaoticNever gate or compare single-event recovery outcomes across machines or domains: evaluate disturbances as survival distributions over phases and seeds, compare longitudinally on one machine, and treat any single-point cliff as unconfirmed until it survives the statistical protocol.
Symptom
Mac evaluation found a hard "0.8 N*s cliff" (0/3 survival) that the training machine flatly contradicted: the identical protocol (0.8 impulse at 8 s, cmd 0.2) survived 3/3 there, and a 0.6/0.8/2/4 cross sweep survived everything.
Context
The verdict became a named lesson ("跨机混沌课文"): whether one specific push at one specific phase is survived depends on a trajectory that diverges across machines from floating-point differences alone - "单次冲量恢复是二值混沌量, 跨机浮点发散可翻结局". The boundary was drawn precisely: the 20-seed statistical gates DO agree across machines (established precedent), but that agreement cannot be extrapolated to single-point recovery tests. Protocol amended: disturbance evaluation uses multiple push phases (8/10/12 s), >=10 seeds, and only same-machine longitudinal comparisons; the Mac-side recommendation built on the unreproducible cliff was not adopted, while its directionally-consistent small-impulse data was kept.
Change
Push evaluation redefined from single-event pass/fail to multi-phase multi-seed statistics, with cross-machine comparison banned for event-level results and allowed for distribution-level ones.
Outcome
A false hardware-relevant "cliff" was prevented from steering the ladder (the s2e push rung decisions were made on same-machine statistics); the chaos lesson was cited again when real push tests were restricted to qualitative cross-domain use.
Mechanism
Perturbation recovery near the viability boundary has sensitive dependence on initial conditions; different BLAS/GPU reduction orders yield different trajectories from identical configs, so a binary outcome at one phase is machine-specific noise. Averaging over phases and seeds restores a quantity whose expectation is machine-stable.
Applies when
- a push/disturbance result differs between machines or sim and real
- designing push-recovery acceptance tests
- a sharp pass/fail cliff appears in a chaotic-regime evaluation
“训练机上 Mac 原协议 (0.8 @8s cmd0.2) 3/3 全活 … 与 Mac 的 +0.8 0/3 直接矛盾。定性: 单次冲量恢复是二值混沌量, 跨机浮点发散可翻结局;统计门 (20-seed 八门) 跨机吻合的先例不能外推到单点恢复测试。协议改判: 抗推评测多相位 (push 时刻 8/10/12s) + ≥10 seed + 只做同机纵向比”
train/README.md § s2e 支线终章 (跨机混沌课文) Removing a foot-spacing wall passed every simulated gate and made the feet collide on the real robot - nothing priced stance width in the two-foot phase, the policy narrowed to the simulator's self-collision floor, and real calibration offsets closed the last millimetres; the wall came back with a gate
removed-wall-returns-on-hardwareWhen a constraint is removed, name what will govern that quantity instead and add a gate for it; never let a simulator's collision floor be the margin, and when a gate is exceeded by a hair, record the exact numbers and hand the release decision to a person instead of quietly passing it.
Symptom
On the first real-robot try of oneleg_v0 (2026-09-16) the two feet collided; the user also judged the folded foot not high enough.
Context
The V0 reward table had dropped the feet_lateral_distance wall because it seemed to conflict with the hip adduction single support needs. In the two-foot command bucket no remaining term governed stance width, so the policy drifted narrower until the simulator's self-collision stopped it; the sim acceptance had no foot-spacing gate, so 40/40 said nothing about it. On the robot, calibration offsets consumed the margin.
Change
V0.1: the wall restored (-10, minimum 0.16 m), re-checked against measured numbers (a swing-phase lateral spacing of ~148 mm costs 0.12 per step, acceptable); fold weight 0.8 -> 2.0; a ninth gate: minimum foot spacing >= 100 mm and zero leg-contact frames. The removal was kept on record.
Outcome
oneleg_v0_1 (V0r2 model_2200) passed 39/40 with the spacing gate 40/40. The single miss (a 15.4 deg tilt transient against a < 15 deg limit during a side switch, steady 6.9 deg, everything else green) was recorded with its numbers and released for the user to overrule.
Mechanism
An unpriced degree of freedom drifts to wherever the simulator stops it; if that stop is the simulator's own collision model, the policy's margin on hardware is whatever the calibration error leaves.
Applies when
- dropping a reward term that looked redundant or conflicting
- hardware shows a failure no simulated gate measures
- a release candidate misses one gate row by a small amount
“V0 撤墙被真机证伪(2026-09-16):双脚桶没有任何项管站宽,策略贴 sim 自碰撞底线收窄,真机标定偏差一吃**双脚相碰**。 … min ≥ 100 mm 且腿碰 0 帧(eval_straight 同判据)—— … V0 真机双脚相碰暴露 sim 门未看脚距的缺口 … L s2 标称 tilt 瞬态 15.4°(门限 <15, 超 0.4°, 稳态 6.9°, 该跑其余全绿)——换侧瞬态蹭线, 判定放行留档, 用户可否决。”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §4 feet_lateral_distance 行 / §6 验收门 ⑨ / §8 核查单 7 Slowing the gait clock at deployment is out-of-distribution and backfires - lower the commanded speed instead, or train the knob
cycle-time-override-is-oodAny deployment override must correspond to a dimension the policy was trained to handle; to make a parameter field-adjustable, randomize it in training and observe it - otherwise use the levers inside the trained envelope (commands) and leave the knob alone.
Symptom
Real-robot feedback "walks very fast and unstable" suggested slowing the gait; a deploy-side --cycle-time override existed, making "just slow the clock" a one-flag temptation.
Context
A sim sweep of the override on walk_v6 @cmd 0.3 showed monotone degradation away from the trained 0.40 s cycle: at 0.50 s tilt jumped 7.9 -> 13.2 deg and landing force 1.52x -> 2.24x; at 0.80 s (half speed) clearance collapsed to 3 mm - dragging again - with 20 deg tilt. Meanwhile the legitimate lever, lowering the commanded speed with the clock untouched, improved everything monotonically: cmd 0.1 gave 104% tracking, 6.7 deg tilt, minimum slip - the most stable operating point. The file distinguishes the two "slows" explicitly: lower command = smaller steps at the same 2.5 Hz rhythm; a slower rhythm itself requires retraining - randomize cycle_time (e.g. 0.40-0.65 s) during training and expose it as an observation, and only then does --cycle-time become a field-adjustable knob.
Change
Deployment guidance: never ship a cycle-time override the policy was not trained under; respond to "too fast/unstable" with lower commands; schedule clock variability as a training-time (contract-level) change if a field knob is wanted.
Outcome
The sweep quantified the trap before hardware paid for it (dragging and 2.2x landing force at slowed clocks); cmd 0.1 documented as the stable demo point.
Mechanism
The policy is a function fitted around the training distribution; a deploy-side override moves an input (phase rate) to values never seen, so behavior degrades unpredictably - the knob LOOKS like a capability because it exists in the code, but capability lives in the training distribution, not the interface.
Applies when
- a deploy tool exposes overrides (clock, scale, gains) beyond the training distribution
- hardware feels "too fast/aggressive" and a quick knob exists
- deciding between a deploy-side tweak and a retrain
“0.80s | 1.25Hz | 0.165 | 3mm(拖地) | 20.0° … 慢一半直接崩 … 策略按 0.40 训练, 别的周期属分布外。… 降指令速度才是有效杠杆 … cmd 0.1 是最稳的工作点。… 要节奏本身变慢必须重训 —— 训练期把 cycle_time 随机化(如 0.40~0.65s)并作为观测的一维, 部署时 --cycle-time 就成了现场可调的旋钮。”
train/WALK_DIAGNOSIS.md § 2026-08-01 追加: 调慢步态时钟(--cycle-time)在仿真里是反效果 No parameter tuning on the floor - a failing config retries once, then it is out; anomalies go back to sim
no-field-tuning-protocolHardware time is for executing and measuring the pre-registered matrix, never for tuning: failing configs get one retry then elimination, anomalies get recorded and reproduced in sim, and contract-check bypass flags stay unused.
Symptom
Hardware sessions create pressure to fix problems live - nudge a gain, tweak a scale - which destroys attribution and risks the robot.
Context
The anomaly-handling section of the acceptance sheet is three fixed plays: (1) falls at start -> retry once at the same settings; falls again -> that configuration is eliminated, "不现场调参" (no on-site parameter tuning); (2) limit cycle or motor screech -> stop immediately, record the gain level and the joint, reproduce in sim before any discussion; (3) systematic disagreement with sim -> record it as a finding (hardware outranks sim) rather than adjusting anything to force agreement. Related guardrails elsewhere in the sheet: never pass --allow-unstamped / --allow-plant-drift to bypass manifest checks - if it errors, something real is wrong, stop and look.
Change
Field sessions restricted to executing the pre-written matrix; every fix path routed through sim reproduction and the normal config/rung process.
Outcome
Sessions stayed interpretable (each run matched a documented config) and safety overrides never became habit; anomalies arrived back in sim as reproducible cases instead of half-remembered floor stories.
Mechanism
Field-tuned values are measured under adrenaline on one floor with no logging or baselines - they contaminate the config lineage and are unattributable afterwards; and every bypass flag that skips a contract check converts a designed safety property into an operator promise.
Applies when
- a config fails or oscillates during a hardware session
- someone reaches for a live gain tweak or a bypass flag
- writing the anomaly-handling section of a deployment runbook
“起步即摔 → 换档重试一次, 仍摔则该档出局, 不现场调参。出现极限环/啸叫 → 立刻停, 记录档位与关节, 回 sim 复现再议。… 不要给 --allow-unstamped / --allow-plant-drift —— 三枚 ONNX 都已盖章 … 真要报错说明有别的问题, 停下来看。”
train/REAL_RUN_S2.md § 4. 异常处置 Export every CAD part in the whole-machine frame so URDF rotations are zero and inertia is exact
urdf-shared-origin-exportGenerate the model so that correctness is structural: shared-origin STL export, zero rotations, subtraction-only origins, and an explicit 1e-9 g*mm^2 -> kg*m^2 conversion - never hand-rotate inertia tensors.
Symptom
Hand-assembled URDFs accumulate per-link rotation/origin errors and unit-conversion mistakes in inertia tensors - silent plant corruption that no later calibration can cleanly fix.
Context
Documented CAD -> URDF -> USD procedure from a successful Isaac Lab deployment, kept as the recipe if Lucen regenerates its model.
Change
(1) In CAD, align the whole robot to Z-up, X-forward (Isaac Lab convention) and ground the assembly; (2) export each STL with other parts hidden but the machine's shared origin kept, so all parts share one origin, every URDF rotation is 0, and inertia matrices equal CAD values directly; (3) units: Fusion 360 gives g*mm^2, URDF wants kg*m^2 - multiply by 1e-9; (4) link origin = negative of the joint position; link COM = CAD COM minus joint position; joint origin = difference of the two joint positions; (5) after URDF -> USD import, open the USD separately and set it instanceable before saving.
Outcome
A URDF whose rotations are all zero and whose inertia tensors are CAD-exact, eliminating an entire class of hand-transcription plant errors.
Mechanism
Keeping one shared origin turns every frame transform into a pure translation computable by subtraction, and leaves inertia tensors in the frame CAD already computed them in - no rotation of inertia tensors, the most error-prone manual step, is ever needed.
Applies when
- building or regenerating URDF/MJCF from CAD
- inertia or frame bugs suspected in the plant model
- importing URDF into Isaac Lab / USD
“导出 STL 时隐藏其他零件但导出整机——这样所有零件共享同一原点,URDF 里所有 rotation 全是 0,惯量矩阵直接等于 CAD 值 / 单位:Fusion 360 给 g·mm²,URDF 要 kg·m²,乘 1e-9 / link origin = 该关节坐标取负 … URDF → USD 导入后必须单独打开 USD 设成 instanceable 再存”
Experience.md § URDF 制作流程 (lines 87-92) A stronger action_rate penalty cut the median torque demand under the gate and left the p99 at 4x the limit - only a hinge on the pre-clip (computed) torque, weighted by comparison with a peer term, collapsed the tail
tail-torque-needs-hinge-on-computed-demandJudge actuator demand against the deployed limit, read the pre-clip demand (applied torque is censored and gives no gradient on the excess), use an L2 rate penalty for the median and a thresholded hinge on computed demand for the tail, and set a new term's weight from its measured steady magnitude next to a peer term rather than from a back-of-envelope estimate.
Symptom
After R0.5 hip_pitch delivered torque sat at its 12 N*m limit in a typical get-up (demand 119-125% of the limit, p99 4.2x) - zero control margin at exactly the moment modelling error matters.
Context
The 12/17/11 N*m limits are deployment limits written into robot.yaml by set_torque (RS06 at 33% of rated), and simulation uses the same effort_limit - so the gate is judged against them, not the 36 N*m rating (an early reading against the rating was retracted). Applied torque is clipped at the limit - censored data - so demand must be read from computed_torque. With the full-range action contract (hip_pitch scale 1.309, kp 30) a single-step action change of 0.306 already saturates hip_pitch, and action_rate penalizes exactly that change.
Change
R3.0: action_rate_l2 -0.01 -> -0.03 (child-run). R3.1: new torque_headroom = sum relu(|tau_computed|/limit - 0.9)^2, normalized so three motor types share a scale. Its weight was first estimated at -0.1, measured in a 12-iteration run at an effective -0.019 (12x smaller - the estimate had mixed a per-episode-peak p99 with a per-step p99, and at 1% of upright it would have been numerically absent), and set to -0.5 so its steady value (-0.095) matched action_rate's (-0.097).
Outcome
R3.0: sum |da|^2 -64%, success 99.6 -> 100%, delivered median = demand median (the clamp no longer fired in a typical episode), gate PASS at worst 79.9% - but p99 unchanged (hip_pitch 419-432% -> 427-436%). R3.1: p99 hip_pitch -> 148-189% (-57 to -65%), knee 422-439% -> 233-234%, saturation duty -60%, success 100%; the worst joint became hip_roll at 67.7%. Its cost appears in torque-penalty-bought-by-leg-bracing.
Mechanism
A squared-rate penalty presses the whole-episode sum and moves the typical step, not rare spikes; the spikes came from the kp term (large targets while a limb is blocked by the ground - velocity alone could not reach them under vel_limit), and a penalty on applied torque cannot see demand above the clip because every excess sample reads as exactly the limit.
Applies when
- torque demand saturates actuator limits in high-effort skills
- a smoothness penalty improves medians but not peaks
- a new reward term's weight is set by estimate alone
“**必须用 `computed_torque` 而不是 `applied_torque`**:后者被 `effort_limit` 削平, 是删失数据,超限样本全被压成"恰好等于限",对超限部分梯度恒为 0。 … 改按同侪定标取 **−0.5**(稳态 ≈ −0.095,与 `action_rate_l2` 的 −0.097 等量)。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §24 R3.1(torque_headroom 力矩需求越限罚) Prone get-up sat at 0/159 until two gated hinge terms moved the seated feet - first sideways (561 -> 360 mm), then fore-aft (-168 -> +56 mm) - and success went to 158/159 with nothing else changed
prone-dead-end-is-foot-placementWhen a stuck state and a successful state differ geometrically, penalize the discriminating quantity with a gated hinge that is exactly zero in the state the policy actually reaches (measure it - not the nominal), then re-probe: flattening one axis can move the discriminant to another.
Symptom
Prone falls always righted and then sat with the feet splayed wide or tucked behind the hips, from where the policy never stood (0% for four generations).
Context
Four lines of evidence pointed at foot position: the configuration probe (ankles 215 mm apart stood 52.3%, 561 mm apart 0.0%); FK showing the action contract's nominal (a = 0) is itself a 465 mm straddle, so the action_rate and still terms were pulling toward the splits; biomechanics (feet tucked under the body cut peak hip-extension torque 148.8 -> 32.7 N*m, -78%); and HoST's foot-displacement term, which this reward table lacked. The earlier "not a reward hole" reading was corrected to "a gradient hole, not a level hole": at the dead point the heaviest term (upright) was saturated with zero gradient, still paid for not moving, and the one live gradient (base_height) pointed at the thigh-horizontal torque barrier. A prone ROM scan had already ruled out pushing up from prone.
Change
R0.4: feet_spread_excess = clamp(ankle distance - 0.215, 0, inf) x upright gate, weight -2.0, plus a height-decay factor added after measuring that the policy's real standing stance was 406 mm, not the 215 mm nominal (the plain version would have taxed every successful stand 0.38/s). R0.5: the same shape on the fore-aft axis, feet_fore_seated = |fore-aft offset - 0.05| x upright gate x height decay, target +50 mm (the measured natural offset of standing postures). One variable per rung.
Outcome
R0.4: seated ankle distance 561 -> 360 mm, supine/side exactly unchanged, prone 0 -> 1.9%, mid 45.9 -> 62.2%; a probe then showed the discriminant had moved to the fore-aft axis (standing starts +42 to +51 mm, the prone seat -168 mm). R0.5: supine 99.4, prone 99.4, side 100, mid 100%, re-falls 0%; both geometry terms collapsed to ~0 near iteration 13,100 as base_height rose, and the prone fore-aft offset went -168 -> +56 mm - the term's own target, closing the causal chain. The cost, unmeasured at the time: action jitter rose 33% (sum |da|^2 6.82 -> 9.06).
Mechanism
An upright-gated hinge is inert while the robot rolls and exactly zero in the achieved stance, so it adds gradient only inside the stuck basin; a seated robot with its feet behind or outside its COM must make a kinematically unfavourable transition to stand, and moving the feet under the body removes it.
Applies when
- a get-up or transition skill fails from one start category only
- successful and failed episodes differ in a measurable geometric quantity
- a shaping term might tax the posture successful episodes already use
“`recovery_r0_5`,唯一变量 = 追加 `feet_fore_seated`(与 R0.4 同形状,只换测量轴)。 … 对照 R0.4 的 prone(3.1%,360 mm,**−168 mm**):前后偏移从 −168 走到 +56, 正是这一项的目标量,**判别量被消掉后成功率随之到顶** —— 因果链完整。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §20 R0.5(前后向脚位置项):成功率门全过 —— prone 0/159 → 158/159 The shipped checkpoint was chosen by scanning checkpoints on the full gate - neighbours 100 iterations apart failed 1 and 38 cells, late checkpoints degraded - never by taking the last one, and training stopped on signals, not on a schedule
checkpoint-choice-is-a-full-gate-scanChoose a release checkpoint by running the full acceptance battery over a band of checkpoints (including the transfer axis), stop training on measured signals rather than a fixed iteration count, and expect adjacent checkpoints to differ sharply.
Symptom
Gate results moved sharply and non-monotonically between checkpoints of the same run, and the last checkpoint was often not the best.
Context
One-leg V0r1: the 2,000 neighbourhood was best; from 2,500 on the nominal gates degraded (late overtraining); 2,000 itself had one real micro-hop (17.7 mm over 5 frames); 2,300 was all green and shipped. V0r2: failed cells per checkpoint 2,000:19, 2,100:38, 2,200:1, 2,300:3, 2,400:27, 2,500:12, 3,000:18 - 2,200 shipped. The recovery line learned the same from the other side: stopping v2_6 early at a scheduled point left a policy whose re-fall rate had spiked to 9-22% before consolidation healed it ("stop on signals, not on the schedule"), and a continuation's transfer decayed checkpoint by checkpoint while Isaac stayed perfect.
Change
The acceptance rule "scan checkpoints, do not look only at the last one" is written into the one-leg gates (called the S1 discipline); release candidates are chosen from the scan.
Outcome
Both one-leg releases were mid-run checkpoints (2,300 and 2,200) chosen by the full 40-cell battery.
Mechanism
PPO keeps changing the policy after the gates saturate; with no gradient toward the gate's conditions, later checkpoints wander, so gate quality is a noisy function of iteration.
Applies when
- picking which checkpoint of a run to export and stamp
- a run is stopped at a fixed iteration budget
- final-checkpoint results are worse than mid-run smoke tests
“Isaac 侧 S1 纪律: 验收扫 checkpoint,不是只看最后一个。 … 扫描判决: 2000 邻域最优——2500+ 标称面退化(⑤③② 散挂, 晚期过训), 2000 有一例真微跳(L s100 μ1.2, 17.7mm/5帧), 2300 全绿。”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §6 验收门 / §8 核查单 5 与 7 When hardware underperforms, audit deployment knobs before prescribing retraining
deploy-knob-attribution-before-retrainingBefore any "retrain it" decision, reproduce the symptom in sim under the exact deployment configuration; if the symptom follows the deployment knob rather than the checkpoint, fix the knob or randomize it in training - never top-up-train the skill.
Symptom
Real-robot feedback after the C4 deployment - "turning is weak" - with two retraining options on the table: top up turn training, or restart from the s1e root.
Context
The sim account showed the policy turned well (75-81% at pw1.0); the robot was deployed at power-scale 0.8. The 3-6 pp difference between C2 and C4 policies at the same power was noise; the 40-50 pp difference between power levels was the entire effect. Both proposed retraining paths would have burned budget on a non-existent training gap, and restarting from s1e would additionally have discarded the sidewalk skill that took four rungs and a coordinate-bug hunt to obtain.
Change
Decision: retrain nothing. (1) Try pw1.0 on hardware first - sim says net gain; (2) only if 1.0 is unacceptable (heat/feel), the correct training fix is power/torque randomization in the S2 plant line (one variable, fixes turn and backward together) - not skill top-up; (3) restart-from-root explicitly ranked worst.
Outcome
The "weakness" was fully explained by the deployment knob; the sim/real signatures matched the earlier power-derating law verbatim ("与 C2 时代 power 衰减主要伤非前进轴 逐字吻合").
Mechanism
The policy's competence is defined under its training plant; deployment knobs (power scale, teleop mapping, command bands) silently define a different plant. Attributing a deploy-plant effect to a training gap produces exactly the wrong fix - more training on the wrong variable.
Applies when
- real robot underperforms a skill that sim says is fine
- proposals on the table include retraining or re-rooting
- deployment uses any override the trainer never saw (power scale, remapped commands, different control rate)
“正确的训练修法不是补训转向,而是训练时加 power/力矩随机化让策略在 0.8 下自己补偿 —— 单变量,属 S2 plant 线,一次同时修好转向与后退;从 s1e 重训是最差选项:丢掉四轮 + 一个指标 bug 才换来的侧走,而 C2 的转向本来就没问题。”
train/C_LADDER_RUN.md § 3p. 三 处置顺序(回答「补训转向 还是 回 s1e 重训」:都不该) Choose the fork root by which candidate's shortfalls are recoverable, not by headline score
fork-root-recoverable-shortfallWhen picking a checkpoint to fork from, rank candidates by whether their weaknesses are trainable-back, not by current headline metrics; prefer the candidate whose deficits the upcoming training directly pays for.
Symptom
Multiple candidate checkpoints for the omni-command ladder root, each best at something different: fric-3000 had the best tracking precision (vx 88-91%) and hardened plant robustness; s1e-500 had lower precision (vx 82%) but was the only candidate that could still walk backward.
Context
Root selection ran as a data probe, not a preference vote: 6 candidates x 8 out-of-distribution omni commands x 20 seeds = 960 cells (probe_omni_0808.json). s1e-500 @pw1.0 survived 20/20 in all eight conditions including backward at 67% tracking; the deep-trained fric lineage scored backward 0-3/20 despite better forward precision.
Change
Decision criterion made explicit: list what each candidate exclusively wins at, then ask which of those wins the loser could train back. fric-3000's exclusive wins (precision, plant robustness) are both retrainable - precision is directly optimized by the reward, plant hardening is a planned later pass. s1e-500's exclusive wins (backward plasticity 20/20 vs 2/20, disturbance margin 159/160 vs 125/160, push chirality symmetry 40/40 vs 17/40) had all been shown unrecoverable - push-level rungs failed twice, chirality never recovered even with mirror augmentation on. Root = s1e-500.
Outcome
s1e-500 carried the whole C ladder; its backward skill was preserved through C2/C4 gates (regress budget <=2/20 enforced), and the final C4 product passed a 260-cell battery at 20/20 everywhere.
Mechanism
Training can re-earn anything the objective directly pays for, but capabilities that earlier training destroyed and never restored (plasticity, symmetry, robustness margins) are empirically one-way doors. The information-bearing comparison is therefore recoverability of each candidate's deficit, which the team stated as "独占项的可恢复性正好相反 —— 这就是判据" (the exclusive items' recoverability is exactly opposite - that is the criterion).
Applies when
- selecting a resume/fork root among several checkpoints
- one candidate is more precise but another retains a skill the rest lost
- planning a task-extension ladder from an existing lineage
“fric-3000 赢在精度(vx 88~91%…)与 plant 鲁棒性 → 两样都训得回来…;s1e-500 赢在可塑性(C1 20/20 vs 2/20)、抗扰余量(159/160 vs 125/160)、手性对称(推 ±6 N·s 40/40 vs 17/40)→ 三样都训不回来 … 独占项的可恢复性正好相反 —— 这就是判据。”
train/C_LADDER_RUN.md § 0. 为什么根是 s1e-500(数据,不是偏好)