Training Coach
Doctrine
A report may cite any of these as doctrine-N.
doctrine-1Contract freeze and fingerprint disciplineThe policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.
Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).
Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.
doctrine-2Attribution by resolved training params - never eval-override knobsCapability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.
Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).
Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.
doctrine-3PASS gates become constraints; FAIL gates become objectivesOnce a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.
Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.
Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".
doctrine-4One variable per ladder rung - counted against what the checkpoint sawA rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.
Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.
Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.
doctrine-5Pre-register risks, readings, and stop criteria before the ladderBefore a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.
Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).
Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.
doctrine-6Plant parameters are measured, never inventedEvery plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.
Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).
Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.
doctrine-7Sim2sim gate before sim2real - under deployment conditionsEvery checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).
Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).
Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.
doctrine-8Observation honesty - the actor's inputs are a hardware contractThe actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.
Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).
Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.
doctrine-9Reward economics are audited in realized currencyReward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.
Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).
Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.
doctrine-10The zero-cost option must be the desired behaviorFor every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.
Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).
Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.
doctrine-11Measurement discipline: independent referees, signs, distributionsA disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.
Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).
Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.
doctrine-12The deployment pipeline is plantIrreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.
Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).
Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.
doctrine-13DR budget is finite; its distribution is the measured supportRobustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.
Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).
Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.
doctrine-14Gates measure what hardware feels: posture, margins, stripped assistsAcceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.
Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).
Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.
doctrine-15Fork and root selection: recoverability, maturity, frozen rewardsChoose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.
Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).
Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.
doctrine-16Curricula: verified engagement, lineage counters, disease-phase gatingAutomatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.
Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).
Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.
doctrine-17Probe before training: feasibility first, hypotheses in tablesAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.
Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).
Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).
doctrine-18External advice is recomputed locally; values transfer as ratiosEvery external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.
Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).
Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.
doctrine-19Hardware sessions are scripted experiments, not tuning sessionsReal-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.
Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).
Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.
doctrine-20Close questions in writing; restart when the debt is structuralAudited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.
Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).
Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.
doctrine-21Name the quantity in the space it lives inA goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.
Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).
Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.
doctrine-22Continuation needs a live gradient; a release is chosen by a scanContinue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.
Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).
Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.
Experience cards
100 cards matching “single-support-gain-authority-probe”.
Audit which joints your imitation term constrains - a task that needs deviation is fighting the reference
imitation-term-scope-auditList which joints your imitation/deviation terms actually constrain and check the new skill's required motion against that list; for balance-coupled joints deliver references as feed-forward residuals, not absolute-position targets - and never assume "reference = 0" is neutral.
Symptom
Sidewalk would not learn despite a dedicated tracking reward; meanwhile the gait-shaping imitation term (joint_pos_ref) computed its error norm over ALL 12 joints while its reference covered only the 6 sagittal joints - roll/yaw reference was constantly 0.
Context
Two prior generations had shown the forward gait itself was taught by joint_pos_ref, not discovered by PPO (v6 halved the shaping and swing height collapsed 35 mm -> 4 mm). So the reference is load-bearing - but sidewalk requires hip_roll to deviate from nominal, and the all-joints norm punished exactly that deviation: "一边悬赏一边罚过程" (posting a bounty while punishing the process). A follow-up experiment (C4-redo3, free_roll=True releasing the 4 roll joints from the norm) raised the regularization headroom 6x -> 27x yet sidewalk stayed flat and released hip_roll wandered, killing other skills - net negative, withdrawn. A --roll-absolute probe showed the converse failure: pinning roll to a clock-driven absolute trajectory drove tilt 6.9 -> 13.7 deg. Conclusion recorded: absolute-position imitation cannot teach actions that must be superimposed on state feedback.
Change
The audit reframed the problem: neither punishing roll deviation nor freeing roll nor absolute roll tracking works; the reference for a balance-coupled joint must be delivered as feed-forward under the policy's residual control (see feedforward-for-phase-locked-skills).
Outcome
free_roll rung: joint_pos_ref term rose 0.887 -> 1.104 (release confirmed effective) but vy stayed flat; regularization hypothesis eliminated by experiment.
Mechanism
An imitation error norm defines a cage: joints inside it are pulled to the reference in absolute position, so any skill requiring systematic deviation is taxed per step; but joints carrying active balance cannot follow absolute references either, since their correct position depends on state. The scope and the delivery mechanism of the reference are therefore design decisions per joint, not defaults.
Applies when
- adding a skill that moves joints your reference sets to zero/nominal
- an imitation or deviation penalty coexists with a new tracking reward
- considering releasing joints from a shaping term mid-lineage
“前进步态也不是 PPO 自己发现的,是 joint_pos_ref 教出来的(v6 砍半塑形 → 抬脚 35 mm 塌到 4 mm…)。而 ref_joint_offset 原本只写 6 个矢状面关节,roll/yaw 参考恒 0 —— 侧走既没被教,roll 一偏离 nominal 反被 joint_pos_ref 扣分。一边悬赏一边罚过程。”
train/C_LADDER_RUN.md § 3e. 为什么首战 FAIL / 3i. 解锁笼子 Never referee a suspect metric with another metric from the same code - they can share the disease
independent-referee-for-metric-disputesTo adjudicate a disputed measurement, compute the quantity by an independent method from raw state; never accept a sibling column from the same pipeline as the tiebreaker.
Symptom
A triple reversal on one question: sidewalk sign diagnosis (correct) was retracted using a second metric from the same script, then the retraction itself had to be retracted when that second metric turned out to be the buggy one - two opposite-direction errors on the same problem in one day, both written into the execution sheet.
Context
The probe's net-displacement metric suggested the sidewalk reference sign was inverted. Worried about yaw-drift pollution of net displacement, the author checked the same table's body-frame vy_mean column (~0.003 everywhere, 20-50x smaller) and retracted the sign diagnosis. But vy_mean came from mj_objectVelocity, which was silently reporting vertical velocity due to a frame bug - the "referee" was the diseased measurement. Re-measured with a truly independent computation (xmat.T @ qvel, world trajectory), the original diagnosis was confirmed: saw -0.5 gave vy +0.058/-0.130 (76%/106%), consistent with the net-displacement values all along (yaw pollution was real but only 10-21 deg, nowhere near reversal-sized).
Change
Lesson written twice, verbatim, as a hard rule: when questioning a measurement, the referee must be an independent algorithm (different code path, different physical derivation), e.g. rotate qvel by the body matrix directly, or inspect the raw world trajectory.
Outcome
With the independent referee in place the frame bug was confirmed, fixed, and the whole C4 line re-scored - revealing sidewalk had been working (see body-frame-velocity-api-audit).
Mechanism
Metrics sharing a code path (or an upstream API) share failure modes; agreement between them is evidence about the code, not the world. Only a measurement with an independent derivation can break the tie, because its errors are uncorrelated with the suspect's.
Applies when
- two metrics of the same quantity disagree
- about to retract a conclusion based on a second readout
- auditing evaluation code after a surprising result
“我用一个坏指标去质疑一个好指标,并把撤回写进了执行单。教训(写死):质疑一个测量时,不能用同一份代码里的另一个测量当裁判 —— 它们可能同源同病。裁判必须是独立算法(这次的裁判应该一开始就是 xmat.T · qvel[:3],或直接看世界轨迹)。”
train/C_LADDER_RUN.md § 3m. 二 我今天犯了两个方向相反的错 / 3n. 五 元教训 Knee swing collapsed because it directly trades against the slip penalty - price the conflict explicitly and clamp what reward cannot hold
knee-swing-vs-slip-pricingWhen a behavior collapses as another metric improves, look for the term pair trading them and set their price ratio deliberately (with escalation and reverse tripwires pre-registered); where the policy actively spends action budget to undo your target, stop paying more reward and clamp the target space structurally.
Symptom
Knee peak-to-peak swing collapsed across generations - v5 33 deg, v10 26-30, v10b 7-8, v11 6.5-8.6 - and rolling the clock back did not recover it, acquitting the clock; the collapse tracked the gated slip penalty instead: "屈膝与不打滑在当前奖励里直接对抗" - v10b's excellent 93 deg slip was purchased with knee amplitude.
Context
Reward-side flexion fixes had failed three times: raising reference amplitude backfired twice (v9/v11), and v11's deep-squat default was actively fought by the policy - it spent 0.68 of action budget pulling the squat straight ("被策略花 0.68 动作拉直反杀"). v12's design accepted the conflict as real and attacked on two tracks: (1) ECONOMICS - a direct knee_swing_amplitude reward (+0.3, target 0.55 rad, capped at 0.6/step = 55% of tracking), explicitly opposed to the slip penalty by design ("显式对立——这正是设计:v12 就是这场对抗的定价实验"), with an escalation ladder (K +0.3 -> +0.5, then slip -0.5 -> -0.3, one layer at a time) and a reverse tripwire (slip telemetry back at v10 levels -> slip weight to -0.8, accept ~20 deg knee compromise); (2) STRUCTURE - knee target bounds [0.2, 0.9] rad so full straightening is physically impossible (straightest 11.5 deg) and the 0.68 fighting budget is released. A bonus falsifiable prediction was attached: phase-lock strength tracks amplitude (v9_probe 48 deg locked 2.5 Hz; v11 low-amplitude 1.36 Hz unlocked), so if K works, hardware phase-lock should return - one change, two verdicts.
Change
knee_swing_amplitude reward + knee target clamp + pre-registered escalation/reverse levers; the failed reward-side-only approach retired.
Outcome
The lineage was frozen before v12 trained (strategic reset), but the diagnosis stands as the walk line's clearest example of two reward terms trading a behavior between them, with the pricing experiment and structural clamp fully designed and calibrated.
Mechanism
When two terms price opposite aspects of one motion (swing amplitude creates yaw momentum that becomes slip), the optimizer settles wherever the price ratio puts it - patching one side moves the equilibrium, not the conflict. Explicit pricing makes the trade a designed quantity; structural clamps remove the regions where the policy spends budget fighting the designer.
Conflicts
The pricing experiment (K vs slip) was designed and calibrated but never trained - the 2026-08-05 reset suspended v12; the collapse attribution table and the 0.68-action counterattack are measured, the remedy's效果 is untested.
Applies when
- one gait quality degrades in lockstep with another's improvement
- the policy visibly fights a default pose or reference
- repeated reward-side fixes for the same behavior have failed
“膝摆塌在 v10→v10b,头号嫌疑是门控滑移罚(四代实测膝 p2p:v5 33° / v10 26~30° / v10b 7~8° / v11 6.5~8.6°;退时钟没救回 → 非时钟)——"屈膝"与"不打滑"在当前奖励里直接对抗 … 奖励侧修屈膝已三败 … v11 深蹲 default 被策略花 0.68 动作拉直反杀”
train/WALK_V12_SPEC.md § 0. 定位 / 2. K —— 膝摆经济(与滑移罚的对偶) The walk phase machine structurally cannot express flight - rebuild the representation for duty < 0.5, teach flight with a mask tax, never a cliff bounty
phase-machine-structural-limitsWhen a new gait changes the contact pattern's structure, audit whether the phase/mask representation can express it and rebuild the representation if not; teach the new contact pattern with graded mask-mismatch pressure and count it in acceptance with artifact-proof definitions (minimum segment length), never with cliff bounties.
Symptom
Running requires both feet airborne, but the walk-era phase machine switches legs by the sign of sin(phase) - with duty < 0.5 the two swing windows must OVERLAP during flight, which a sign-switching representation cannot express at all.
Context
The run phase machine was re-architected rather than patched: per-leg phases (left = phi, right = phi+0.5 mod 1) with leg_phase < duty defining stance, aligned to the walk sin convention at duty=0.5 so the machines agree where their domains overlap. Flight is taught by the SAME mechanism that once cured foot-dragging, direction reversed: in the two planned flight windows the contact mask is (0,0) and feet_contact_number_duty charges -0.3 per foot still on the ground - a mild ~0.16/step tax, deliberately NOT a cliff: "悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文". The reference shape (half-sine bump over swing progress) is zero at window boundaries by construction, eliminating the clearing-window that the C4 probe measured to cost 13-19% on non-sinusoidal references. The shape self-check ("零代价选项是什么") was run on three behaviors: standing pays ref everywhere (known cmd=0 stepping risk, booked), walking pays only the flight-window tax, proper running collects full marks.
Change
New rewards.py run section (leg_phase_duty / stance_mask_duty / ref_run / clearance_run / contact_duty) with the walk versions untouched byte-for-byte; flight acceptance metric defined with a segment-length floor (>=40 ms to count) so numeric contact flicker cannot fake flight.
Outcome
Flight became expressible and taught by a calibrated mild pressure; the walk lineage's phase code stayed frozen as its own contract.
Mechanism
A phase representation defines which contact patterns exist in the reward's vocabulary; duty cycling below 0.5 introduces states (double-flight) outside a half-period sign convention's language, so no weight tuning can teach them. And rare desirable events taught by cliff-shaped bounties invite hacks (jumping in place); the graded mask tax prices the planned pattern without creating a jackpot.
Applies when
- extending a walking stack to running/jumping (duty < 0.5)
- a desired contact pattern never appears despite reward increases
- defining flight/contact acceptance metrics
“duty<0.5 时摆动窗 (1−duty)T > T/2,两腿摆动窗在腾空段重叠 —— sin 符号切腿的机制结构上表达不了"双脚同时在空中"。… feet_contact_number_duty 对"还踩着地"持续 −0.3/脚 —— 与 walk 治拖地同一机制,方向相反。不腾空的税 ~0.16/步 … 梯度温和不构成悬崖(悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文)。”
train/RUN_V0_SPEC.md § 5. 相位机设计 When hardware underperforms, audit deployment knobs before prescribing retraining
deploy-knob-attribution-before-retrainingBefore any "retrain it" decision, reproduce the symptom in sim under the exact deployment configuration; if the symptom follows the deployment knob rather than the checkpoint, fix the knob or randomize it in training - never top-up-train the skill.
Symptom
Real-robot feedback after the C4 deployment - "turning is weak" - with two retraining options on the table: top up turn training, or restart from the s1e root.
Context
The sim account showed the policy turned well (75-81% at pw1.0); the robot was deployed at power-scale 0.8. The 3-6 pp difference between C2 and C4 policies at the same power was noise; the 40-50 pp difference between power levels was the entire effect. Both proposed retraining paths would have burned budget on a non-existent training gap, and restarting from s1e would additionally have discarded the sidewalk skill that took four rungs and a coordinate-bug hunt to obtain.
Change
Decision: retrain nothing. (1) Try pw1.0 on hardware first - sim says net gain; (2) only if 1.0 is unacceptable (heat/feel), the correct training fix is power/torque randomization in the S2 plant line (one variable, fixes turn and backward together) - not skill top-up; (3) restart-from-root explicitly ranked worst.
Outcome
The "weakness" was fully explained by the deployment knob; the sim/real signatures matched the earlier power-derating law verbatim ("与 C2 时代 power 衰减主要伤非前进轴 逐字吻合").
Mechanism
The policy's competence is defined under its training plant; deployment knobs (power scale, teleop mapping, command bands) silently define a different plant. Attributing a deploy-plant effect to a training gap produces exactly the wrong fix - more training on the wrong variable.
Applies when
- real robot underperforms a skill that sim says is fine
- proposals on the table include retraining or re-rooting
- deployment uses any override the trainer never saw (power scale, remapped commands, different control rate)
“正确的训练修法不是补训转向,而是训练时加 power/力矩随机化让策略在 0.8 下自己补偿 —— 单变量,属 S2 plant 线,一次同时修好转向与后退;从 s1e 重训是最差选项:丢掉四轮 + 一个指标 bug 才换来的侧走,而 C2 的转向本来就没问题。”
train/C_LADDER_RUN.md § 3p. 三 处置顺序(回答「补训转向 还是 回 s1e 重训」:都不该) Run acceptance under measured contact parameters - honest condim/torsional-friction flipped a false PASS into a real-matching FAIL
eval-plant-honesty-contact-paramsPin the evaluation plant's contact model to measured values (contact dimension, torsional/rolling friction, mu) before trusting any gate that involves slip, impact, or drift - a gate can only fail a policy for physics its simulator contains.
Symptom
walk_v5 passed the old acceptance battery yet failed on the real robot (footfall force, drift, kicking) - the evaluation plant was flattering the policy.
Context
The battery was re-run under "honest contact parameters" - condim 4 (adding torsional contact) with measured torsional friction 0.035 - and v5 then FAILED exactly the rows corresponding to its real problems: heading 185 deg (limit 30), support-foot yaw slip 284 deg (limit 80), landing force 1.72x (limit 1.5x), suspended tilt 45.9 deg (limit 10). The slip physics depends on torsional friction, which the default contact model (condim 3) does not even simulate - a slip problem is invisible to an evaluator that cannot represent yaw friction at the foot. Term-sizing measurements for the new rewards were likewise taken under the same honest parameters (cmd 0.45, skipping the 5 s start transient).
Change
Acceptance harness pinned to condim 4 / torsion 0.035 (measured); verdicts issued under defaults declared non-citable for these rows.
Outcome
Sim acceptance verdicts began agreeing with hardware ("现在失败, 与真机一致"); the v6 fixes could be developed and validated against an evaluator that could actually see the disease.
Mechanism
An evaluator is a plant model too: contact dimensionality and friction values decide which failure modes exist in the simulation at all. Evaluating under default contact parameters tests the policy in a world where its real failure is physically impossible, producing structurally false PASSes.
Applies when
- sim acceptance passes policies that fail on hardware
- slip/drift/impact gates run under default simulator contact settings
- setting up a cross-simulator evaluation harness
“accept_v2.py 已加三条判据, walk_v5 在诚实的接触参数下(--condim 4 --torsion 0.035)现在失败, 与真机一致:直行 15s 航向累计 <30° | 185° ✗ … 落脚力峰值 <1.5× 体重 | 1.72× ✗”
train/WALK_V6_MINIMAL.md § 5. 验收 Prove an armless get-up exists before training it - a connected static domain, 25% torque on the cheapest path, an 8 mm hand-over gap, a static roll-over - and write down what each scan cannot represent
get-up-feasibility-accounts-before-trainingBefore training a get-up or any multi-contact skill, compute the quasi-static accounts - connectivity of the static domain, torque along the cheapest path, hand-over gaps, COM shift available for rolling - and state which configurations each scan cannot represent; when a policy gets stuck in one of those, extend the scan before blaming the reward.
Symptom
A torso-and-legs robot has no arms to push off the ground; whether it can get up from the floor at all was unknown when the line opened.
Context
recovery_feasibility.py ran three accounts before any training (the run line's "hard accounts first" discipline): a sagittal quasi-static scan (0.05 rad grid, 44,520 configurations, MuJoCo FK, flat-foot assumption). (1) The static standing domain (COM over the feet, torques in limit) has 29,586 cells, flood-fill connected with no islands, from a 0.097 m deepest squat to the 0.384 m stand. (2) The minimum-torque path peaks at 25% of the limits (knee 2.9/12, ankle 3.7/17 N*m) - a 4x margin. (3) All 508 ground-contact configurations have the contact behind the COM; the smallest gap to pure foot support is 8 mm. A roll-over account: swinging both straight legs to one side shifts the COM 96 mm against a 62 mm torso half-width - 1.6x, so rolling needs no momentum. Three conclusions were written down for later attribution: the legs are 80% of the mass (swinging them moves the COM), prone has no flat-foot hand-over face (merge into a supine/side sit first), and supine needs no sit-up (hip flexion is limited to 75 deg).
Change
The accounts gated opening the line and were cited in every later argument about what the robot can physically do.
Outcome
They held where they applied: in V1.0 every fall category was righted under a hard rate limit, which the spec records as the quasi-static roll-over account verified by training, and the 25% torque path was the basis for pursuing a slow get-up. They also misled once: account (3) is sagittal, and on 08-09 the spec corrected its scope - it cannot represent the splayed W-sit where the policy actually stalled. A follow-up prone hip-ROM scan (471,625 cells) found 3,912 two-foot-contact cells and none with both soles within 25 deg of level (best 40.2 deg): a flat-footed push-up from prone is infeasible on this robot, so the fix became where the feet go after sitting up.
Mechanism
A get-up needs a connected path through statically feasible configurations and enough torque along it; quasi-static accounts bound both cheaply, and momentum can only make the real problem easier. A reduced-dimensional scan, though, only speaks for the configurations it can express.
Conflicts
In R0.1-R0.2 the spec read account (3)'s "prone has no front hand-over" as "prone lacks the roll-over skill"; R0.3's confusion matrix showed prone had righted its torso 159/159, and the spec then restricted account (3) to the sagittal configurations it models.
Applies when
- opening a get-up, recovery or climbing skill on a new robot
- a robot lacks arms or other obvious contact options
- a policy stalls in a configuration a feasibility scan never modelled
“本机 **torso + legs、无手臂可撑地**,开训前先证明存在不依赖手臂的物理解 … 双腿同侧直腿摆最大横移 **96 mm = 1.6×** … **静态摆腿即可翻身,无需动量**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §1 机械可行性判决(2026-08-09,recovery_feasibility.py 三笔账) Three times a joint-angle stand-in for a foot-level quantity was gamed or lied - the absolute ankle roll sold stance width to buy flat feet, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart
joint-space-proxy-for-task-space-quantityExpress foot-level (task-space) goals and acceptance criteria in task space - link attitude and lateral spacing from world poses - never through joint angles that assume other joints are at zero, and measure the task-space value before trusting a joint-space estimate of it.
Symptom
The line's first real-robot get-up (v2_6, 2026-08-11) was "fairly stable", but after standing the feet were too close and the robot slid into the splits and fell several times; later reports added that it also got up through a split posture.
Context
(1) flat_feet penalized sum |q_ankle_roll|, but a flat foot is ankle roll compensating hip roll; the proxy taxed the compensated wide solution, and the untaxed combination was hips straight plus ankles at zero - flat and narrow. On hardware the lateral support shrank to hip width plus centimetres, lateral balance rested on 11 N*m ankle motors, and the feet slid apart. (2) The next rung's acceptance criterion "hip roll >= 20 deg" assumed zero hip yaw; at 50 deg of yaw the lateral contribution is x cos 50 ~ 0.64 - the same substitution again, inside a criterion. (3) A new task-space diagnostic reading the foot links' world poses measured v2_6c's stance at 0.159 m where the kinematic audit from joint angles had said 0.271 m (0.271 x cos 47 ~ 0.17).
Change
Rule written into the spec: task-space quantities are never expressed through joint-space proxies. V3.1's stance terms were all task-space: flat_feet_task from the foot links' world orientation, lateral foot spacing in metres, stand_pose stripped of both roll joints.
Outcome
From scratch with task-space terms (V3.1 P1c): lateral stance 0.355 m, foot residual tilt median 0 deg / P75 2.0 deg, all six criteria passing, mu 1.0-0.4 all 100%.
Mechanism
A joint proxy bundles the goal with everything else those joints do; the optimizer finds the combination the proxy does not tax, and a joint-based criterion silently assumes the other joints sit at their nominal.
Applies when
- rewarding flat feet, stance width, foot placement or end-effector pose
- an acceptance criterion is written in joint angles for a geometric goal
- joints with large yaw or coupled axes are involved
“**病根 = 关节空间代理**:`flat_feet` 罚 Σ|q_ankle_roll|(§41 取的简易口径)。 "脚掌平"的运动学正解是 **踝滚补偿髋滚**(q_ankle_roll ≈ −q_hip_roll) … 代理把"脚平"和"站距"绑死在一起卖了。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §43 真机首试入账(2026-08-11)判读:flat_feet 的代理口径错误 With no term on foot attitude the robot stood on the edges of its feet (ankle roll -26 deg) from R0.5 to V2.5b; pricing it fixed that and turned stance width into the next free variable - the user saw both on video before any metric flagged them
unpriced-foot-attitude-is-a-free-variableList every posture quantity the hardware cares about (foot attitude, stance width, yaw, knee flexion) and make sure each is priced by a term with real gradient at the observed error; pricing one exposes the next, so re-inspect the feet-level video after every change.
Symptom
Watching the v2_5 video the user said the ankle roll after standing looked very strange - the feet were not flat. The feet view showed the right foot standing on its outer edge; median ankle roll at t = 8 s was -/+26 deg (limit +/-35), mirrored. This was the physical form of the ankle_roll saturation criterion that had failed since R0.5.
Context
No term priced foot attitude: feet_contact_upright counts an edge contact as contact, and stand_pose's wide exp kernel (sigma^2 = 9) gives almost no gradient at 0.45 rad. HoST carries an "ankle parallel" term (+20); this table never had one. Edge standing was made a blocking precondition for hardware (continuous ankle load, unstable contact, wear).
Change
V2.6 (single variable): flat_feet = (|q_l_ankle_roll| + |q_r_ankle_roll|) x upright gate x height gate, a linear hinge with a 5 deg margin, weight -2, continued from V2.5b.
Outcome
Ankle roll -/+26 -> 5.3/5.2 deg, ankle_roll saturation 6.6% (< 10%), feet flat on the feet-view video; success 99.6%, re-falls 0-1%. Then the user watched v2_6: hip yaw constantly tense and the legs very close together. The numbers: hip roll -/+4.9/4.8 deg against a nominal 25 - feet flat and hips open 25 deg cannot coexist without ankle compensation, the new term taxed that compensation, and nothing priced stance width. "Foot attitude as a free variable" was fixed and "stance width became the new free variable" - which the real robot then exposed as splits.
Mechanism
An optimizer spends every posture degree of freedom no term prices; closing one reallocates the slack to the next unpriced one.
Applies when
- a standing or landing posture looks wrong on video while gates pass
- a saturation criterion keeps failing on one joint
- a new posture term was just added
“用户看 v2_5 视频:"起身之后 ankle_roll 非常奇怪,脚根本不是平着站立"。 … 机理:奖励表**无任何脚掌姿态项** —— feet_contact_upright 边缘接触也算触地, stand_pose 的 exp 核(σ²=9)对 26°=0.45 rad 梯度≈0。脚掌姿态是自由变量。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §41 预注册 V2.6:flat_feet —— ⑥ 老账的物理形态被用户目视锁定 Record-high training reward hid a fully-failing DR subgroup - aggregate metrics average over draws, gates must test per condition
aggregate-metrics-mask-subgroup-failureNever gate on metrics aggregated across DR draws: evaluate at fixed representative conditions (especially the deployment-critical stratum), and if a difficulty axis matters, ramp it on measured per-stratum success rather than sampling the full range from iteration zero.
Symptom
omni_s1e trained under constant-wide latency DR (0, 0.06 s) posted the lineage's highest-ever Isaac reward (129) - while the --delay 2 smoke evaluation showed 3/3 falls from iter 1500 onward, persisting to early stop; the usable checkpoint window shrank to iters 500-1000.
Context
Diagnosis written plainly: "聚合奖励掩盖重延迟尾部子群体失败" - the aggregated reward averages over latency draws, so the majority of light-delay environments can mask the total failure of the heavy-delay tail. The remedy for the training side was a survival-gated ratchet curriculum (survival_gated_latency): the sampling cap starts at 0.02 s and rises +0.01 only when a 4096-reset window's survival (time_out share) reaches >=90%, capped at 0.06, ratchet up-only - "增益与延迟耐受一起长,不升到策略撑不住的 地方" (gain and delay tolerance grow together; never raise past what the policy can hold). The detection side was already in place from the noise-crutch episode: the per-condition smoke curve, not the training reward, is the health readout.
Change
Latency exposure made curriculum-gated on measured subgroup survival instead of uniform-from-zero; per-condition (--delay 2) smoke evaluation kept as the authoritative curve; watcher scoring adjusted (survival weighted 3x) so recovery during hard phases is not early-stopped away.
Outcome
The failure mode was caught by the smoke curve within one generation; the follow-up redesign (deterministic staged latency) superseded the ratchet, but the aggregate-masking lesson held through both.
Mechanism
Expected-return training weights each DR draw by probability, so a subgroup can contribute bounded loss while being catastrophically failed; any scalar averaged over the randomization cannot distinguish "uniformly decent" from "great on easy draws, dead on hard ones". Only conditioning the evaluation on the stratum reveals the split, and curricula should raise difficulty on measured stratum success, not on schedule.
Applies when
- training reward hits records while a fixed-condition eval degrades
- wide DR on an axis where deployment sits at one known value
- designing curricula for difficulty axes (delay, push, terrain)
“常量 latency DR (0,0.06) 从零训被证伪——Isaac reward 129 历代最高,但 --delay 2 冒烟 iter1500 起 3/3 全摔持续到早停(聚合奖励掩盖重延迟尾部子群体失败,可用窗口只剩 500/1000)。… 采样上限 0.02 起步 … ≥90% 才 +0.01s,0.06 封顶,棘轮只升不降。”
train/OMNI_V0_SPEC.md § 3. S1.5(s1e 训练塌方复盘) When the training reset distribution changes, freeze the acceptance distribution separately and pin its seed - the same checkpoint measured twice differed by 3.8 and 10.2 points
frozen-acceptance-distribution-and-pinned-seedAcceptance distributions are frozen artifacts, decoupled from whatever the training distribution becomes and evaluated with a pinned seed, and every gate row carries its sample size so that a small-n row is never read as a regression.
Symptom
R0.3 added easier roll-arc start states to the training resets, and the acceptance script shared the training category table; separately, the same checkpoint scored side 3.8 points and mid 10.2 points differently on two runs of the same script.
Context
Acceptance ran 512 parallel envs drawn from the fall categories. Had it kept following the training table, 15% of acceptance samples would have landed on states easier than prone - inflated scores and generations that could not be compared. The two same-checkpoint runs had identical per-item height medians, so the ruler had not changed; the spread was reset resampling noise (side n ~ 169, sigma 2.3%; mid n ~ 35, sigma 8.3%). The spec's "20 seeds" had always meant controlled seeds.
Change
accept_recovery.ACCEPT_CATEGORIES pinned to the four R0-R0.2 categories and decoupled from the training FALL_CATEGORIES; --seed 20260809 pinned, after which two consecutive runs were bit-identical. The mid row (n ~ 35) was labelled the bluntest gate.
Outcome
Every later generation (R0.3 through V3.1) was scored on the frozen distribution and seed, which is what let R0.3's intermediate state be read as "no measurable gain" (62.7 -> 62.1%) and R0.2's mid drop be booked as noise rather than a regression.
Mechanism
An acceptance set that follows the training distribution measures a moving target, and an unpinned reset draw adds sampling noise that small-n rows cannot absorb.
Applies when
- the training reset or command distribution changes between generations
- repeated evaluations of one checkpoint disagree
- a small category drives a pass/fail decision
“**分布冻结**:`accept_recovery.ACCEPT_CATEGORIES` 钉死 §5 四类 … 与训练侧 `FALL_CATEGORIES` **解耦**。 … **side 差 3.8 点、mid 差 10.2 点**(h 中位逐项一致,证明不是尺子变了)—— 纯 reset 重采样噪声 … 钉死后两次连跑逐位相同。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §14 验收尺子的两处加固(R0.3 起生效,向后兼容) FK-verify a borrowed reference's structure, then size its amplitude by the reference's job - it pins phase, the policy adds lift
reference-structure-fk-amplitude-divisionWhen borrowing a reference trajectory: verify its structural claim against your own kinematics (an invariant like flat-foot), assign it the phase-pinning job, and size amplitude low enough that the policy contributes the lift - moving toward a proven foreign value in halves, not jumps.
Symptom
walk_v4 had big knee swing (40-46 deg) but only 18-24 mm foot lift - amplitude without hip/knee/ankle phase coordination; later, walk_v5's real-robot swing ballooned to 73.6 deg (sim 55.7) with violent footfalls - amplitude over-driven by the reference.
Context
Structure first: Humanoid-Gym's 1:2:1 hip:knee:ankle reference was verified on the local model before adoption - the ratio exactly satisfies the locally derived flat-foot constraint hip - knee + ankle = 0, FK-tested at multiple amplitudes with sole pitch 0.00 deg throughout. Amplitude second, and here the first reasoning failed honestly: FK said shorter legs need LARGER reference scale (0.30 for 30 mm lift), and the FK was correct - but the premise was wrong ("FK 没错, 但前提错了"): it assumed foot lift must come from the reference. HighTorque Pi, same scale, uses 0.08 with a 0.02 m foot-height target - proof that lift is added by the policy ON TOP of the reference, whose actual job is pinning the phase relationship. Scale 0.30 made the reference the entire gait: over-constrained and over-driven. The correction went to 0.15, deliberately not Pi's 0.08: "一次只走一半, 留退路" (walk half the distance, keep a retreat).
Change
target_joint_pos_scale 0.30 -> 0.15 as one of v6-minimal's three changes, treating both the footfall force and the lateral kicking (yaw momentum scales with leg swing amplitude).
Outcome
v6 improved landing force 1.72x -> 1.55x, suspended tilt 45.9 -> 23.0 deg, turn-gain asymmetry 70% -> 19%; the later v6-halved-shaping experiment (35 mm -> 4 mm collapse) confirmed the reference still carries the gait's existence on this machine - the division of labor is real but machine-specific.
Mechanism
A joint-space reference plays two separable roles: encoding structure (phase relations that keep the foot flat) and injecting amplitude (energy). Structure transfers across robots and is checkable by FK against an invariant; amplitude is a negotiation with the policy, and over-assigning it to the reference removes the policy's freedom to modulate lift with state.
Applies when
- importing a reference gait / imitation target from another codebase
- reference amplitude reasoning based on leg length alone
- real swing amplitude far exceeds sim's under a strong reference
“FK 没错, 但前提错了。我默认抬脚必须由参考轨迹产生。HighTorque Pi 同尺度机器人 … 用 0.08, 而它 target_feet_height = 0.02 m —— 说明抬脚是策略在参考之上加出来的, 参考只负责钉住髋/膝/踝的相位配合。我们取 0.30 等于让参考本身就是整个步态, 过约束 + 过驱动”
train/WALK_V6_MINIMAL.md § ① target_joint_pos_scale 0.30 → 0.15 A proposal in the runbook - torque and action limits as versioned safety tiers (classroom / research / expert) written to motor RAM and read back, separate from the reward's effort penalty - recorded as a proposal, its implementation unrecorded
safety-limits-are-a-layer-not-a-rewardKeep hardware limits as an explicit, versioned safety layer (tiers written and read back at start, the persisted default the safest one) and the effort penalty as a behaviour layer; when a skill needs more torque, change tier deliberately rather than trading one layer against the other.
Symptom
Running and jumping need more torque than the deployed limits allow, and the temptation is to trade the training-side effort penalty against the hardware limit, or to hand a new user a robot "tuned however the last person left it".
Context
A message pasted into the operator runbook (undated, citing Berkeley's practice of storing the full motor configuration as JSON with write and read-back scripts) proposes: configuration is a versioned artifact, not a verbal agreement; three safety tiers in robot.yaml beside the gain and policy profiles - classroom (RS06 limited to 10 N*m, lateral joints clamped: "however bad the policy, it only moves awkwardly"), research (14 N*m, clamps at twice the measured need, the default) and expert (the 36 N*m rating, joint limits only, requiring an explicit flag); deploy writes the tier to motor RAM at start and reads it back, while the stored copy stays classroom so a power cut returns to the safest state. It frames limits as the safety layer and the effort penalty as the behaviour layer - more torque for running means switching tier, not weakening the penalty.
Change
None recorded: the message ends by asking which to do first, a rollback or the tiers.
Outcome
The sources do not record the tiers being implemented; the deployed limits stayed at 12/17/11 N*m through the recovery and one-leg lines (the one-leg spec treats raising the RS06 limit as a separate, unapproved hardware decision). Related and recorded elsewhere: torque limits were written to RAM only in a scripted, self-reversing field experiment.
Mechanism
Hardware limits bound the damage any policy can do; reward terms shape what a policy prefers. Mixing them either weakens safety to buy behaviour or distorts behaviour to buy safety.
Applies when
- a new skill needs more torque than the deployed limits
- robots are handed to students or new users
- motor configuration lives in people's heads or in the firmware only
“配置是版本化的产物,不是口头约定。 … deploy_policy 启动时按档写进电机 RAM 并读回校验(落盘的那份永远保持 classroom,断电自动回到最安全状态)。 … 限幅是安全层,dof_torques_l2 是行为塑造层,它们在不同的层,不冲突。跑步要更大力矩就换档,而不是去动训练里的省力惩罚。”
RL系统/FOLLOW THIS copy 2.md § 面向 developer / 教育机构该怎么做 (pasted proposal, undated) The +/-50 mm lateral COM randomization meant to spread the legs coincided with legs pulling IN - rolled back per its own pre-registered contract
com-dr-rollback-on-symptomWhen adopting a DR value that covers no local measurement, write its intent and rollback trigger into the config at adoption time; roll it back as the control arm the moment the symptom contradicts the intent, and promote the symptom's metric into the acceptance battery.
Symptom
After v7 adopted the reference developer's oversized lateral COM randomization (+/-50 mm) explicitly to force leg spread, the real robot's legs narrowed instead - lateral mean 154 mm / closest 107 mm in sim (nominal 214.5), narrower still on hardware with occasional leg contact.
Context
The rollback was clean because the adoption had been honest: the robot.yaml comment recorded the intent AND that the +/-50 value covered no local measurement (only a 16/7 mm measured offset existed; even the prior widening to +/-20 was subjective), plus the reference's own reported side effect (base sway) and the note "这一项要单独跑、 单独归因". When the opposite symptom appeared, v8 returned y to +/-20 mm as the control arm ("要么没起作用、要么帮了倒忙 … 按约定退回做 对照"), kept x/z untouched (a noise-level difference not worth another variable), and named the second suspect: the landing penalty itself, via the reference's own three-link chain (landing penalty -> stance narrows -> spacing penalty needed). A gate lesson was booked in the same table: v7's sim numbers had ALREADY crossed the line (154/107 vs v5's 182/147) - "这个指标本可拦下 v7" - so foot-distance became a standing acceptance row (min >120 mm, zero leg-leg contacts).
Change
base_com_offset_m y: 0.050 -> 0.020 (x/z kept), regenerated through the export tool rather than hand-editing derived files; foot-distance acceptance row added.
Outcome
A borrowed DR lever with no local measurement basis was retired the moment its symptom contradicted its purpose, at single-variable cost; the metric that would have caught it pre-hardware entered the gate.
Mechanism
DR ranges shape behavior through the policy's robustness strategy, which is jointly determined with every reward term; a lever that forces stance width on one robot can be dominated by a stronger narrowing pressure (landing softness) on another. Levers adopted without local measurement must carry their own rollback trigger, because there is no nominal to argue from when they misbehave.
Conflicts
Causality is not fully closed in the source: the narrowing may come from the landing penalty rather than the COM lever ("腿距的第二嫌疑人是 ④ 本身"); the rollback is the pre-agreed control experiment, not a verdict that the lever caused the narrowing.
Applies when
- importing DR ranges or behavioral-forcing randomizations from references
- a DR lever's observed effect contradicts its documented purpose
- a sim metric existed that would have caught a shipped regression
“⑥ 的本意 … 是逼策略把脚分开;真机结果是脚向内收且偶发相碰——要么没起作用、要么帮了倒忙。… 注释当时就写了"这一项要单独跑、单独归因"。现在症状出现了,按约定退回做对照。”
train/WALK_V8_SPEC.md § 3. 改动 C — 质心随机化退回(撤销 v7-⑥ 的 y 项) Close a question with an audit, then freeze the wording - later symptoms may not reopen it without new hard evidence
frozen-verdicts-semantic-boundariesWhen an audit closes a hardware-vs-policy question, record the closing evidence, freeze a citable wording for future recurrences, and set the reopening bar explicitly; separate robustness perturbations from plant-truth questions so a DR rung's failure can never silently reopen a closed measurement.
Symptom
Recurring directional bias on the robot kept re-suggesting "maybe the hardware/COM/mechanics are asymmetric", threatening to re-litigate questions that audits had already closed - burning attention each time a descendant policy leaned or drifted.
Context
Two boundary decisions were written as permanent: (1) semantic separation - "S2④ COM ±20mm = 纯鲁棒性扰动,不再承担「解释真机后仰」任务" - if the COM-DR rung degrades, the ONLY allowed conclusion is "policy insufficiently robust to COM uncertainty"; reopening "is the CAD COM wrong" is forbidden because the mass audit was completed and closed (@63f9212). (2) a frozen wording for chirality, to be quoted verbatim whenever left/right bias appears in later rungs: observed directional bias = policy-level spontaneous symmetry breaking; plant asymmetry = no supporting evidence after the mass + model symmetry audit; mitigation candidate pi_sym queued, not blocking. The evidential basis was quantitative: the root policy was perfectly symmetric under +/-6 N*s pushes (40/40) while descendants broke (17/40, 13/40) - "手性是 S2 训练中获得的, 根没有; 机械侧已双 PASS 关案, 不重开".
Change
Closed questions carry (a) the audit commit that closed them, (b) a frozen citable wording for recurrences, and (c) an explicit evidence bar for reopening ("无新硬证据不得重开").
Outcome
Later chirality observations (C2's 15 pp turn gap, hip_roll drift bias) were handled as policy-lineage properties with policy-side mitigations, without a single hardware re-audit cycle.
Mechanism
Symptom classes recur under different guises; without a frozen verdict each recurrence re-runs the same expensive investigation and risks a different (worse-informed) conclusion. Freezing verdict plus wording converts recurring symptoms into citations, while the evidence bar keeps the closure honest rather than dogmatic - the root/descendant symmetry comparison is what makes "it's the training, not the machine" checkable at any time.
Applies when
- a recurring symptom keeps suggesting an already-audited hardware cause
- writing conclusions for a completed calibration/audit
- a DR rung's degradation invites re-measuring the plant
“若 S2④ 退化,结论只能是「当前 policy 对 COM 不确定性不够鲁棒」,不得重开「CAD COM 是不是错了」… 手性冻结表述 … Plant asymmetry: no supporting evidence after mass + model symmetry audit … 无新硬证据不得重开机械不对称”
train/OMNI_V0_SPEC.md § 4. 语义分界与手性冻结表述(2026-08-07 用户定,永久) Before adding a command mode, compute what ignoring it costs - the lazy optimum must lose
reward-cost-of-ignoring-auditPrice the do-nothing policy for every new command or objective: compute reward-per-step for "comply" vs "ignore" from the actual table, and only train once ignoring is decisively unprofitable.
Symptom
A new command mode can be silently unlearnable if the reward table makes "ignore the command entirely" nearly free compared to the tracking reward available elsewhere.
Context
For each C rung the team computed the per-step cost of completely ignoring the new command versus ignoring forward: ignoring vx=0.25 costs 1.264/step; ignoring wz=0.20 costs 0.984/step (78% of forward - gradient sufficient, so C2 was certified "zero reward surgery"); but ignoring vy=0.10 cost only 0.020/step - 50-60x weaker, because vy entered the table only as an L2 tax, not a tracking term.
Change
Rule instituted: a rung may claim "no reward change needed" only after this arithmetic shows the ignore-cost is the same order as forward's. For C4 the audit failed, so track_lin_vel_y_exp (+2.0, std 0.15, same form as vx) was added - raising the ignore-cost at cmd_vy=0.10 from 0.020 to 0.718/step (36x).
Outcome
C1/C2/C3 proceeded with zero reward edits, keeping single-variable attribution clean; C4's needed surgery was identified before training instead of after a failed run.
Mechanism
PPO converges to whatever costs least; if the reward margin for obeying a new command is a rounding error against existing terms, the "ignore" policy is the optimum and no amount of training fixes it. The audit prices the lazy optimum explicitly before spending compute.
Applies when
- adding a command axis or task mode to an existing reward table
- a new skill trains flat while other skills stay healthy
- certifying a rung as "no reward change"
“cmd wz 0.20 → 0.984(coarse .473 + fine .491 + L2 .020)… 对照:忽略 vx=0.25 = 1.264(本级 78%,同量级);忽略 vy=0.10 = 0.020(弱 50 倍——那才是 C4 必须加 track_lin_vel_y_exp 的原因)。本级不动奖励表。”
train/C_LADDER_RUN.md § 4. 原地转级(C2)② 奖励梯度已验够 Training-log reward values and fixed-command eval values live on different distributions - comparing them once claimed a 44% improvement that was really 6-10%
same-distribution-reward-comparisonQuote reward-term values only with their distribution attached (command range, DR on/off, environment), and compare across runs only when those match; re-measure in a common environment before claiming any improvement percentage.
Symptom
A v6-era analysis concluded slip had dropped 44% by comparing the training log's Episode_Reward against values calibrated in a fixed-command play environment; a same-condition re-measurement showed the true improvement was 6-10%.
Context
The training log's reward is an expectation over the training command distribution (vx 0.15-0.5, yaw +/-0.6, with pushes and domain randomization); play-environment calibrations are taken at a single fixed command with DR off. Subtracting one from the other compares apples to oranges - the warning was written into the v7 pre-flight: "奖励数值只能在同一指令分布下比较 … 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次)".
Change
Rule adopted: any before/after reward-term comparison must hold the command distribution, DR state, and evaluation environment fixed; training-log values compare only against training-log values of runs with identical command/DR configs.
Outcome
The phantom 44% improvement was retracted; later term-level accounting (e.g. the C4 ignore-floor work) consistently specified its distribution before quoting numbers.
Mechanism
A reward term's expectation depends on the visited-state distribution as much as on the policy; changing the command distribution or DR moves every term's baseline. Cross-distribution differences therefore measure the distributions, not the policy change.
Applies when
- comparing reward telemetry across training runs or vs play evals
- claiming improvement percentages from training logs
- term-level reward accounting for diagnosis
“奖励数值只能在同一指令分布下比较。训练日志的 Episode_Reward 是在训练指令分布上算的(vx 0.15~0.5 / 偏航 ±0.6 / 带推力与域随机化), 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次: 据此以为滑移降了 44%, 同条件对拍只有 6~10%)。”
train/WALK_V7_SPEC.md § 3. 开训自查 ⚠️ A stronger action_rate penalty cut the median torque demand under the gate and left the p99 at 4x the limit - only a hinge on the pre-clip (computed) torque, weighted by comparison with a peer term, collapsed the tail
tail-torque-needs-hinge-on-computed-demandJudge actuator demand against the deployed limit, read the pre-clip demand (applied torque is censored and gives no gradient on the excess), use an L2 rate penalty for the median and a thresholded hinge on computed demand for the tail, and set a new term's weight from its measured steady magnitude next to a peer term rather than from a back-of-envelope estimate.
Symptom
After R0.5 hip_pitch delivered torque sat at its 12 N*m limit in a typical get-up (demand 119-125% of the limit, p99 4.2x) - zero control margin at exactly the moment modelling error matters.
Context
The 12/17/11 N*m limits are deployment limits written into robot.yaml by set_torque (RS06 at 33% of rated), and simulation uses the same effort_limit - so the gate is judged against them, not the 36 N*m rating (an early reading against the rating was retracted). Applied torque is clipped at the limit - censored data - so demand must be read from computed_torque. With the full-range action contract (hip_pitch scale 1.309, kp 30) a single-step action change of 0.306 already saturates hip_pitch, and action_rate penalizes exactly that change.
Change
R3.0: action_rate_l2 -0.01 -> -0.03 (child-run). R3.1: new torque_headroom = sum relu(|tau_computed|/limit - 0.9)^2, normalized so three motor types share a scale. Its weight was first estimated at -0.1, measured in a 12-iteration run at an effective -0.019 (12x smaller - the estimate had mixed a per-episode-peak p99 with a per-step p99, and at 1% of upright it would have been numerically absent), and set to -0.5 so its steady value (-0.095) matched action_rate's (-0.097).
Outcome
R3.0: sum |da|^2 -64%, success 99.6 -> 100%, delivered median = demand median (the clamp no longer fired in a typical episode), gate PASS at worst 79.9% - but p99 unchanged (hip_pitch 419-432% -> 427-436%). R3.1: p99 hip_pitch -> 148-189% (-57 to -65%), knee 422-439% -> 233-234%, saturation duty -60%, success 100%; the worst joint became hip_roll at 67.7%. Its cost appears in torque-penalty-bought-by-leg-bracing.
Mechanism
A squared-rate penalty presses the whole-episode sum and moves the typical step, not rare spikes; the spikes came from the kp term (large targets while a limb is blocked by the ground - velocity alone could not reach them under vel_limit), and a penalty on applied torque cannot see demand above the clip because every excess sample reads as exactly the limit.
Applies when
- torque demand saturates actuator limits in high-effort skills
- a smoothness penalty improves medians but not peaks
- a new reward term's weight is set by estimate alone
“**必须用 `computed_torque` 而不是 `applied_torque`**:后者被 `effort_limit` 削平, 是删失数据,超限样本全被压成"恰好等于限",对超限部分梯度恒为 0。 … 改按同侪定标取 **−0.5**(稳态 ≈ −0.095,与 `action_rate_l2` 的 −0.097 等量)。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §24 R3.1(torque_headroom 力矩需求越限罚) Teleop fed the sidewalk axis a command beyond its training band - feet clipped; give each axis its own speed setting
teleop-command-band-per-axisGive every command axis its own teleop scale, clamped to that axis's training band, and reproduce any hardware incident in sim with the exact deployed command values before touching training.
Symptom
Robot stepped on its own foot when sidewalking left under teleop - and only when going left.
Context
The teleop tool used one speed setting for all axes: --teleop-speed 0.20 applied to A/D sent cmd_vy = 0.20, above the training band's top (0.08-0.18) where foot-spacing margin is thinnest. Sim reproduction of the incident (product policy, pw0.8, 5 seeds x 20 s, true collision threshold = single foot width 104 mm): at vy 0.20 the minimum foot distance was 111-115 mm - 7-11 mm from self-collision - vs 147 mm at vy 0.10. Left was 4x more dangerous than right (25% vs 6% of time inside the 160 mm soft wall at vy 0.10), matching the left-only symptom; the margin did not degrade over time (pressing more just lengthened exposure).
Change
deploy_policy gained --teleop-side (default 0.10), separating the lateral speed from the forward speed so each axis's teleop command sits inside its own trained band.
Outcome
Command now inside the band with 43 mm margin at default; the incident became a quantified, reproduced, closed account rather than a mystery.
Mechanism
The policy's competence envelope is the training command distribution per axis; teleop mappings that share one scalar across axes silently command out-of-band inputs on the weakest axis. Asymmetric risk (left vs right) came from the policy's own chirality bias, so a symmetric command produced an asymmetric hazard.
Applies when
- wiring a joystick/teleop layer over a learned policy
- a hardware incident occurs on one command direction only
- training bands differ across command axes
“A/D 一直与 W/S 共用速度档,所以按 A 下发的是 vy = 0.20 —— 既超训练带(0.08~0.18)上沿 … 0.20(遥控实际值)| 111~115 mm | 7~11 mm … 且左比右危险 4 倍 … 处置:deploy_policy 新增 --teleop-side(默认 0.10),侧移与前进档分开。”
train/C_LADDER_RUN.md § 3p. 一 向左走踩到自己 → --teleop-speed 0.20 同时喂给了 vy Under continuous 3-axis uniform sampling, pure straight-line walking is a zero-measure event the policy never trained
zero-measure-commands-need-mode-samplingEnumerate the exact command points users will actually issue (straight, stop, in-place turn) and give each explicit probability mass via mode sampling with off-axes pinned to zero - never assume a continuous sampler covers its measure-zero subsets.
Symptom
"The robot drifts even in sim when told to walk straight" persisted across reward tunings - because with commands drawn as vx in [0.15,0.5] x vy ~ U(+/-0.2) x wz ~ U(+/-0.6), the event vy=0 AND wz=0 has probability zero: pure straight-line walking was never sampled even once.
Context
Restart evidence item #3: "纯直行是零测度点 … 'sim 里直行就漂'是分布的 必然,不是 reward 没调好" - the drift metric was legitimately drowned by commanded turning (v11's own comment self-documented this). The structural fix is discrete mode sampling: a custom ModeVelocityCommand that first draws a mode by share (stand/forward/back/turn/side/mixed), then draws values only on that mode's axes with all others pinned to exact zero - which is also what preserves single-variable discipline in the C ladder (native 3-axis uniform "采不出'离散模式桶' … 把 C1~C4 的单变量纪律直接毁掉"). The mixed mode later got an ellipsoid constraint rather than a cube for the same reason in reverse - corner combinations of a cube are unrepresentative extremes.
Change
Command generation moved from independent per-axis uniforms to mode-bucket sampling with pinned-zero off-axes (plus 20% rel_standing); acceptance likewise evaluates per mode.
Outcome
Straight-line behavior became a trained, testable mode instead of a measure-zero hope; the C ladder could add one mode per rung with provable isolation.
Mechanism
A policy optimizes expected reward under the command distribution; events of probability zero contribute nothing to the objective, so exact-zero-command behaviors (straight walk, stand, in-place turn) are only learned if the sampler gives them mass. Product-of-uniforms distributions concentrate mass on mixtures and give none to the pure behaviors users actually command.
Applies when
- a "simple" command (straight, stop) underperforms mixtures in sim
- designing command distributions for velocity-tracking tasks
- a ladder needs per-mode isolation for attribution
“纯直行是零测度点:最终 command 为 vx∈[0.15,0.5] × vy∈U(±0.2) × wz∈U(±0.6) 连续均匀,vy=0∧wz=0 从未被专门采样 —— "sim 里直行就漂"是分布的必然,不是 reward 没调好 … Isaac 原生 UniformVelocityCommand 是三轴各自 uniform,采不出"离散模式桶"”
train/OMNI_V0_SPEC.md § 0. 为什么从零 (3) / 三件前置 (1)