Training Coach
Doctrine
A report may cite any of these as doctrine-N.
doctrine-1Contract freeze and fingerprint disciplineThe policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.
Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).
Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.
doctrine-2Attribution by resolved training params - never eval-override knobsCapability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.
Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).
Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.
doctrine-3PASS gates become constraints; FAIL gates become objectivesOnce a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.
Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.
Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".
doctrine-4One variable per ladder rung - counted against what the checkpoint sawA rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.
Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.
Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.
doctrine-5Pre-register risks, readings, and stop criteria before the ladderBefore a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.
Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).
Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.
doctrine-6Plant parameters are measured, never inventedEvery plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.
Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).
Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.
doctrine-7Sim2sim gate before sim2real - under deployment conditionsEvery checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).
Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).
Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.
doctrine-8Observation honesty - the actor's inputs are a hardware contractThe actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.
Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).
Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.
doctrine-9Reward economics are audited in realized currencyReward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.
Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).
Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.
doctrine-10The zero-cost option must be the desired behaviorFor every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.
Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).
Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.
doctrine-11Measurement discipline: independent referees, signs, distributionsA disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.
Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).
Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.
doctrine-12The deployment pipeline is plantIrreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.
Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).
Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.
doctrine-13DR budget is finite; its distribution is the measured supportRobustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.
Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).
Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.
doctrine-14Gates measure what hardware feels: posture, margins, stripped assistsAcceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.
Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).
Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.
doctrine-15Fork and root selection: recoverability, maturity, frozen rewardsChoose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.
Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).
Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.
doctrine-16Curricula: verified engagement, lineage counters, disease-phase gatingAutomatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.
Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).
Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.
doctrine-17Probe before training: feasibility first, hypotheses in tablesAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.
Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).
Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).
doctrine-18External advice is recomputed locally; values transfer as ratiosEvery external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.
Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).
Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.
doctrine-19Hardware sessions are scripted experiments, not tuning sessionsReal-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.
Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).
Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.
doctrine-20Close questions in writing; restart when the debt is structuralAudited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.
Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).
Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.
doctrine-21Name the quantity in the space it lives inA goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.
Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).
Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.
doctrine-22Continuation needs a live gradient; a release is chosen by a scanContinue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.
Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).
Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.
Experience cards
66 cards matching “com-randomization-forces-leg-spread”.
Oversized lateral COM randomization (+/-5 cm) deliberately forces leg spread
com-randomization-forces-leg-spreadDR ranges can be behavior-shaping tools, not just robustness padding: oversize a randomization axis to force a strategy the reward struggles to express - and expect a compensating behavior to appear as the cost.
Symptom
Feet drift toward the centerline and even collide; policy has no incentive to keep a lateral support base.
Context
COM randomization ranges were chosen asymmetrically by axis: lateral +/-5 cm ("比常规大,故意的" - larger than usual, on purpose), fore-aft +/-2 cm, vertical +/-2 cm. The oversized lateral range is not robustness padding but a behavioral forcing function. Lucen logged it as directly relevant to its own roll-channel / sideways leg-kick symptom.
Change
Set COM randomization to lateral +/-5 cm, fore-aft +/-2 cm, vertical +/-2 cm, with the lateral band intentionally oversized to make narrow stances fail during training.
Outcome
Effective at separating the feet on the reference robot; side effect - the base began swaying left-right, which then required a foot-centerline distance penalty (see reward-chain-foot-height-landing-spacing).
Mechanism
Randomizing COM laterally makes narrow-stance policies fall for some draws, so PPO discovers wide stances as the only strategy robust across the band - DR used as an implicit reward. The sway side effect appears because the policy hedges against unknown COM by active lateral correction.
Applies when
- feet too close / self-collision in a learned gait
- roll-axis instability suspected to come from narrow stance
- choosing COM or mass-offset DR ranges
“两脚太近甚至互撞 → 先试质心横向随机化 ±5 cm,逼迫策略把脚分开;有效但引发新问题——基座开始左右摇摆 … 横向 ±5 cm(比常规大,故意的,用来逼出分腿)/ 前后 ±2 cm / 垂直 ±2 cm”
Experience.md § 质心随机化范围 (lines 75, 84-86) The +/-50 mm lateral COM randomization meant to spread the legs coincided with legs pulling IN - rolled back per its own pre-registered contract
com-dr-rollback-on-symptomWhen adopting a DR value that covers no local measurement, write its intent and rollback trigger into the config at adoption time; roll it back as the control arm the moment the symptom contradicts the intent, and promote the symptom's metric into the acceptance battery.
Symptom
After v7 adopted the reference developer's oversized lateral COM randomization (+/-50 mm) explicitly to force leg spread, the real robot's legs narrowed instead - lateral mean 154 mm / closest 107 mm in sim (nominal 214.5), narrower still on hardware with occasional leg contact.
Context
The rollback was clean because the adoption had been honest: the robot.yaml comment recorded the intent AND that the +/-50 value covered no local measurement (only a 16/7 mm measured offset existed; even the prior widening to +/-20 was subjective), plus the reference's own reported side effect (base sway) and the note "这一项要单独跑、 单独归因". When the opposite symptom appeared, v8 returned y to +/-20 mm as the control arm ("要么没起作用、要么帮了倒忙 … 按约定退回做 对照"), kept x/z untouched (a noise-level difference not worth another variable), and named the second suspect: the landing penalty itself, via the reference's own three-link chain (landing penalty -> stance narrows -> spacing penalty needed). A gate lesson was booked in the same table: v7's sim numbers had ALREADY crossed the line (154/107 vs v5's 182/147) - "这个指标本可拦下 v7" - so foot-distance became a standing acceptance row (min >120 mm, zero leg-leg contacts).
Change
base_com_offset_m y: 0.050 -> 0.020 (x/z kept), regenerated through the export tool rather than hand-editing derived files; foot-distance acceptance row added.
Outcome
A borrowed DR lever with no local measurement basis was retired the moment its symptom contradicted its purpose, at single-variable cost; the metric that would have caught it pre-hardware entered the gate.
Mechanism
DR ranges shape behavior through the policy's robustness strategy, which is jointly determined with every reward term; a lever that forces stance width on one robot can be dominated by a stronger narrowing pressure (landing softness) on another. Levers adopted without local measurement must carry their own rollback trigger, because there is no nominal to argue from when they misbehave.
Conflicts
Causality is not fully closed in the source: the narrowing may come from the landing penalty rather than the COM lever ("腿距的第二嫌疑人是 ④ 本身"); the rollback is the pre-agreed control experiment, not a verdict that the lever caused the narrowing.
Applies when
- importing DR ranges or behavioral-forcing randomizations from references
- a DR lever's observed effect contradicts its documented purpose
- a sim metric existed that would have caught a shipped regression
“⑥ 的本意 … 是逼策略把脚分开;真机结果是脚向内收且偶发相碰——要么没起作用、要么帮了倒忙。… 注释当时就写了"这一项要单独跑、单独归因"。现在症状出现了,按约定退回做对照。”
train/WALK_V8_SPEC.md § 3. 改动 C — 质心随机化退回(撤销 v7-⑥ 的 y 项) A torque-tail penalty was paid for by bracing the legs against each other - the second simulator's leg-contact count caught it, and the first explanation ("the trainer can't see self-collision") was retracted from the run's own config
torque-penalty-bought-by-leg-bracingWhen a penalty lowers a demand metric, look for what the policy traded to get there - keep self-contact frames and foot spacing as standing sim2sim readouts - and check any "the trainer cannot see X" explanation against the run's resolved config before it enters the record.
Symptom
After R3.1's torque_headroom term collapsed the demand tail, MuJoCo success fell 100 -> 98% and leg-leg contact frames at mu 1.0 rose 750 -> 2,190 (worst rollout 177 -> 450). The one failure (prone seed 2) had the legs crossed, one foot on the other leg, trapped at 0.067 m - visible on video.
Context
Across the ten prone seeds, foot spacing and leg-leg contact frames were monotonically anti-correlated, and the failure was the extreme of the series. Pulling the legs toward the midline shortens the hip_roll lever arm and lowers torque demand. At the time the spec explained it as Isaac training without self-collisions ("a free lunch in a simulator without self-collision").
Change
Leg-leg contact frames and foot spacing were tracked in every MuJoCo gate; R3.2's candidates were "train with self-collision on" or "a minimum leg spacing term" - not stacked.
Outcome
The next rung's joint-velocity penalty incidentally erased the dependency (2,190 -> 86 frames). On 08-10 the runs' logged env.yaml showed enabled_self_collisions true in both r3_1 and v2_2 (inherited from walk v10): the tangle was physically learned bracing, visible to both simulators, and the Isaac/MuJoCo contact-count gap was mesh and contact fidelity. The "self-collision debt" narrative was withdrawn for the whole line.
Mechanism
A penalty on demand rewards any configuration that lowers demand; legs pressed together act as a mutual support that fails when contact geometry shifts slightly.
Conflicts
§24 attributes the dependency to self-collisions being disabled in training; §36 retracts that from the runs' env.yaml ("§24's mechanism explanation was wrong") and keeps the older sections unedited as history.
Applies when
- a torque, impact or energy penalty improves its metric and cross-sim success drops
- legs or links approach each other after a regularization change
- an explanation relies on a simulator setting nobody checked in the run config
“prone 十个 seed 逐条看,脚距与腿-腿接触帧数单调反相关, 而唯一失败的那条正是最极端的一条 … 机制上说得通:把腿收到身体中线附近能缩短 `hip_roll` 力臂、降低力矩需求”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §24 代价:MuJoCo 成功率 100% → 98%,病因是两腿卡住 Borrow reward values from other robots as ratios (to tracking weight, to leg length) - never as absolute numbers
transfer-ratios-not-absolutesWhen importing any numeric from another robot's config or paper, identify its natural normalizer (tracking weight, leg length, sqrt(g*L), body mass) and transfer the dimensionless ratio; sanity-check against a same-scale robot when one exists.
Symptom
Published configs offered tempting absolute values (swing height target 0.06-0.08 m, weight -20) that would have been wrong for a robot with half the leg length and a different tracking weight.
Context
The cross-check normalized before transferring: G1's feet_swing_height weight -20 against tracking +1.0 is a 20x ratio, so with local tracking at 1.5 the equivalent is -30, not -20. G1's 0.06 m target on a ~0.70 m leg scales to ~28 mm on the local 0.325 m leg (Humanoid-Gym converts to ~23 mm), confirming the locally chosen 0.03 m and explicitly rejecting copying 0.05-0.08 absolutes. A same-class robot (Menlo's 16 kg) was used to sanity-check feet_air_time (+0.5 vs the local 2.0, flagged over-high). A nondimensional check of the same kind later validated the sidewalk speed target (v/sqrt(gL) = 0.081 vs hardware-verified 0.101 - 0.104 - inside the envelope, conservative).
Change
All borrowed values converted through ratios (weight/tracking-weight, height/leg-length, dimensionless speed) before entering the config.
Outcome
The scaled values worked (0.03 m target matched both the scaling law and measured 22-23 mm baseline); no cross-robot absolute was ever copied raw.
Mechanism
Reward economies are scale-relative (only ratios to the tracking term matter to the optimum) and kinematic quantities are morphology-relative (clearance scales with leg length, speed with sqrt(g*L)); absolutes encode the source robot's scale, ratios encode the design intent.
Applies when
- copying reward weights/targets from open-source configs or papers
- setting clearance heights, speed targets, or impact thresholds
- comparing your weights to published tables
“G1 的 feet_swing_height 是 tracking 的 20 倍(−20 vs +1.0)。我们 tracking 是 1.5,按同比例应为 −30 … G1 目标 0.06 m / 腿长 ~0.70 m,换算到我们 0.325 m 腿长约 28 mm;Humanoid-Gym 换算约 23 mm。故 target 取 0.03 m 是对的 … 不必抄 0.05~0.08 的绝对值。”
train/WALK_DIAGNOSIS.md § 修正 ②(权重放大) / 修正 ③(目标高度按腿长缩放) Before training a one-leg stand the spec named the cheapest cheats - hopping on the support foot, a raised foot resting unloaded, a leg tripod - and gave each a countermeasure and a gate; one still appeared and was caught by exactly those gates
enumerate-cheapest-cheats-before-trainingBefore training, list the cheapest behaviours that would satisfy each reward term without doing the task, give each a countermeasure in the reward and a gate in acceptance, and prove the intended behaviour is reachable with a probe - then treat any gate the policy games as evidence about the reward, not the gate.
Symptom
The literature's single-leg benchmark reports eight state-of-the-art general policies holding a clean one-leg stand 0 times out of 90 - they survive by sneaking steps and hops - so the task's first adversary was the policy's own cheating.
Context
The spec's shape self-check ("what is the zero-cost option?") listed, for the one-foot bucket: the cheapest cheat, a foot resting on the ground without load, countered by a 5 N contact threshold plus positive swing income; the second cheapest, small hops on the support foot to reset balance, countered by a continuous support-air penalty plus a gate of zero support-foot flight segments. The probe that preceded training had already seen a third: early low-lift postures "survived" by pressing the swing foot at 78-95 N, a leg tripod, removed by folding the shank back. The two-foot bucket was checked too: its zero-cost behaviour is ordinary standing, with no odd base state.
Change
Countermeasures and gates written before training: swing-contact and support-air penalties, gate 2 (zero swing-foot contact frames above 5 N), gate 3 (zero support-foot flight segments).
Outcome
The first run still found the unloaded-foot cheat (a binary reward band gave it no gradient to lift) - and it was caught, by the contact gates and the cross-simulator comparison, not discovered on hardware. The retrained V0 passed all gates 40/40, including zero support-foot flight after the flight detector was corrected.
Mechanism
A policy optimizes the reward, not the intent; the cheapest behaviours that satisfy the reward are predictable from the reward's structure, and a gate written for each before training turns a silent cheat into a failed row.
Applies when
- designing rewards for balance, contact or "hold still" tasks
- benchmark policies are known to cheat the task
- writing acceptance gates for a new skill
“文献里 8 个 SOTA 通用策略在单脚站基准上 0/90 干净保持, 全靠偷步偷跳活命,这是本任务的第一反作弊对象 … 单脚桶下最便宜的作弊是"脚虚放地上不受力"——接触判定 >5 N 力阈(沿用),配 swing_height_band 正收入拉开。 … 第二便宜是"支撑脚小跳重置"——support_air_penalty 连续罚 + 验收门支撑脚腾空段=0 双保险。”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §0 目标口径 / §5 形状自检(零成本选项是什么) Train with self-collisions ON (filtering nested-link ghost pairs) - the reward wall prevents, the physics makes cheating impossible
self-collision-physics-plus-reward-wallNever train a contact-risk behavior with self-collisions disabled; enable them with an audited filter list for nested/overlapping pairs (zero contacts across a pose sweep), record the fps cost, and keep a calibrated distance penalty as the preventive layer on top.
Symptom
walk_v8 logged 107 frames of leg-on-leg contact while still earning 0.751 tracking score - because training-side self-collisions were OFF, leg clipping was literally imperceptible to the policy ("碰腿在训练里 根本感知不到").
Context
Enabling self-collisions naively is its own trap: an Isaac audit had shown PhysX auto-filters adjacent bodies (base-hip clean for free) but nested links generate ghost forces - calf and ankle_roll overlap 65 mm at the zero pose, producing 12x body-weight phantom forces. The v10 recipe: enable self-collisions, explicitly filter only the two nested pairs (l/r calf-ankle_roll), then run a zero-contact audit at three poses (nominal stand, walk crouch, swing-extreme) requiring contact count = 0, adding any residual pair to the filter and re-auditing; a 500-iter sanity run for NaN and an fps-cost record (measured -8.8%). Redundancy with the reward-side foot-distance wall was argued, not assumed: "N2 离得远(奖励侧预防),SC 碰了疼(物理侧兜底)" - the reward keeps distance at range, the physics makes contact hurt - so the v8-style "clip legs and still score" outcome becomes physically impossible.
Change
enabled_self_collisions=True + 2-pair filter + three-pose zero-contact audit (re-verified at 0.00 N after the later mass update) + fps budget recorded.
Outcome
Leg contact entered the training signal; the audit protocol caught the nested-pair ghost-force hazard before it corrupted training; combined with the calibrated distance wall, later versions held contact = 0 on hardware and in sim.
Mechanism
A hazard absent from the training physics cannot be learned about, no matter the reward; but collision meshes that interpenetrate at rest inject large fictitious forces if enabled blindly. Filtered enabling plus a pose-swept zero-contact audit gives true contact physics with no phantom energy - and layering prevention (reward) with consequence (physics) covers both learning and enforcement.
Applies when
- real robot self-contacts while training scored it healthy
- enabling self-collisions on a model with nested collision meshes
- deciding between reward-side and physics-side fixes for clipping
“PhysX 自动过滤相邻体(base↔hip_pitch 免费干净),幽灵力只在 calf↔ankle_roll(零位嵌套 65mm,12 倍体重)。… 与 N2 互补不冗余:N2 离得远(奖励侧预防),SC 碰了疼(物理侧兜底)—— v8 那种 107 帧互碰拿 0.751 跟踪分的事从此物理上不可能。”
train/WALK_V10_SPEC.md § 4. SC —— 训练侧自碰撞(范围已探明,比想象便宜) Removing a hand trim re-exposed the plant offset it had been silently compensating - and a slope scan told bias from sensitivity
hand-trims-hide-plant-offsetsTreat hand-tuned trims as undocumented plant measurements: before deleting one, find what it compensates and re-house that knowledge in the model or the reward budget; diagnose posture errors with a sensitivity sweep to distinguish constant bias from gain problems.
Symptom
After switching from the old hand-trimmed default to the clean geometric zero, the retrained stand policy's only regression was torso lean: 1.8 deg -> 4.1 deg backward.
Context
The old default's ankle-pitch trim (-0.0489/+0.0628) had been pre-compensating a fore-aft COM mismatch; removing the trim removed the hidden compensation, and the posture reward alone was too weak to win it back. A COM sensitivity scan settled what kind of problem this was: sweeping base COM offset -50 to +50 mm gave nearly identical slopes for old and new policies (~0.026 deg/mm) - "不是质心敏感度问题, 是恒定偏置" (not a sensitivity problem, a constant bias). Fix landed in stand_v1b: posture corrected to +0.24 deg while keeping symmetry (<=0.1 deg) and low effort (0.259), disturbance rejection better than both predecessors. Model credibility was checked the honest way: v0's sim prediction at the real COM position (-22 mm) was -2.31 deg lean vs real measured 2.2-3.1 deg - "预测精准命中" - which is what licensed trusting v1b's -0.52 deg prediction. (Side flag from the same file: a sign convention had been documented wrongly in early comments - gravity_base[0] > 0 is forward lean.)
Change
Trims retired in favor of explicit modeling: symmetric geometric default plus a posture-reward budget sized to carry the real COM offset; the offset itself known (real COM ~22 mm behind model).
Outcome
stand_v1b passed acceptance as the standing lineage's final version; the walk-line requirement "加大躯干姿态惩罚权重" was upgraded from suggestion to mandatory, since walking amplifies what standing tolerates (real walk_v1 hit 26 deg lean vs sim 7.4).
Mechanism
Hand trims are plant knowledge stored in the wrong place - invisible, asymmetric, and stale after recalibration; removing them re-exposes the raw plant error. A sensitivity sweep separates the two possible diagnoses (slope change = control problem; parallel offset = constant plant bias), each with a different fix.
Applies when
- cleaning up hand-tuned offsets/trims in defaults or calibration
- a posture bias appears after a default or calibration change
- deciding whether a lean is a COM-sensitivity or constant-offset issue
“两者斜率几乎相同(≈0.026°/mm),v1 只是整体多后仰约 2.4° —— 不是质心敏感度问题,是恒定偏置。成因:旧 default 的踝俯仰 trim(−0.0489/+0.0628)本就预补偿了前后质心偏差,换成零位 default 后这份补偿没了 … v0 在真机质心处(−22 mm)的 sim 预测为 −2.31° 后仰,真机实测 2.2~3.1° 后仰 —— 预测精准命中。”
train/RETRAIN_v2.md § 4b. stand_v1 独立验证结果 / 4c. stand_v1b 验收结果 Before training a one-leg stand, the accounts and a probe showed the default gains could not hold it at all - kp 20 needs 0.39 rad of error to carry the static roll moment, more than the whole adduction range - so per-joint gains came first, and thermal limits set the session length
single-support-gain-authority-probeBefore training a posture that loads one joint statically, compute the steady tracking error load/kp and the series stiffness against m*g*h, and prove with a simple hand-written controller that the posture can be held under the deployment gains - change the gains first if it cannot; then size session length from the thermal account.
Symptom
The one-leg line (standing on one foot, the other folded back, no hopping) had to decide whether the existing gain profile could hold single support before any reward was designed.
Context
Hardware accounts (9.792 kg, COM 0.234 m high, 170 x 80 mm feet, legs 80% of the mass): moving the COM over one foot needs 107 mm of shift and the 20 deg hip-roll adduction range gives 131 mm - geometrically enough. The static frontal moment is 7.8-9 N*m, within RS02's 17 N*m - torque is enough. But at kp 20 carrying 7.8 N*m needs 0.39 rad of tracking error, more than the entire adduction range, and the real robot had already shown it: commanded +0.17, actual -0.04 (0.21 rad droop) under load, 0.0008 rad hanging - load, not the motor. A probe (probe_oneleg.py) then showed open-loop PD cannot hold single support on physics grounds, so the criterion became "an equilibrium exists and a hand-written 4-gain COM feedback can hold it": single-support roll stiffness is hip and ankle in series and must exceed m*g*h_com = 22.5 N*m/rad; ankle kp 12 in series with hip kp 80 gives only 10.4 (open loop 16/16 fell), ankle 60 with hip 80 gives 34.3 (52% margin).
Change
A per-joint gain profile (rl_oneleg: hip_roll kp 80, ankle_roll kp 60, the rest as rl_default) - which needed per-joint gain support in robot.yaml, the bridge, deploy and the trainer's actuator groups - decided before training. Thermal account: single support makes hip_roll the dominant heat load (about 7.8 N*m against a 7 N*m continuous rating), so acceptance and demos run in segments of at most 60 s with a temperature check.
Outcome
Under rl_oneleg the hand-written feedback held six cells cleanly for 6 s (hip_roll steady torque 2.1-3.4 N*m, half the thermal budget); under rl_default the same feedback on the same cells fell 0/4. The trained V0 policy then passed its 40-cell acceptance.
Mechanism
With PD position control, the steady error needed to carry a static load is load/kp; when that error exceeds the joint's range the posture is unreachable whatever the policy does, and series compliance between joints lowers the effective stiffness below the gravity stiffness that single support demands.
Applies when
- single-support, crouched or one-arm-load postures on PD actuators
- a joint "droops" under load on hardware but tracks well when hanging
- deciding whether a new skill needs its own gain profile
“但 kp=20 时撑住 7.8 N·m 需要 **0.39 rad 跟踪误差 > 整个内收行程**。真机已实测: 命令 +0.17 实际 −0.04(droop 0.21 rad),悬挂时 0.0008 rad——是负载不是电机。 … 单支撑滚转是 hip/ankle **串联**刚度,必须 > m·g·h_com = 22.5 N·m/rad;ankle kp12 串 hip80 只有 10.4(开环 16/16 全摔),60 串 80 = 34.3(裕 52%) … **rl_default 同反馈同格 0/4 全摔**(增益档必要性对照)”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §1-1 单脚站: 几何可行,卡点是 hip_roll 增益权限 / §2 A 线增益 / §5 probe 定谳 The run policy never left the ground and fell in the second simulator from the frontal plane - its DR (gains and latency only) covered the actuator axis, not the frontal-plane contact and inertia disturbances the doubled stride amplified; "is DR on" is the wrong question
thin-dr-judged-by-channel-coverageJudge a DR recipe by whether its randomized terms cover the channel where the skill can lose stability, not by whether DR is enabled; when a new skill lengthens single support or enlarges motion in one plane, add disturbances in the plane it destabilizes before training.
Symptom
run R1 (6,000 iterations, 78 min): no flight phase ever appeared, and every one of 13 checkpoints failed the eight-gate MuJoCo smoke. In Isaac: zero terminations in 6,000 iterations, 4.2 deg tilt. In MuJoCo at delay 2: 1/6 survived, falls within 1.9-6.2 s at 50.8-58.7 deg, the most saturated joints all roll joints.
Context
The run contract doubled sagittal travel (knee action scale 0.9, knee swing peak 1.14 rad) with a 0.60 s period and 0.40 duty - long single support - while roll/yaw scales were deliberately left at 0.5. DR copied the s1e recipe: kp/kd (0.9, 1.1) and latency on; mass, COM, joint friction and push all off; ground friction pinned at (1.0, 1.0). Flight was read two independent ways: Isaac's per-foot contact reward stayed 0.845-0.857, never above 0.87 - the arithmetic ceiling of a gait with zero flight - and 30 of 36 MuJoCo seeds had flight fraction exactly 0 (the nonzero six were all tumbling falls). Foot lift itself worked (46-59 mm against a 50 mm design point): the walk-era "not enough travel" failure did not recur.
Change
Verdict FAIL, with the pre-registered first knob (exploration noise 1.0 -> 1.2) explicitly rejected as aimed at a different axis. The lesson was generalized and applied at the next line's design review: the one-leg spec made push, body mass, base COM and friction DR mandatory for its permanent single support and banned the thin recipe.
Outcome
The run line did not continue past R1 in the sources. The one-leg V0 with the wider DR passed its friction-variant gate (mu 0.4 and 1.2) inside a 40/40 acceptance.
Mechanism
Randomizing gains and latency covers the actuator's axis; a skill whose failure lives in frontal-plane contact and inertia needs randomization on that channel (push, mass, COM, friction), or the trainer's exact plant becomes the only one the policy can stand on - the omni_s1 transfer trap a second time, this time with DR switched on.
Applies when
- a policy is flawless in the trainer and falls immediately in a second simulator
- reusing a DR recipe from a skill with a different support pattern
- failures concentrate on one axis (roll, yaw) the DR does not touch
“**机理**: 矢状面行程翻倍 (膝摆动峰 1.14 rad) + T 0.60 + duty 0.40 的长单支撑, 把额状面扰动放大了一个量级; 而 roll/yaw 通道按 §3 **刻意没有放大** (仍 0.5), DR 又是 s1e 复刻的薄配方 (mass/COM/关节摩擦/push **四关全关**, 地面摩擦钉死 (1.0, 1.0))。 … 说明**薄 DR 的判据不能只看"有没有开 DR"**, 要看**开的那几项 是否覆盖失稳所在的通道** —— kp/kd 与延迟是执行器轴向的, 对额状面接触/惯性 扰动零覆盖。 … 0.87 正是「零腾空的走路步态」的天花板算术”
git:Lucen V2@origin/run-line:train/README.md § run R1 FAIL (2026-08-09, run 21-30-30_run_r1): 腾空零, 但病根在额状面不在探索 Prove an armless get-up exists before training it - a connected static domain, 25% torque on the cheapest path, an 8 mm hand-over gap, a static roll-over - and write down what each scan cannot represent
get-up-feasibility-accounts-before-trainingBefore training a get-up or any multi-contact skill, compute the quasi-static accounts - connectivity of the static domain, torque along the cheapest path, hand-over gaps, COM shift available for rolling - and state which configurations each scan cannot represent; when a policy gets stuck in one of those, extend the scan before blaming the reward.
Symptom
A torso-and-legs robot has no arms to push off the ground; whether it can get up from the floor at all was unknown when the line opened.
Context
recovery_feasibility.py ran three accounts before any training (the run line's "hard accounts first" discipline): a sagittal quasi-static scan (0.05 rad grid, 44,520 configurations, MuJoCo FK, flat-foot assumption). (1) The static standing domain (COM over the feet, torques in limit) has 29,586 cells, flood-fill connected with no islands, from a 0.097 m deepest squat to the 0.384 m stand. (2) The minimum-torque path peaks at 25% of the limits (knee 2.9/12, ankle 3.7/17 N*m) - a 4x margin. (3) All 508 ground-contact configurations have the contact behind the COM; the smallest gap to pure foot support is 8 mm. A roll-over account: swinging both straight legs to one side shifts the COM 96 mm against a 62 mm torso half-width - 1.6x, so rolling needs no momentum. Three conclusions were written down for later attribution: the legs are 80% of the mass (swinging them moves the COM), prone has no flat-foot hand-over face (merge into a supine/side sit first), and supine needs no sit-up (hip flexion is limited to 75 deg).
Change
The accounts gated opening the line and were cited in every later argument about what the robot can physically do.
Outcome
They held where they applied: in V1.0 every fall category was righted under a hard rate limit, which the spec records as the quasi-static roll-over account verified by training, and the 25% torque path was the basis for pursuing a slow get-up. They also misled once: account (3) is sagittal, and on 08-09 the spec corrected its scope - it cannot represent the splayed W-sit where the policy actually stalled. A follow-up prone hip-ROM scan (471,625 cells) found 3,912 two-foot-contact cells and none with both soles within 25 deg of level (best 40.2 deg): a flat-footed push-up from prone is infeasible on this robot, so the fix became where the feet go after sitting up.
Mechanism
A get-up needs a connected path through statically feasible configurations and enough torque along it; quasi-static accounts bound both cheaply, and momentum can only make the real problem easier. A reduced-dimensional scan, though, only speaks for the configurations it can express.
Conflicts
In R0.1-R0.2 the spec read account (3)'s "prone has no front hand-over" as "prone lacks the roll-over skill"; R0.3's confusion matrix showed prone had righted its torso 159/159, and the spec then restricted account (3) to the sagittal configurations it models.
Applies when
- opening a get-up, recovery or climbing skill on a new robot
- a robot lacks arms or other obvious contact options
- a policy stalls in a configuration a feasibility scan never modelled
“本机 **torso + legs、无手臂可撑地**,开训前先证明存在不依赖手臂的物理解 … 双腿同侧直腿摆最大横移 **96 mm = 1.6×** … **静态摆腿即可翻身,无需动量**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §1 机械可行性判决(2026-08-09,recovery_feasibility.py 三笔账) Reward fixes come in causal chains - foot height, then landing impact, then foot spacing
reward-chain-foot-height-landing-spacingPlan reward shaping as a chain, not a point fix: when you patch a degenerate gait behavior, pre-register which adjacent behavior the optimizer will exploit next and watch for it.
Symptom
Three problems appeared strictly in sequence: (1) swing feet lifted too low; (2) after fixing that, feet slammed down - "实际比视频里暴力得多" (far more violent in person than on video); (3) after fixing that, feet drifted too close together and collided.
Context
Each reward fix removed one degenerate optimum and exposed the next. The fix for foot spacing (COM lateral randomization +/-5 cm to force leg spread) itself caused base side-to-side sway, requiring a further foot-to-centerline distance penalty. Lucen had just solved its own foot-height problem (19mm -> 40mm swing height) and logged landing impact and foot spacing as the predicted next two problems.
Change
Chain of additions - (1) penalty when swing foot below 5 cm; (2) landing vertical-velocity penalty at touchdown; (3) COM lateral randomization +/-5 cm, then foot-centerline distance penalty to cancel the induced sway.
Outcome
Reference robot progressed through each stage; each individual fix worked and predictably surfaced the successor problem. For Lucen the chain served as a pre-registered roadmap of what breaks next.
Mechanism
Locomotion rewards are coupled through contact dynamics: raising swing height adds potential energy that must go somewhere at touchdown (impact); penalizing impact and forcing robustness to COM shifts changes lateral support strategy (spacing/sway). The optimizer always exploits the cheapest unpenalized channel, so fixing one channel routes the exploit to its neighbor.
Applies when
- adding a foot-height / clearance reward
- feet slam or landing impact grows after a clearance fix
- feet converge toward the centerline or self-collide
- any single-reward fix to a coupled gait behavior
“抬脚太低 → 加惩罚:摆动足低于 5 cm 就扣分 / 加完之后砸脚 → 抬起来了但落地极猛,"实际比视频里暴力得多" → 加落地速度惩罚 … / 两脚太近甚至互撞 → 先试质心横向随机化 ±5 cm … 有效但引发新问题——基座开始左右摇摆 → 再加足-中心线距离惩罚 … 这三条是串联的:每个修复都会暴露下一个问题。”
Experience.md § 三个问题的解法链 (lines 72-77) Isaac splits Coulomb friction into static and dynamic columns - wiring only static means zero loss during motion, silently discarding the identified value
sim-api-friction-columnsWhen installing identified actuator parameters, map each measured quantity to the simulator's exact API column for the operative regime (dynamic for moving loss, viscous for damping), verify per joint after landing, and audit how randomization intervals fall on each column's nominal.
Symptom
The hardware-identified Coulomb friction (tau_c) was about to be installed into the trainer through the friction= field alone - which in Isaac 5 populates only STATIC friction, so during motion the joints would lose no torque at all: "只给 static 则运动中不损耗, 辨识的 τ_c 走路时等于没接" (the identified tau_c would effectively not be connected while walking).
Context
The v12 integration wired all three columns deliberately: armature= and friction= from the 2026-08-04 hardware identification, PLUS dynamic_friction= (Isaac 5 splits Coulomb into static/dynamic; the moving-loss column is dynamic) and viscous_friction= (= the measured joint damping 0.02, aligned to MJCF's damping). Each value was re-checked per joint after landing. A DR interaction was audited and booked rather than hidden: randomize_joint_parameters jitters ALL friction columns with ONE interval - the [-0.05, +0.10] band was calibrated against the Coulomb nominal, and landing on the viscous nominal 0.02 it becomes [0, 0.12], "偏宽但保守" (wide but conservative), accepted with the note that pre-viscous behavior was already [0, 0.10] on a base of 0.
Change
Measured actuator parameters installed across all applicable API columns (armature, static, dynamic, viscous), with the DR side effect on shared randomization intervals audited and recorded.
Outcome
The first generation where the identified plant actually acts during motion in the trainer; the silent-column failure mode documented before it cost a training run.
Mechanism
Physics engines decompose "friction" differently (single coefficient vs static/dynamic/viscous columns); a measured parameter is only installed when it reaches the column the solver reads in the regime that matters (motion, not stiction). Randomizers that share one interval across columns rescale the band by each column's nominal - a hidden unit change.
Applies when
- installing identified friction/armature into any trainer
- porting plant parameters between simulators or engine versions
- joint losses in sim do not match bench measurements during motion
“并额外传 dynamic_friction=(Isaac 5 把库仑拆 static/dynamic 两列,只给 static 则运动中不损耗,辨识的 τ_c 走路时等于没接)与 viscous_friction=(= joint_damping 0.02,对齐 MJCF damping)。… randomize_joint_parameters 用同一个 friction 区间抖三列, [-0.05,+0.10] 是按库仑标称标的,落到粘滞标称 0.02 上成了 [0,0.12]”
train/WALK_V12_SPEC.md § 7. 核查单 (Isaac 接 V.ACTUATORS 新字段) Friction DR was demoted after a measurement (94% success at mu 0.4 with no friction randomization) and promoted again when the action contract changed and mu 0.4 fell to 76% - DR priorities belong to a plant and contract, not to a task
friction-priority-re-measured-after-plant-changeRe-measure transfer along the friction axis for every new action contract or plant, not once per task; a DR priority settled under one action parameterization does not carry to the next.
Symptom
Getting up is all scraping and pushing against the ground, and training pinned friction at 1.0, so friction looked like the first thing to randomize.
Context
The MuJoCo gate on R0.5 (5 categories x 10 seeds x 4 friction levels) measured 100/100/98/94% at mu 1.0/0.8/0.6/0.4: degradation showed first as time (prone 3.2 -> 5.3 s), not failure, so friction DR was demoted and the DR budget earmarked for mass/COM. After the switch to the beta-anchored action space, V2.2 read 90/94/90/76%: mu 0.4 was now the weak row.
Change
V2.3 (single variable): friction DR static (1.0, 1.0) -> (0.2, 2.0), dynamic (0.15, 1.6), the HiFAR range keeping the base dynamic/static ratio; restitution untouched. Continued from v2_2.
Outcome
Isaac nominal 99.8% (DR did not hurt the nominal plant); MuJoCo 98/98/96/92% - mu 0.4 76 -> 92%, mu 1.0 back to R3.1's 98% with bounded torque.
Mechanism
How much a policy leans on friction depends on how it moves; the spec records that the sensitivity rose after the action contract changed but does not establish why.
Applies when
- changing the action space, gains or authority of an existing skill
- deciding which DR axis to spend the next rung on
- an earlier sweep justified leaving an axis unrandomized
“**μ 砍到 0.4(训练值的 40%)仍有 94%**,退化先体现在**用时**(prone 3.2→5.3 s) 而不是成败。μ≥0.8 完全无损。→ **§17 曾把"摩擦随机化提到 R4 第一项"当作优先 事项,这条实测把它降级了**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §21 MuJoCo 复核门 ② 摩擦依赖 Close a question with an audit, then freeze the wording - later symptoms may not reopen it without new hard evidence
frozen-verdicts-semantic-boundariesWhen an audit closes a hardware-vs-policy question, record the closing evidence, freeze a citable wording for future recurrences, and set the reopening bar explicitly; separate robustness perturbations from plant-truth questions so a DR rung's failure can never silently reopen a closed measurement.
Symptom
Recurring directional bias on the robot kept re-suggesting "maybe the hardware/COM/mechanics are asymmetric", threatening to re-litigate questions that audits had already closed - burning attention each time a descendant policy leaned or drifted.
Context
Two boundary decisions were written as permanent: (1) semantic separation - "S2④ COM ±20mm = 纯鲁棒性扰动,不再承担「解释真机后仰」任务" - if the COM-DR rung degrades, the ONLY allowed conclusion is "policy insufficiently robust to COM uncertainty"; reopening "is the CAD COM wrong" is forbidden because the mass audit was completed and closed (@63f9212). (2) a frozen wording for chirality, to be quoted verbatim whenever left/right bias appears in later rungs: observed directional bias = policy-level spontaneous symmetry breaking; plant asymmetry = no supporting evidence after the mass + model symmetry audit; mitigation candidate pi_sym queued, not blocking. The evidential basis was quantitative: the root policy was perfectly symmetric under +/-6 N*s pushes (40/40) while descendants broke (17/40, 13/40) - "手性是 S2 训练中获得的, 根没有; 机械侧已双 PASS 关案, 不重开".
Change
Closed questions carry (a) the audit commit that closed them, (b) a frozen citable wording for recurrences, and (c) an explicit evidence bar for reopening ("无新硬证据不得重开").
Outcome
Later chirality observations (C2's 15 pp turn gap, hip_roll drift bias) were handled as policy-lineage properties with policy-side mitigations, without a single hardware re-audit cycle.
Mechanism
Symptom classes recur under different guises; without a frozen verdict each recurrence re-runs the same expensive investigation and risks a different (worse-informed) conclusion. Freezing verdict plus wording converts recurring symptoms into citations, while the evidence bar keeps the closure honest rather than dogmatic - the root/descendant symmetry comparison is what makes "it's the training, not the machine" checkable at any time.
Applies when
- a recurring symptom keeps suggesting an already-audited hardware cause
- writing conclusions for a completed calibration/audit
- a DR rung's degradation invites re-measuring the plant
“若 S2④ 退化,结论只能是「当前 policy 对 COM 不确定性不够鲁棒」,不得重开「CAD COM 是不是错了」… 手性冻结表述 … Plant asymmetry: no supporting evidence after mass + model symmetry audit … 无新硬证据不得重开机械不对称”
train/OMNI_V0_SPEC.md § 4. 语义分界与手性冻结表述(2026-08-07 用户定,永久) The real robot's right-leg kicking was over-trained-delay times loop gain - irreducible pipeline latency is plant, model it fully from day one
pipeline-latency-is-plant-not-drMeasure the end-to-end action pipeline delay and build it into the nominal plant and every acceptance gate from day one; treat power/scale deratings that "fix" oscillation as gain-reduction crutches flagging an unmodeled delay, and expect higher-feedback-gain policies to be MORE delay-fragile.
Symptom
On hardware, s1c/s1d at action scale 1.0 always kicked wildly with the right leg (s1c only ran as SOTA at power 0.8; s1d only at 0.7) - while sim showed nothing under default evaluation.
Context
Sim reproduced the incident item by item once the real pipeline delay was injected: s1d@1.0 with --delay 1 fell at 10.2 s, --delay 2 at 5.2 s; s1c@1.0 stressed (r_hip_roll saturation 5 -> 16%; "右脚" = the policy's chirality makes the right leg its high-gain leg); and the combos that worked on hardware (s1c@0.8+delay2, s1d@0.7+delay2) all survived in sim. Mechanism: the real pipeline is ~1-2 ticks (BusWorker next-cycle pickup + CAN round trip) but S1.1 trained only to 1 tick - "超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定" (delay beyond training x full loop gain = oscillation; the power derating had been buying stability by compressing loop gain). s1d was MORE fragile than s1c because its yaw 3-layer stack had learned higher feedback gain - higher gain, lower delay tolerance. Three changes: latency DR widened to cover reality; acceptance gates and smoke runs moved permanently to --delay 2 ("门必须在真机条件下预测 真机"); and the doctrine written twice-paid: "不可约的管线属性(延迟、 限速)不是'随机化选项',是 plant 本体,第一天就该全额建模" - S1's nominal-then-robust staging falsified by hardware for the second time. The later s1e hardware run at power 1.0 (no kicking, normal force) closed the loop: "0.8 = 旧代拐杖" - the derating had been a crutch for the under-modeled delay, not a real requirement.
Change
Latency modeled as plant from day one of any lineage (measured 1-2 ticks covered, bridge-layer rate limits likewise modeled by default); every gate and smoke evaluation issued under --delay 2.
Outcome
Kicking reproduced, explained, and eliminated in the s1e generation at full scale and full power; the deploy-side crutches (0.7/0.8) retired for the new lineage.
Mechanism
Feedback oscillation onset is a product of loop gain and phase lag; a policy trained below the real delay learns gains that sit past the real stability margin, and any output derating masks it by scaling gain down. Since pipeline delay is deterministic hardware property - not an uncertainty - it belongs in the nominal plant, and every evaluation must include it or the gate predicts a robot that does not exist.
Applies when
- hardware oscillation/kicking that sim only reproduces with added delay
- a policy only runs on hardware at reduced power/scale
- defining what belongs in the nominal plant vs the DR list
“真实链路延迟 ~1~2 拍 … S1.1 只训到 1 拍——超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定。s1d 比 s1c 更脆 = yaw 三层栈学出更高反馈增益,增益越高延迟容忍越低。… 教训入账:S1「先标称后鲁棒」第二次被真机证伪——不可约的管线属性(延迟、限速)不是"随机化选项",是 plant 本体,第一天就该全额建模。”
train/OMNI_V0_SPEC.md § 3. S1.4(真机右脚乱踢事故强制) A DR tail the robot never has is pure cost - stage deterministic plant levels instead of one wide uniform
dr-tail-plant-continuationSet every DR range from the measured deployment distribution and cut tails that hardware cannot produce; when an axis changes the controller's character (delay, major gain regimes), prefer staged deterministic levels with gates over one wide uniform.
Symptom
Two consecutive lineages (s1e, s1f) trained under uniform latency DR (0, 0.06 s = 0-3 frames) both converged to drag-glide gaits - buying survival under heavy delay by giving up swing (3.6 mm) - even though the real pipeline never exceeds ~2 frames.
Context
The account: roughly 1/3 of training quality was spent on the >2 frame tail that hardware never presents ("uniform 尾部 ~1/3 训练质量 花在真机不出现的 >2 帧上"). The deeper reading came from the user: uniform 0-3 frames is not merely tail-heavy - it "把性质不同的控制系统 混进同一 PPO batch" (mixes qualitatively different control systems into one PPO batch); a 0-frame and a 3-frame plant demand different controllers, and one policy trained on the mixture serves neither. The S2 v2 ladder therefore redefined latency "从「随机化参数」重新定 义为 actuator/control plant 的一部分": deterministic FIFO levels, staged 1 frame then 2 frames (lo=hi so fractional interpolation degenerates to exact N frames, synonymous with the harness --delay N), each level gated by the fixed acceptance battery - a plant continuation, not a randomization.
Change
Latency DR replaced by staged deterministic levels covering the measured 1-2 tick reality with no tail; each stage a separate continuation rung with the standard gate and rollback.
Outcome
The s2_lag1 rung showed the clean-signal benefit immediately (survival 20/20, heading 6x recovery) with the swing cost booked honestly (21 -> 12 mm, half-pass, ladder paused for adjudication); the drag-glide attractor from uniform tails did not recur.
Mechanism
DR asks one policy to cover a plant family; when part of the family is fictitious, the policy pays real capability for fictitious robustness, and when family members demand structurally different controllers, gradient averaging produces a compromise controller optimal for none. A measured, discrete plant set matches the actual deployment support and keeps each rung's training signal coherent.
Applies when
- policies converge to degenerate gaits that buy worst-case survival
- a DR range extends well past the measured hardware range
- choosing between wide randomization and a staged ladder on an axis
“两轮实证(s1e/s1f)宽尾延迟 DR 逼出拖地滑行 … uniform 0~3 帧不止尾重,而是把性质不同的 控制系统混进同一 PPO batch;1→2 帧确定性分级 = plant continuation,训练信号干净得多—— latency 从「随机化参数」重新定义为 actuator/control plant 的一部分。”
train/OMNI_V0_SPEC.md § 4. v2 阶梯 (2026-08-07 用户定) A single-signal contact detector lied in both directions - foot height flagged 40% false flight on a walking gait, contact force alone flagged false flight during low-friction slip - so flight became force < 5 N AND sole height > 5 mm
contact-detector-single-signal-liesDefine contact and flight events from two independent signals (force and geometry) in conjunction, validate the detector on a behaviour known not to contain the event before using it as a gate, and match the trainer's threshold when comparing across simulators.
Symptom
The run line needed a flight-fraction gate. The first MuJoCo version, "sole higher than 2 mm", measured 40% flight on a walking policy that never flies. Five weeks later the one-leg gate, using contact force alone, reported support-foot "flight" segments at mu 0.4 for a policy that was not hopping.
Context
During a walking step the toe lifts or the heel strikes with the foot pitched, so the ankle-roll origin rises a few mm while part of the sole still touches - height alone calls that flight. Switching to contact force < 5 N (the same threshold Isaac's contact reward uses) zeroed the false flight on walking. In the one-leg re-test every force-only "flight" segment had a measured sole height of 0.0 mm: the normal force chattered while slip corrections played out on low friction.
Change
Run line: flight = contact force < 5 N ("lift-off must be judged by contact force"). One-leg line (2026-09-16): flight = force < 5 N AND sole height > 5 mm, recorded as the same measurement lesson in the opposite direction; the gate's behavioural meaning was unchanged.
Outcome
With force-based detection the walking policy read 0.0 flight and the run policy's zero flight was confirmed by two plants; with the conjunctive definition the one-leg support-foot gate stopped reporting false hops.
Mechanism
Each signal has its own failure: geometry moves without leaving the ground (foot pitch), and forces drop without leaving the ground (slip chatter); requiring both removes both families of false positives.
Applies when
- writing a flight, lift-off, hop or slip detector for a gate
- a gate reports an event the video does not show
- reusing a detector on a different gait or floor friction
“首版用**足底高度>2mm** 判离地, 在 omni_s1e 走路策略上测出 40% 假腾空 … 改用**接触力 <5N**(与 Isaac feet_contact_number 同源阈值)后 空检归零 (walk 策略 flight_frac 0.0)。**课文: 离地判定必须用接触力, 高度判 会把脚的俯仰当腾空**”
train/README.md § run R1 立项 (2026-08-09): Mac 侧新工具 + 一次空检抓获 A binary reward band on the swing knee had zero gradient everywhere below it, so the one-leg policy parked in an unloaded "fake touchdown" that Isaac's 5 N threshold scored as success and MuJoCo showed as real pressing - a capped constant-gradient ramp, retrained from scratch, passed 40/40
binary-band-reward-fake-touchdownShape approach-to-target rewards as capped ramps with gradient from the starting posture, never as bands or indicators; and compare contact-based terms across simulators, because a policy riding just under a force threshold looks perfect in one and wrong in the other.
Symptom
At iteration 1,000 of the first one-leg run the swing foot never lifted: the policy stood with the "raised" foot resting lightly on the ground. In Isaac the contact-match term paid 96% of full marks; the same policy in MuJoCo pressed that foot on the ground for 450 frames.
Context
The swing-leg goal was "shank folded fully back" (knee 1.5-1.95 rad), rewarded as a binary band: +0.8 inside [1.5, 1.95], zero elsewhere. From knee 0.05 to 1.5 rad the term was flat. Contact is judged at a 5 N force threshold, so a foot carrying less than 5 N counts as lifted. The walk line had hit the same disease with a binary indicator (v4) and fixed it with a capped ramp (knee_swing_amplitude).
Change
swing_knee_fold changed from the binary band to a ramp clamp(|q|/1.5, 0, 1) - a constant gradient capped near 86 deg - and the policy was retrained from scratch (V0r1). After the first real-robot try showed the fold still too low, its weight went 0.8 -> 2.0 (V0.1).
Outcome
V0r1 model_2300 passed the full acceptance 40/40 (swing knee 1.72 rad, about 98.5 deg) and was stamped as oneleg_v0.onnx; the cross-simulator disagreement is recorded as the thing that caught the cheat.
Mechanism
A reward that is flat until the target is reached gives no gradient to approach it, so the policy settles for the nearest state other terms reward - here, a foot that satisfies the contact threshold without lifting; a second simulator with different contact force resolution exposes such threshold-riding.
Applies when
- rewarding a posture target with an in-band / out-of-band indicator
- a contact threshold decides whether a foot counts as lifted
- trainer-side contact terms are near full marks while the video looks wrong
“初版二值带 [1.5,1.95] 在膝 0.05→1.5 全程零梯度,策略停在"卸力虚点地"(Isaac 5N 阈下 contact_match 96% 满分 / MuJoCo 同策略 450 帧实压——跨仿真器互证抓作弊);v4 二值指示同型病,按 knee_swing_amplitude 判例改常数梯度封顶 ramp,从零重训 … **oneleg_v0.onnx = V0r1 model_2300, 40/40 PASS**”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §4 奖励表 swing_knee_fold 行 / §8 核查单 5 Privileged signals (true velocity, foot force, foot height) go to the critic only
observation-honesty-critic-onlyTreat the actor observation vector as a hardware contract: every element must exist on the real robot with realistic noise; privileged simulator truths belong in the critic only.
Symptom
Policies trained on ground-truth base linear velocity work in sim and fail on hardware, where only a drifting IMU and encoders exist - the policy has learned to depend on a signal that does not survive deployment.
Context
Many open-source locomotion stacks feed simulator ground-truth linear velocity to the actor. The reference team refused: the real robot has no ground-truth velocity. Asymmetric actor-critic keeps the training benefit of privileged information without deploying the dependency.
Change
Route ground-truth velocity, foot contact forces, and foot heights to the critic only; the actor observes exclusively signals that exist on hardware (IMU-derived quantities, encoders, commands, previous actions).
Outcome
Recorded as adopted doctrine in Lucen's experience log; the trained actor's input contract matches what the real robot can actually produce.
Mechanism
The critic is discarded at deployment, so it may consume any privileged state to reduce value-estimation variance; the actor's observation set is a deployment contract - anything in it that hardware cannot supply (or supplies with different noise/drift) becomes a train/deploy distribution shift the policy was never trained to handle.
Applies when
- designing actor/critic observation spaces
- reviewing a config where the actor sees base_lin_vel or contact forces
- sim policy is strong but real robot drifts, oscillates, or falls without obvious actuator cause
“很多开源代码库把真值线速度喂给策略,Asimov 团队没有,因为真机上没有真值速度,只有会漂的 IMU 和编码器;用完美速度训练出来的策略会依赖它,然后在硬件上失效。真值速度、足底力、足高统统只给 critic”
Experience.md § 观测空间的诚实性 (line 7) The first real-robot get-up was "very violent, kicking on the floor, dangerous" - a sim-perfect policy with no reason to be slow, unbounded absolute targets, no domain randomization and a rate limiter that filtered nothing; the task was restated as "safe, slow, transferable"
first-real-get-up-violent-stage-one-policyDo not put a get-up policy on hardware until its action is bounded (hard bound or state-anchored targets), smoothed, randomized and tested at the real pipeline's latency, and say explicitly that the task is "safe, slow and transferable" - a simulation-perfect policy optimizes only "gets up".
Symptom
On 2026-08-09 the user ran a V0-lineage recovery policy on the real robot and stopped it: very violent, kicking on the floor, dangerous. The planned next rung (a heavier torque_headroom) was never started.
Context
The spec had pre-registered that R0/R1 products stay in simulation and that the real-robot precondition was the R3 smoothing rungs plus a bridge-slew check plus a hanging protocol; the robustness (DR) rungs had not run. In simulation the policy passed 100% with a get-up of about a second. Which ONNX, which gain profile and whether a torque/joint log existed were left "to be recorded later" and never were.
Change
The V0 ladder was stopped at its best product (R3.1, sim only) and a re-rooting proposal was put to the user. The spec's four-layer account: style (the reward pays for standing early and nothing pays for slowness - HumanUP's "Stage I" get-up, "fast but unsafe ... infeasible for real-world deployment"); impact (full-range absolute targets with no hard bound, raw |a| up to 4.77, action saturation 100%, a single-step change of 0.306 saturating hip_pitch); transfer (zero DR, friction pinned at 1.0, the learned leg bracing); link (the bridge's RL slew equals vel_limit, 0.2-0.66 rad per step, while the real pipeline has 1-2 steps of time-varying latency and acceptance ran at delay 0).
Outcome
The line was re-rooted twice (training-side rate limit, then the beta-anchored action space) and gained a hang protocol before the next real attempt; on 08-11 a beta-anchored policy produced the line's first real get-up.
Mechanism
A task reward that pays for standing early selects the fastest feasible get-up; with absolute full-range targets every large target jump is a torque impulse bounded only by the clip; zero DR and braced-leg solutions do not transfer; and a limiter set at the velocity limit does nothing at 50 Hz.
Conflicts
The four layers are the spec's reconstruction from simulation probes and the literature; the real run's policy file, gain profile and log were never recorded, so no layer was confirmed against hardware data.
Applies when
- a first hardware trial of a high-effort skill is being scheduled
- sim success is high but the policy saturates actions or torques
- pre-registered hardware preconditions are not all met
“用户真机反馈:**非常猛、地上乱踢、危险**,叫停(R3.3 torque_headroom 加档已选型 weight −0.5→−1.5,未启动)。真机细节(哪个 onnx、什么档、有无 τ/q log)**待补记** … 任务从"能起来"变成 **"安全、慢、可迁移"** … **链路层**:桥层 slew RL 档 = vel_limit(10/20/33 rad/s ≈ 每拍 0.2~0.66 rad), 对 recovery 形同虚设;真机 1~2 拍时变延迟,验收默认 delay 0。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 真机叫停与换根判决(2026-08-09) The latency DR range must cover the measured deployment pipeline - 0-20 ms could not even reach the real 1-2 control steps
latency-dr-covers-measured-pipelineMeasure end-to-end action latency in control steps on your own stack (including cross-process queue boundaries), set the DR range to cover it with margin, and never import a delay count without its control frequency.
Symptom
Action latency was randomized over 0-20 ms (0-1 control step at 50 Hz), but the measured deployment path is 1-2 steps: the deploy process writes the target, an independently running BusWorker picks it up on its NEXT cycle, plus CAN round-trip - the training range could not cover the robot's actual latency at all.
Context
Fix: widen action_latency_s to 0-0.06 (0-3 steps). The external reference's "uniform 6 steps" was explicitly NOT copied - that number depends on his unknown control frequency; locally, a sweep at 0/1/2/3 steps showed walk_v5 survives all with insensitive metrics, so 6 steps "在我们这里没有依据" (has no local basis). The range was set from the measured pipeline with margin, not from a foreign constant.
Change
action_latency_s (0, 0.02) -> (0, 0.06), justified by pipeline analysis (writer/worker cycle boundary + bus time) and bounded by the local latency sweep.
Outcome
The DR band now brackets the true deployment latency; the policy trains against the delay it will actually face instead of a fictional sub-step world.
Mechanism
Latency DR only immunizes against delays inside its support; a range below the physical pipeline guarantees an untrained distribution shift at deployment. The correct range comes from tracing the pipeline's worst case (queueing boundaries + transport), and foreign step-counts are meaningless without the control rate they were measured at.
Applies when
- setting or auditing action-delay randomization
- deployment uses a separate bus/worker process from the policy loop
- importing delay-modeling numbers from other projects
“现行 0~20 ms = 0~1 个 50Hz 控制步, 而实测部署链路是 1~2 步(deploy 写 STATE.target 后, 独立跑的 BusWorker 下一轮才取走下发, 再加 CAN 往返)——现在的区间覆盖不到真机的实际延迟。… 不照抄参考来源的"统一 6 步": 那取决于他的控制频率(未知), 而我们扫过 0/1/2/3 步 … 6 步在我们这里没有依据。”
train/WALK_V7_SPEC.md § ⑤ action_latency_s 0~0.02 → 0~0.06 A frame-history observation under zero DR memorizes the trainer's plant fingerprint - the estimator must see variation to learn estimation
history-obs-needs-plant-variationIf the observation carries history (stacked frames, RNN), keep at least minimal plant variation (gain/latency jitter) on from the first iteration - "nominal first, robust later" is structurally invalid for estimator-bearing contracts.
Symptom
omni_s1 (fresh 215-dim contract with a 5-frame history window, trained with DR fully off): training all green, yet the MuJoCo gate scored 0/20 on all eight doors - falls within 2 s, seven checkpoints, not one transferred.
Context
The history window exists precisely to let the actor implicitly estimate line velocity and actuator dynamics (the actor is denied base_lin_vel by observation honesty). Under a constant plant that implicit estimator has nothing to estimate - it learns the trainer's exact response fingerprint instead, and any other simulator's micro-differences are out-of-distribution: "5 帧窗按设计就是隐式估计器, plant 恒定时它学到 Isaac 精确响应的指纹". The planned "nominal-first-robust-later" staging was declared STRUCTURALLY incompatible with history observations: "估计器要见过变化才学估计, 否则学背诵" (an estimator must see variation to learn estimation, otherwise it learns recitation). Honest confound note kept: this is mixed with "zero DR does not transfer, period" - but both attributions prescribe the same fix, so no control was run.
Change
S1.1: minimum actuator jitter turned on from day one - kp/kd +/-10%, latency 0-1 frame (friction/COM/mass still nominal, no push - those stay for the S2 ladder).
Outcome
Transfer restored: survival 0/20 -> 20/20, speed 19/20, foot distance 20/20 (remaining failures moved to gait quality, a different disease); the staging doctrine was amended - history-carrying contracts never train under a frozen plant.
Mechanism
A recurrent/history channel fits whatever temporal structure minimizes loss; with a deterministic plant the cheapest structure is the plant's own impulse-response signature, yielding features that are simulator-specific rather than physics-general. Plant variation forces the channel to carry state-estimation features that transfer.
Conflicts
Attribution is explicitly confounded with the simpler "zero DR never transfers" reading ("与「零 DR 本身就不迁移」混杂 … 两种归因处方相同, 不做对照") - the source chose not to spend a control run separating them.
Applies when
- adding frame stacking or recurrence to an actor observation
- a nominal-plant policy fails a cross-simulator gate within seconds
- planning DR staging for a new contract
“frame_hist × 零 DR = plant 指纹过拟合——5 帧窗按设计就是隐式估计器, plant 恒定时它学到 Isaac 精确响应的指纹, MuJoCo 的微小差异即 OOD, 2 s 内摔, 七个 checkpoint 无一迁移。「先标称后鲁棒」的分段与历史观测结构性冲突:估计器要见过变化才学估计,否则学背诵。”
train/OMNI_V0_SPEC.md § 3. S1.1 修订记录 ① A policy's gain profile is part of its contract - the one-leg policy needs per-joint gains the default profile lacks, and the manifest refused an evaluation under the default once; the recovery contract's beta was never stamped, a known gap not to repeat
gain-profile-belongs-in-the-stampStamp everything that defines the closed loop a policy was trained in - gains included - into its manifest, and make every consumer refuse a mismatch; a profile field that is not in the stamp is a silent misconfiguration waiting for an operator to forget a flag.
Symptom
A policy trained with hip_roll kp 80 and ankle_roll kp 60 behaves differently, or falls, under the default kp 20/12 profile - and the gain profile is a command-line flag an operator can forget.
Context
The one-leg line added a gain_profile field to the contract so the stamped manifest carries it; the spec's deployment note says the manifest guard blocks rl_default and that it had already bitten once in simulation (an evaluation run without the one-leg profile). The same spec states the general rule - any new profile field must be synced into the manifest builder - and names the counter-example: the recovery line's beta was never put into the manifest. The recovery line itself had decided that its anchored authority is computed from the base rl gains and written into the contract so it cannot drift with the gain flag, and that the older kp x 0.9 profile chosen in the V0 era does not match the beta contract and must not be used.
Change
Gain profile as a contract field checked at load; per-contract gain choices written into the run sheets.
Outcome
Evaluations and hardware runs of the one-leg policy run under rl_oneleg or are refused; the recovery beta gap stayed recorded as known.
Mechanism
A policy is trained against a closed loop whose gains are part of the plant; running it under other gains is an out-of-distribution plant, exactly like a wrong observation scale.
Applies when
- a skill introduces per-joint or skill-specific gains
- deployment gains are chosen by a command-line flag
- adding any new field to a policy profile
“增益档 `--profile rl_oneleg` 必须给 —— manifest 防线会拦 `rl_default`(sim 已咬合一次) … (recovery 的 β 未进 manifest 是已知缺口,不再复制)”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §9b AGX 真机手顺 要点 / §3 契约 A constant-value plant rung passed every binary gate with record scores - and shipped 60% thinner posture margins that hardware exposed
constant-value-dr-overfits-marginRandomize deployment-critical axes over a narrow band spanning the measured real support - never a single value, never a fictitious tail - and report graded margin quantities (tilt margin) next to binary gates, because saturated gates rank thin-margin and thick-margin policies identically.
Symptom
s2_lag1 (trained at constant 1-frame latency) posted the strongest sim gate sheet in history (20/20 everywhere) yet was unstable on hardware, while s1e (trained across the full 0-3 frame band) was the every-run-stable SOTA at the same power.
Context
The sim autopsy (new --delay-jitter harness modeling the BusWorker's time-varying phase drift): 18 runs across constant and time-varying delays ALL survived - time variation alone does not kill - but the tilt-margin ordering reproduced hardware exactly: s1e 7.7-9.1 deg (thickest) < s2_lag1 10.7-15.0 < s1c 16.2-18.2. Attribution: constant-value training permits precise specialization to that one value; s1e's band diversity forced cross-value robustness - "恒定 1 帧训练 vs s1e 的 0~3 帧全带——分布多样性逼出跨值鲁棒,恒定值允许精确 特化" - so the constant-rung policy's margins were ~60% thinner, fine in sim's clean world, pushed over the line by real-world disturbances. Tool lesson booked: "存活门二值饱和后掩盖裕度差" - binary survival gates saturate and hide margin differences; graded margin columns (tilt-max) belong in the report. The synthesis with the opposite failure (wide tails cause drag-glide): the proposed resolution was a NARROW uniform band (0.02, 0.04) covering exactly the real 1-2 ticks - diversity inside the measured support, no tail, no single point. s1e's root selection later leaned on the same property: its full-band latency training "预装" the delay rungs and delivered "全工况稳定裕度" that survived power derating.
Change
DR-on-an-axis design refined to a three-way distinction: no wide fictitious tails (drag), no single constant values (thin margins), but a narrow band spanning the measured real support; acceptance reports gained graded margin columns alongside binary gates.
Outcome
The tilt-margin column entered the standard report; the s1e root (band-trained) carried the C ladder while the constant-value branch was archived with its three contributions credited.
Mechanism
Robustness margins are shaped by the diversity of the training distribution, not just its support: a point-mass distribution lets the optimizer trade margin for on-point performance, while a band forces solutions that keep margin across the band - and binary survival metrics cannot see the difference until the margin is spent on hardware.
Conflicts
The narrow-band (0.02,0.04) resolution was a pending recommendation ("裁决建议(待用户)") at the time of writing; the lineage instead moved root to s1e whose full-band training predated the staged ladder - the deterministic-staging card and this card record the two failure modes the final design must avoid simultaneously.
Applies when
- a rung trained at a fixed plant value aces sim but wobbles on hardware
- binary acceptance gates are all saturated across candidates
- choosing between constant, banded, and wide DR on one axis
“18 跑全活,时变性单独不足以击杀;但 tilt_max 裕度排序完整复现真机:s1e 7.7~9.1°(最厚)< s2_lag1 10.7~15.0 … 恒定 1 帧训练 vs s1e 的 0~3 帧全带——分布多样性逼出跨值鲁棒,恒定值允许精确特化 … 存活门二值饱和后掩盖裕度差(s2_lag1 sim 门 20/20 史上最强却真机不稳)”
train/README.md § s2_lag1 真机不稳 × s1e 稳的 sim 对拍(2026-08-07,时变延迟实验) Push DR helped one lineage and hurt another at the same dose - robustness budget is conserved and gets borrowed, not created
push-dr-conditional-budget-conservationBefore opening a disturbance-DR rung, measure whether the untrained policy already meets the spec; if training it anyway, expect the benefit to be conditional on the lineage's existing DR load, grade the intensity, and audit retained margins - budget spent elsewhere will be borrowed back.
Symptom
The push rung's outcome flipped with the lineage: direct +/-0.6 m/s push failed outright on first attempt (base walking collapsed - kd1.2 scan 0/3 from iter 3300, sim2sim self-falls with pushes OFF - no PASS point existed); staged +/-0.3 then gave the narrow-kd single-working-point lineage real gains (push survival 1/5 -> 4/5) while the SAME dose made the dual-working-point balanced-band lineage WORSE (20-seed survival 18 -> 12/20 plus across-the-board push regression).
Context
The four-ladder verdict ("四梯定案", s2e/s2f at both intensities) named the pattern: "push DR 收益条件性" - the benefit is conditional on how much robustness budget the lineage has already spent. The law candidate: "DR 总预算守恒, 平衡带鲁棒性从抗扰余量借" - total DR budget is conserved; a lineage already covering a wide plant band pays for push tolerance out of its disturbance margin. Both S2 ladders therefore closed at the friction rung, with the decisive numerator: untrained push tolerance already met the 4-6 N*s requirement, so the rung was not needed at all ("⑥ push 不训(收益条件性,免训 ±0.6 已达 标)"). The same accounting later justified the C-before-S2 ordering ("push/μ 两轮已实证 DR 预算有限且会被重分配") and trimmed the second S2 pass to three rungs.
Change
Push removed from the standing ladder; graded intensity retained as the method IF a lineage ever needs push training; "does the untrained policy already meet the disturbance spec" instituted as the first check before opening any disturbance rung.
Outcome
Two rungs (push, ground mu) deleted from the second S2 pass on measured grounds; the ladder's real yield was re-stated honestly as precision, not robustness (speed gate 0 -> 20/20, zero-command drift 0.98 -> 0.06 m, but push 159 -> 125/160).
Mechanism
A fixed-capacity policy allocates representation and margin across the training distribution; adding a disturbance axis to a lineage that already spans a wide plant family forces reallocation - the new tolerance is bought with existing margins. Lineages with narrow plant coverage have free budget, so the identical DR dose lands as gain. Benefit is a property of (dose x lineage state), never of the dose alone.
Applies when
- proposing push/perturbation training on a hardened lineage
- the same DR rung helped one lineage and hurt another
- accounting where a ladder's robustness gains actually came from
“push DR 收益条件性 —— s2e⑥a (单工作点血统 kd 窄带) ±0.3 得抗推 1/5→4/5; s2f⑥ (双工作点平衡带血统) 同档反而 20-seed 存活 18→12/20 且抗推全面倒退。规律候选: DR 总预算守恒, 平衡带鲁棒性从抗扰余量借。两阶梯均以 ⑤ 摩擦级收官 … 抗推 4~6 N·s 免训已达标。”
train/OMNI_V0_SPEC.md § 4. ⑥ push 四梯定案 (2026-08-07) Record-high training reward hid a fully-failing DR subgroup - aggregate metrics average over draws, gates must test per condition
aggregate-metrics-mask-subgroup-failureNever gate on metrics aggregated across DR draws: evaluate at fixed representative conditions (especially the deployment-critical stratum), and if a difficulty axis matters, ramp it on measured per-stratum success rather than sampling the full range from iteration zero.
Symptom
omni_s1e trained under constant-wide latency DR (0, 0.06 s) posted the lineage's highest-ever Isaac reward (129) - while the --delay 2 smoke evaluation showed 3/3 falls from iter 1500 onward, persisting to early stop; the usable checkpoint window shrank to iters 500-1000.
Context
Diagnosis written plainly: "聚合奖励掩盖重延迟尾部子群体失败" - the aggregated reward averages over latency draws, so the majority of light-delay environments can mask the total failure of the heavy-delay tail. The remedy for the training side was a survival-gated ratchet curriculum (survival_gated_latency): the sampling cap starts at 0.02 s and rises +0.01 only when a 4096-reset window's survival (time_out share) reaches >=90%, capped at 0.06, ratchet up-only - "增益与延迟耐受一起长,不升到策略撑不住的 地方" (gain and delay tolerance grow together; never raise past what the policy can hold). The detection side was already in place from the noise-crutch episode: the per-condition smoke curve, not the training reward, is the health readout.
Change
Latency exposure made curriculum-gated on measured subgroup survival instead of uniform-from-zero; per-condition (--delay 2) smoke evaluation kept as the authoritative curve; watcher scoring adjusted (survival weighted 3x) so recovery during hard phases is not early-stopped away.
Outcome
The failure mode was caught by the smoke curve within one generation; the follow-up redesign (deterministic staged latency) superseded the ratchet, but the aggregate-masking lesson held through both.
Mechanism
Expected-return training weights each DR draw by probability, so a subgroup can contribute bounded loss while being catastrophically failed; any scalar averaged over the randomization cannot distinguish "uniformly decent" from "great on easy draws, dead on hard ones". Only conditioning the evaluation on the stratum reveals the split, and curricula should raise difficulty on measured stratum success, not on schedule.
Applies when
- training reward hits records while a fixed-condition eval degrades
- wide DR on an axis where deployment sits at one known value
- designing curricula for difficulty axes (delay, push, terrain)
“常量 latency DR (0,0.06) 从零训被证伪——Isaac reward 129 历代最高,但 --delay 2 冒烟 iter1500 起 3/3 全摔持续到早停(聚合奖励掩盖重延迟尾部子群体失败,可用窗口只剩 500/1000)。… 采样上限 0.02 起步 … ≥90% 才 +0.01s,0.06 封顶,棘轮只升不降。”
train/OMNI_V0_SPEC.md § 3. S1.5(s1e 训练塌方复盘) Deployment power derating damages non-forward axes far more than forward - sweep it in sim before deploying
power-scale-hurts-nonforward-axesTreat deployment power/torque scaling as a plant parameter: evaluate the policy in sim at the exact deployment scale, expect non-dominant axes to degrade first under derating, and either deploy at the training power or train with power randomization.
Symptom
Policies deployed at power-scale 0.8 (a safety derating of commanded torque) looked fine walking forward but were weak at backward and turning, inviting the wrong diagnosis "the skill was not trained well".
Context
Measured repeatedly: on s1e, going 1.0 -> 0.8 cost forward 18% but backward 58%; on C4-ff800, turn tracking was +25%/+40% at pw0.8 vs +75%/+58% at pw1.0, backward 51-52% vs 97-103%, while forward stayed 96-98% at both. Sim evaluation numbers in the plan were all pw1.0, but the robot was being run at 0.8.
Change
Pre-deploy protocol added: sweep the exported policy across power in sim (for PW in 0.8 0.9 1.0: eval_c_matrix --power $PW --seeds 20) and deploy at the first level where both turn directions reach >=50%. For C4 the recommendation was raise the robot to pw1.0 - the sweep showed it nearly free (saturation 47%->33%, left foot-clipping danger zone 25%->6%, cost only tilt 6.7->8.3 deg).
Outcome
Turning "weakness" resolved without any retraining; the sim sweep correctly predicted the real-robot signature at both power levels.
Mechanism
Forward walking is the reward-dominant, torque-cheapest skill with the most margin; backward/turn/sidewalk live closer to the torque envelope, so a uniform torque derating consumes their margin first. Training ran at power 1.0 (the trainer does no power scaling), so deploying at 0.8 is a systematic underactuation the policy never experienced.
Applies when
- deploying with any torque/power derating or safety scale
- secondary skills (backward, turn, lateral) underperform on hardware while forward walking looks fine
- choosing the deployment power level for a new policy
“power 衰减对非前进轴的伤害远大于前进轴(s1e:前进 1.0→0.8 掉 18%,后退掉 58%)。转向是非前进轴,0.8 下很可能明显跟不动。”
train/C_LADDER_RUN.md § 3c. A-2 上机前先定部署力度档 / 3p. 二 FK-verify a borrowed reference's structure, then size its amplitude by the reference's job - it pins phase, the policy adds lift
reference-structure-fk-amplitude-divisionWhen borrowing a reference trajectory: verify its structural claim against your own kinematics (an invariant like flat-foot), assign it the phase-pinning job, and size amplitude low enough that the policy contributes the lift - moving toward a proven foreign value in halves, not jumps.
Symptom
walk_v4 had big knee swing (40-46 deg) but only 18-24 mm foot lift - amplitude without hip/knee/ankle phase coordination; later, walk_v5's real-robot swing ballooned to 73.6 deg (sim 55.7) with violent footfalls - amplitude over-driven by the reference.
Context
Structure first: Humanoid-Gym's 1:2:1 hip:knee:ankle reference was verified on the local model before adoption - the ratio exactly satisfies the locally derived flat-foot constraint hip - knee + ankle = 0, FK-tested at multiple amplitudes with sole pitch 0.00 deg throughout. Amplitude second, and here the first reasoning failed honestly: FK said shorter legs need LARGER reference scale (0.30 for 30 mm lift), and the FK was correct - but the premise was wrong ("FK 没错, 但前提错了"): it assumed foot lift must come from the reference. HighTorque Pi, same scale, uses 0.08 with a 0.02 m foot-height target - proof that lift is added by the policy ON TOP of the reference, whose actual job is pinning the phase relationship. Scale 0.30 made the reference the entire gait: over-constrained and over-driven. The correction went to 0.15, deliberately not Pi's 0.08: "一次只走一半, 留退路" (walk half the distance, keep a retreat).
Change
target_joint_pos_scale 0.30 -> 0.15 as one of v6-minimal's three changes, treating both the footfall force and the lateral kicking (yaw momentum scales with leg swing amplitude).
Outcome
v6 improved landing force 1.72x -> 1.55x, suspended tilt 45.9 -> 23.0 deg, turn-gain asymmetry 70% -> 19%; the later v6-halved-shaping experiment (35 mm -> 4 mm collapse) confirmed the reference still carries the gait's existence on this machine - the division of labor is real but machine-specific.
Mechanism
A joint-space reference plays two separable roles: encoding structure (phase relations that keep the foot flat) and injecting amplitude (energy). Structure transfers across robots and is checkable by FK against an invariant; amplitude is a negotiation with the policy, and over-assigning it to the reference removes the policy's freedom to modulate lift with state.
Applies when
- importing a reference gait / imitation target from another codebase
- reference amplitude reasoning based on leg length alone
- real swing amplitude far exceeds sim's under a strong reference
“FK 没错, 但前提错了。我默认抬脚必须由参考轨迹产生。HighTorque Pi 同尺度机器人 … 用 0.08, 而它 target_feet_height = 0.02 m —— 说明抬脚是策略在参考之上加出来的, 参考只负责钉住髋/膝/踝的相位配合。我们取 0.30 等于让参考本身就是整个步态, 过约束 + 过驱动”
train/WALK_V6_MINIMAL.md § ① target_joint_pos_scale 0.30 → 0.15 Low-friction robustness traced to kd DR bandwidth, not friction training - by digging resolved params across 8 lineages, 3840 cells
kd-bandwidth-mu-law-attributionAttribute capability differences by tabulating every lineage's resolved training params and eliminating zero-variance and non-aligned columns first; never let an eval-side override knob serve as the explanation axis, and never write a mechanism into a law before it survives a targeted test.
Symptom
Lineages differed wildly in low-ground-friction survival, and the intuitive explanation - "some trained ground friction, some didn't" - was about to steer the ladder toward a ground-mu training rung.
Context
The attribution ran as a full parameter-vs-result cross: 8 lineages x 4 eval kd levels x 6 mu levels x 20 seeds = 3840 cells, with each lineage's RESOLVED training params dug out and compared item by item. First kill: all 8 lineages had ground mu pinned at (1.0,1.0) - zero variance - so low-mu differences cannot come from friction training at all. The only training parameter aligned with the mu score was kd DR bandwidth: narrow (<=0.24) lineages scored 19.9/19.5/19.5, wide (>=0.40) scored 17.1/15.2/14.6/14.2/12.8 - the two groups completely non-overlapping. Every rival was excluded item by item (kd center no; kp band no; COM small-beneficial non-driving; friction rung a clean double null 19.5->19.5 and 15.2->14.6; iteration count non-monotonic), and the one clean single-variable causal link confirmed it: the s2e-3 kd surgery (0.7,1.3)->(1.08,1.32) moved the score 17.1->19.5. Counter-proof against "each best at its own operating point": the narrow-band lineage evaluated OUT of band (18.2) still beat the wide-band lineage at its own band center (9.2). Two axes were ordered never to be conflated (the first attribution's own error): training kd bandwidth is a parameter axis / lineage property; the eval-side --kd-scale knob is a plant axis (more damping physically helps on slippery floors for ALL policies) - "plant 轴只能当部署缓解,不能当 归因". A tempting mechanism story ("drag vs step attractor") was tested and falsified, and explicitly kept OUT of the law: "机制未定, 不入定律".
Change
The planned ground-mu training rung was recommended closed ("建议 不开") in favor of a kd band-narrowing rung (0.8,1.2)->(0.9,1.1) centered on the deployed value - with a pre-registered risk that the law demands "bandwidth = measured dispersion" and the real robot's kd dispersion was not yet measured; if it exceeds +/-10%, narrowing sacrifices real coverage and the rung must yield.
Outcome
A whole training rung was deleted from the ladder by attribution alone (the second S2 pass dropped mu and push, 5 rungs -> 3); floor material became a deployment-selection input (mu <~0.6 -> deploy the kd1.2 gain profile) rather than a training target.
Mechanism
Cross-lineage performance differences must be attributed over the actual training-parameter table, not over eval knobs or plausible stories: eval knobs act on the plant for every policy (a physical effect), while lineage properties come only from training-time parameters. Zero-variance columns are free eliminations, and one clean single-variable rung is worth more than any correlation.
Applies when
- explaining why lineages differ on a robustness axis
- an eval-side knob (gain scale, power) changes results and invites misattribution
- deciding whether to open a DR rung for an axis never actually varied in training
“8 血统地面 μ 训练带全部钉 (1.0,1.0) 零方差,低 μ 差异与「训没训地面摩擦」无关,是 kd DR 带宽的副产物 … 宽 ≤0.24 → 19.9/19.5/19.5;宽 ≥0.40 → 17.1/15.2/14.6/14.2/12.8, 两组完全不重叠。… 训练 kd 带宽 = 参数轴/血统属性;评测部署 --kd-scale = plant 轴 … plant 轴只能当部署缓解, 不能当归因。… 机制未定, 不入定律。”
train/OMNI_V0_SPEC.md § 4. 地面 μ 鲁棒性 = kd DR 带宽的副产物 (2026-08-08) Before resuming a checkpoint, diff the current cfg against what the checkpoint was trained with
resume-state-dr-audit"One variable per rung" counts variables against what the checkpoint actually experienced: audit the checkpoint's logged training config and align every unintended difference before resuming.
Symptom
Two consecutive rungs (C1 back-mode, C2' forward-turn) failed from the same root with the same full-regression signature despite adding different new modes - so the mode was not the cause.
Context
Both runs resumed s1e-500 with the then-current cfg, which carried PD band (0.8,1.2) plus three DR events (base_com, joint_friction, push_robot) accumulated by later lineages. Verified on the training machine from the source of truth (the run's logged params/env.yaml): s1e-500's actual training state was PD +/-10% (0.9,1.1) and all three DR events None. Resuming it under the new cfg meant eating 4 new plant variables plus a new mode at once - the intended "1 variable" was actually 5. A worse variant (c1_redo from s2e_pd-1400) added push +/-0.3 to a root that had never seen it: near-total collapse within +100 iters.
Change
C2 aligned the cfg to the checkpoint's training state before resuming (PD back to (0.9,1.1), three DR events off) - making the new mode the only true variable. Permanent rule recorded: compare the checkpoint's training-time DR with the current cfg before any resume.
Outcome
C2 trained successfully from the same root that had "failed" twice (wz 20/20 with genuine sign-antisymmetric response by iter 700-800); the A/B falsification ("两个不同模式同签名崩") plus the env.yaml verification closed the attribution.
Mechanism
A resumed policy is instantly evaluated (and its value function trained) under whatever plant distribution the cfg specifies; every DR term the checkpoint never adapted to is a distribution shift applied on day one, compounding with the intended change. Single-variable discipline is therefore a property of (cfg diff) x (checkpoint history), not of the cfg diff alone.
Applies when
- resuming or forking any checkpoint under an evolved config
- a resumed run degrades broadly within the first few hundred iterations
- two different changes from the same root fail with the same signature
“A/B 定谳:两个不同模式同签名崩 → 病因不是模式,是「从 s1e-500 续训」。… s1e-500 训练态 = kp/kd ±10% (0.9,1.1),base_com / joint_friction / push_robot 全 None;而 cfg 里带着 (0.8,1.2) + … 三个 DR —— 从它续训等于一次吃 4 个新 plant 变量 + 新模式 … 永久教训:续训前必须比对 checkpoint 的训练态 DR 与现行 cfg。单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」。”
train/C_LADDER_RUN.md § 3b. 这不是重复实验 —— 前两次的病根已定位并修掉 Prone get-up sat at 0/159 until two gated hinge terms moved the seated feet - first sideways (561 -> 360 mm), then fore-aft (-168 -> +56 mm) - and success went to 158/159 with nothing else changed
prone-dead-end-is-foot-placementWhen a stuck state and a successful state differ geometrically, penalize the discriminating quantity with a gated hinge that is exactly zero in the state the policy actually reaches (measure it - not the nominal), then re-probe: flattening one axis can move the discriminant to another.
Symptom
Prone falls always righted and then sat with the feet splayed wide or tucked behind the hips, from where the policy never stood (0% for four generations).
Context
Four lines of evidence pointed at foot position: the configuration probe (ankles 215 mm apart stood 52.3%, 561 mm apart 0.0%); FK showing the action contract's nominal (a = 0) is itself a 465 mm straddle, so the action_rate and still terms were pulling toward the splits; biomechanics (feet tucked under the body cut peak hip-extension torque 148.8 -> 32.7 N*m, -78%); and HoST's foot-displacement term, which this reward table lacked. The earlier "not a reward hole" reading was corrected to "a gradient hole, not a level hole": at the dead point the heaviest term (upright) was saturated with zero gradient, still paid for not moving, and the one live gradient (base_height) pointed at the thigh-horizontal torque barrier. A prone ROM scan had already ruled out pushing up from prone.
Change
R0.4: feet_spread_excess = clamp(ankle distance - 0.215, 0, inf) x upright gate, weight -2.0, plus a height-decay factor added after measuring that the policy's real standing stance was 406 mm, not the 215 mm nominal (the plain version would have taxed every successful stand 0.38/s). R0.5: the same shape on the fore-aft axis, feet_fore_seated = |fore-aft offset - 0.05| x upright gate x height decay, target +50 mm (the measured natural offset of standing postures). One variable per rung.
Outcome
R0.4: seated ankle distance 561 -> 360 mm, supine/side exactly unchanged, prone 0 -> 1.9%, mid 45.9 -> 62.2%; a probe then showed the discriminant had moved to the fore-aft axis (standing starts +42 to +51 mm, the prone seat -168 mm). R0.5: supine 99.4, prone 99.4, side 100, mid 100%, re-falls 0%; both geometry terms collapsed to ~0 near iteration 13,100 as base_height rose, and the prone fore-aft offset went -168 -> +56 mm - the term's own target, closing the causal chain. The cost, unmeasured at the time: action jitter rose 33% (sum |da|^2 6.82 -> 9.06).
Mechanism
An upright-gated hinge is inert while the robot rolls and exactly zero in the achieved stance, so it adds gradient only inside the stuck basin; a seated robot with its feet behind or outside its COM must make a kinematically unfavourable transition to stand, and moving the feet under the body removes it.
Applies when
- a get-up or transition skill fails from one start category only
- successful and failed episodes differ in a measurable geometric quantity
- a shaping term might tax the posture successful episodes already use
“`recovery_r0_5`,唯一变量 = 追加 `feet_fore_seated`(与 R0.4 同形状,只换测量轴)。 … 对照 R0.4 的 prone(3.1%,360 mm,**−168 mm**):前后偏移从 −168 走到 +56, 正是这一项的目标量,**判别量被消掉后成功率随之到顶** —— 因果链完整。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §20 R0.5(前后向脚位置项):成功率门全过 —— prone 0/159 → 158/159 The walking lines' safety setting, power-scale 0.8, broke the recovery policy's full-range contract - it cut the ends of the joint travel (4/50 could not get up) and left the torque spikes untouched; a kp x 0.9 gain profile inside the trained kp band did the job
power-derating-cuts-full-range-contractA deployment derating knob means something only relative to the action contract: before reusing a line's "safe setting" on a new skill, check what it does to that skill's reachable range and to the term that makes the spikes, prefer a gain change inside the band the policy was randomized over, verify it in simulation, and re-decide when the contract changes.
Symptom
After the violent first real get-up (2026-08-09), the recovery policy needed a gentler setting for its next hardware test, and the walking and omni lines' standard derating - deploying at power-scale 0.8 - was the obvious candidate.
Context
The V0 recovery contract maps actions to absolute targets over the full joint range: a = +/-1 lands exactly on the URDF limits, and standing puts the knee at the clip. Candidates were compared on R3.1 in MuJoCo (5 categories x 10 seeds) on 2026-08-10 before any hardware time was spent.
Change
A new gain profile, rl_kp090 (kp x 0.9, kd unchanged), recorded in robot.yaml as the recovery hardware-test setting, with power-scale 0.8 explicitly banned for recovery.
Outcome
kp x 0.9: 48/50 got up; median torque demand on hip_pitch/knee fell from 120-125% to 100-104% of the deployment limit; leg-leg contact frames 2,152 -> 1,095; the change sits inside the +/-10% kp randomization the policy trained with. power-scale 0.8: 4/50 could not get up, because under the full-range contract it removes the ends of the travel (the deep squat's tucked legs, the straight standing knee), and the torque spikes (kp x error) did not fall at all. When the line moved to the beta-anchored contract, rl_kp090 was declared a V0-era choice that does not fit (beta is calibrated at kp 30) and deployment returned to rl_default; the deploy switch applies power scaling to the walking side only.
Mechanism
A power scale multiplies the action, which under an absolute full-range mapping shrinks the reachable workspace instead of softening the actuator; the spikes come from the proportional term on large errors, which only a gain change reduces - and a gain change inside the trained randomization band stays in distribution.
Conflicts
The undated operator runbook still carries an R3.1 "B comparison" command at power-scale 0.8 beside the rl_default baseline; the sources do not say whether it was written before the ban or was ever run.
Applies when
- reusing a power, torque or action scale from one skill on another
- a policy whose actions map to absolute targets over the full joint range
- choosing a gentler setting for a first or second hardware trial
“kp×0.9 / kd 不动 —— recovery_r3_1 成功 48/50, τ 需求中位 hip_pitch/knee 120~125% -> 100~104% 部署限, 腿-腿接触 2152 -> 1095 帧; ±10% 在训练 kp DR 带内. ⚠️ power-scale 0.8 对 recovery **禁用**: 全 ROM 契约下 0.8 砍的是行程 端点 (深蹲收腿/站直够不到), 实测 4/50 起不来, 且尖峰 (kp·err) 一点不降 —— 它是 walk/omni 的安全档, 不是 recovery 的.”
git:Lucen-recovery@origin/recovery:robot.yaml § gain_profiles 注释: recovery 真机测试安全档 (2026-08-10) / rl_kp090 The fallen-state reset was designed, not sampled from SO(3) - fixed category shares with jitter, a low drop that settles physically, equal left/right shares for mirror augmentation, and a numeric check before training
fallen-pose-reset-distributionBuild a fallen-start distribution from named, physically plausible categories with jitter and a settle phase, keep mirrored categories at equal probability, and check the realized shares and penetration numerically before spending a training run on it.
Symptom
A get-up policy can only learn from the fallen states its resets produce; uniformly random orientations produce ground-penetrating and limit-jammed states the robot can never be in.
Context
R0 reset_root_fallen: supine 30%, prone 30%, side_l 15%, side_r 15%, mid (random axis 50-125 deg) 10%, +/-15 deg jitter, full yaw, dropped from 0.28-0.40 m and left to settle under physics, joints uniform inside the soft limits with a 5% margin plus small random velocities. Random SO(3) was rejected (the advisor agreed). side_l and side_r must have equal probability because mirror augmentation turns a left fall into a right fall. The advisor had proposed supine and prone only for R0; the spec included side and mid because the feasibility accounts showed physical solutions for all of them, and wrote "narrow back to supine+prone" down as the first fallback. With no display on the training box the reset was checked numerically instead of by eye.
Change
Category mix as above; realized shares, settle height and penetration measured over 512 envs before the first run. A fallen-state bank (real falls, settled and stored) was pre-registered for R2.
Outcome
Realized shares 29.3/31.6/16.4/17.8% against the config, settle +0.262 m, final penetration 0/512 (a 0.10 m peak at the write instant, ankle links only, pushed out within 80 ms because the 0.28 m drop floor is shorter than a fully extended leg). R0's failure was a reward basin, not a reset artifact. The fallen-state bank stayed unbuilt through V3.1 (checklist item open); R0.3 later re-sliced the prone share into roll_l/roll_r bands, which is what forced the acceptance distribution to be frozen separately.
Mechanism
A category-structured, physically settled start distribution keeps training on states the robot can actually occupy, and equal mirrored shares keep mirror augmentation a pure doubling of data rather than a bias.
Applies when
- designing reset distributions for get-up, recovery or multi-contact skills
- mirror/symmetry augmentation is on and the task has chiral start states
- no viewport is available to inspect resets on the training machine
“角度 jitter ±15°、yaw 全域、0.28~0.40 m 低空放下由物理沉降,关节软限位内 均匀(留 5% 余量)+ 小随机速度。**不用 random SO(3)**(会采出穿地/极限卡死 等现实不可能状态,顾问同判) … side_l/side_r **概率必须相等**(镜像增强的样本同分布前提)”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §3 R0 任务定义 / §9 核查单 An auto-curriculum ratchet capped out at iter 248 and never engaged - stage difficulty manually or verify engagement
auto-curriculum-engagement-checkPrefer manually staged difficulty with gated transitions; if you use an automatic curriculum, instrument its internal state and alarm when it stops engaging - a saturated curriculum is constant DR wearing a curriculum's name.
Symptom
A curriculum mechanism intended to grow difficulty adaptively (s1f's ratchet) hit its cap at iteration 248 and never bit again - for 96% of the run its effect was equivalent to constant DR, i.e. the curriculum existed in name only.
Context
When external advice suggested graded wz bands (start ±0.15, then ±0.30), the team agreed with grading but explicitly rejected automatic curriculum, citing the s1f episode. The same logic had already been paid for with push grading: ±0.6 failed twice, ±0.3 was feasible - grading matters, but the grade transitions were made by hand at verified checkpoints.
Change
Ladder policy: difficulty staged manually, one band per rung, each transition gated by the acceptance battery; automatic ratchets not used unless their engagement is monitored and demonstrated.
Outcome
Every C-ladder band change (wz ±0.12-0.25 first, wider later) was an explicit, attributable rung; no silent constant-DR-in-disguise runs recurred.
Mechanism
Adaptive curricula couple their own state machine to noisy training metrics; a ratchet that saturates early stops adapting but keeps its name, so the operator believes difficulty is progressing when it is frozen. Manual staging costs more decisions but each decision is observable and reversible.
Applies when
- choosing between auto-curriculum and staged bands for a new skill
- a curriculum's difficulty parameter plateaus early in training
- post-hoc attribution of what difficulty a lineage actually saw
“C2 的 wz 分级(±0.15 → ±0.30,不要一上来 ±0.6)—— 与我们付过学费的 push 分级同形(±0.6 两轮 FAIL,±0.3 才可行)。但必须手动分级,不做自动课程 —— s1f 的课程化棘轮 iter 248 封顶未咬合,96% 时长等价常量 DR”
train/C_LADDER_RUN.md § 1. 采纳 3 条 (C2 的 wz 分级) Raising a command bucket's share does not strengthen its per-state gradient - it only starves the other modes
bucket-share-is-not-a-gradient-leverWhen a skill is not learning, first prove its per-state signal is nonzero (ignore-floor and probe checks); only rebalance sampling shares to fix genuine sample starvation, and account the regression risk to the diluted modes before doing it.
Symptom
Sidewalk was not learning, and the reflex proposal was to give the side bucket a larger share of sampled commands.
Context
The C4-redo3 rung explicitly kept the 20/40/20/20 bucket (stand/forward/turn/side) with the reasoning written out: PPO computes advantages per state, so bucket proportion does not change the per-state gradient of side states; at 4096 envs x 20% x 24 steps the rollout already contained ~19.7k sidewalk states - sample count was not the bottleneck. And the cost side was already measured: cutting forward from 60% to 40% had made vx+0.30 die at +400 in an earlier run - more cuts would only collapse it sooner.
Change
Bucket proportions held constant across the entire C4 redo series; the actual bottlenecks (metric frame bug, reward variance penalty, exploration form) were pursued instead.
Outcome
Sidewalk was eventually fixed with zero bucket changes (feed-forward delivery, +100 iters); forward/turn skills never suffered starvation-induced regressions during the redo series.
Mechanism
Policy-gradient credit is assigned per visited state; oversampling a mode multiplies its states in the batch but not the informativeness of each, so if the per-state gradient is ~0 (behavior unreachable or reward indifferent), N times zero is still zero - while the displaced modes genuinely lose data and regress.
Applies when
- proposing to oversample a failing task/command mode
- a majority mode regresses after share rebalancing
- budgeting env count vs mode share for a multi-skill policy
“比例不动:PPO 逐状态算优势,桶占比不改变单状态梯度;4096 env × 20% × 24 = 每 rollout 已有 1.97 万个侧走状态,样本数不是瓶颈;而 forward 60%→40% 已实测让 f30 在 +400 处死掉,再加码只会更早塌。”
train/C_LADDER_RUN.md § 3i. 桶 20/40/20/20 不动(比例不动) Guessed joint friction was 2.5x low and damping 5x high - measure, then DR around nominal
friction-measured-not-guessedMeasure frictionloss and damping separately (they need different rigs), put the measured value at DR center, and express DR as an additive band around that nominal - a DR range around a guessed value can exclude the real robot entirely.
Symptom
Old MJCF friction values were invented, not measured; when finally measured, every guessed value was wrong by a large factor in some direction.
Context
Joint friction split into Coulomb (frictionloss, tau_c) and viscous (damping b). Measured with the robot hung from a crane (吊机测) while armature was measured no-load; the two measurements are deliberately separated. Old MJCF: frictionloss 0.05, damping 0.1, DR joint_friction range [0, 0.1] "凭空拍的" (made up out of thin air).
Change
Replace guessed values with measured ones - frictionloss: RS06 0.15 / RS02 0.12 / RS00 0.13 N*m (old 0.05, i.e. 2.5x too low); damping: 0.02 N*m*s/rad on all three motor types (old 0.1, i.e. 5x too high). DR reshaped from an absolute made-up range [0, 0.1] to an additive band around measured nominal: joint_friction_add [-0.05, +0.10].
Outcome
"摩擦定稿(与 armature 一起, plant 参数第一次全部来自实测)" - friction frozen as part of the first fully-measured plant; DR now brackets a measured truth instead of spanning an invented interval.
Mechanism
Coulomb friction and viscous damping have opposite behavioral signatures (constant-torque threshold vs velocity-proportional drag); guessing both wrong in opposite directions gives a plant that is simultaneously too easy to start moving and too hard to move fast. DR centered on a wrong nominal makes the policy robust to a family of plants that does not contain the real one.
Applies when
- plant friction/damping values have no measurement provenance
- DR ranges are absolute intervals rather than bands around a nominal
- policy is over- or under-damped on hardware relative to sim
“测关节摩擦。吊机测,而armature应该空机测试。… frictionloss τ_c (N·m) │ 0.15 │ 0.12 │ 0.13 │ 0.05(低 2.5×) … damping b (N·m·s/rad) │ 0.02 │ 0.02 │ 0.02 │ 0.1(高 5×) … DR │ joint_friction_add: [−0.05, +0.10] 叠标称 │ 旧 [0, 0.1] 凭空拍的”
Experience.md § 摩擦定稿表 (lines 12-25) Foot dragging is an attractor, not a low amplitude - and joint damping is the mode switch, adjustable at deploy time
swing-bistability-damping-switchWhen a quality metric is bimodal, stop treating it as an amplitude to be trained up: map the modes against initial conditions and plant parameters, find the parameter that switches basins, apply it first as a deployment lever, and only then bake it into the training distribution (as a plant-family shift, never as an execution-mapping change).
Symptom
s2e_pd-1400's swing height "median 12.1 mm" hid a perfect bimodal distribution: 20 seeds split into a drag mode (2.6-4.9 mm) and a step mode (19.3-24.0 mm) with NOT ONE seed in between - the median sat in the empty gap, and "swing debt -11 mm" really meant "50% probability of falling into the drag attractor".
Context
Two designed experiments closed the mechanism. Test A (nominal plant, 40 seeds): step 42% / drag 58% / middle 0 - at nominal gains, initial conditions alone pick the mode, both modes 100% survivable. Test B (fixed init, kp x kd grid): kd is the mode SWITCH - at kd 1.3 all surviving cells step (13-22 mm), at kd 0.7 nearly all drag (2.7-4.3), only at kd 1.0 does init get a vote; kp >= 1.2 is dangerous (5/6 falls). Global verification at kd x1.3 (20-seed, delay 2): survival 20/20 at ZERO cost, step share 42 -> 80%, swing median 12.1 -> 18.4 mm, slip record low 334, thicker tilt margin - costs: vx 85 -> 78%, saturation +5 pp. A Pareto sweep then priced the knob: step share 42/72/75/88/82/90 across kd 1.00-1.30 with a linear vx tax of -2.3 pp per 0.1 kd - the basin gain is fully collected at kd 1.20 ("1.30 是 over-damping 纯多付税"). Mechanism: low damping leaves a landing micro-oscillation / ground-slide channel the policy can exploit to drag; damping plugs the channel.
Change
Deployment lever adopted: kd-scale 1.20 (conservative 1.15) as the legitimate successor to the power-0.8 crutch ("前者削幅度保稳,后者堵 拖地通道换步态,且不牺牲存活"); training-side prescription: move the DR band to nominal-1.2 x (0.9,1.1) = [1.08,1.32], deleting the [0.7,1.0) drag-teaching zone - a contract-level change requiring digest re-baselining, gated on measuring the real robot's actual kd dispersion first.
Outcome
The kd surgery rung (s2e_kd) delivered basin 8 -> 11/20, slip 405 -> 331, vx 81 -> 85% with no out-of-band fragility (below-band check 20/20) - "拐杖烧进分布的正确姿势", explicitly contrasted with the failed s1g amplitude version: this one changes the plant family the policy has seen, that one changed the execution mapping the policy would have to relearn.
Mechanism
The gait's swing behavior is a bistable dynamical system whose basin boundaries are set by plant parameters; a policy trained across a kd band that includes the drag basin has learned to inhabit it. Shifting the deployed (and then trained) damping moves the system into the step basin without touching the policy - a plant-side fix for what looked like a training deficiency.
Applies when
- a gait quality metric splits into distinct modes across seeds
- deciding between more training and a gain/damping change
- converting a deployment crutch into a training-distribution change
“20-seed 里拖地模式 2.6~4.9mm 与迈步模式 19.3~24.0mm 各半,中间一个不落 … kd 是模式开关——kd1.3 下 6/6 存活格全迈步 … kd0.7 下几乎全拖地 … swing 债的解(至少大半)在部署端阻尼档,不在训练端 … 机理:低阻尼下落脚微振荡/贴地滑给了策略顺势拖行的通道,加阻尼堵之。”
train/README.md § swing 双稳态定性 + kd 部署杠杆 (2026-08-07, 用户设计 Test A/B) Fix the task first, harden the plant second - DR budget spent on a dying task is wasted
task-shaping-before-plant-hardeningFreeze the task/command distribution before spending DR budget on plant robustness; if the task will still change, schedule plant hardening as a final pass and book the interim robustness gap explicitly.
Symptom
Tempting default ordering was to keep the plant-hardened (S2) lineage and teach it new commands; but the S2 plant adaptation had been earned on the straight-walk task, and the new omni tasks (sidewalk, in-place turn) use completely different contact patterns.
Context
The team had direct evidence that DR robustness is a budget that gets reallocated when the data distribution changes ("push/μ 两轮已实证 DR 预算有限且会被重分配") - robustness trained under one task/command distribution does not persist when training continues under another.
Change
Ladder order set to: first C (task shaping - add command modes until the task family is final), then a second S2 pass (plant hardening) on the C product. The plant-robustness gap this creates mid-ladder is accepted and booked explicitly ("此处不欠账" - the debt is assigned to the second S2 pass, not denied).
Outcome
The first S2 pass was not wasted: its laws (kd bandwidth <-> low mu, push need not be trained, ground mu need not be trained, bistability) let the second pass drop from five rungs to three. The C ladder itself ran on the softer plant band without incident.
Mechanism
DR robustness is carried by the policy's visited-state distribution; changing the task changes that distribution, so robustness bought under the old task partially dissolves. Hardening before the task is final means paying for robustness on states that will no longer be visited - "给一个即将不存在的任务花预算" (spending budget on a soon-to-not-exist task).
Applies when
- deciding ordering between skill/command expansion and DR hardening
- a hardened lineage is proposed as the root for a task change
- robustness regressions appear after adding new command modes
“S2 的 plant 适应是为直行步态调的,C4 侧走/C3 原地转是完全不同的接触模式,先硬化再改任务 = 给一个即将不存在的任务花预算(push/μ 两轮已实证 DR 预算有限且会被重分配)。故顺序改为 先 C(任务定型)→ 再 S2(plant 硬化)。”
train/C_LADDER_RUN.md § 0. 决策逻辑 = 短板可不可恢复 (末段) Export every CAD part in the whole-machine frame so URDF rotations are zero and inertia is exact
urdf-shared-origin-exportGenerate the model so that correctness is structural: shared-origin STL export, zero rotations, subtraction-only origins, and an explicit 1e-9 g*mm^2 -> kg*m^2 conversion - never hand-rotate inertia tensors.
Symptom
Hand-assembled URDFs accumulate per-link rotation/origin errors and unit-conversion mistakes in inertia tensors - silent plant corruption that no later calibration can cleanly fix.
Context
Documented CAD -> URDF -> USD procedure from a successful Isaac Lab deployment, kept as the recipe if Lucen regenerates its model.
Change
(1) In CAD, align the whole robot to Z-up, X-forward (Isaac Lab convention) and ground the assembly; (2) export each STL with other parts hidden but the machine's shared origin kept, so all parts share one origin, every URDF rotation is 0, and inertia matrices equal CAD values directly; (3) units: Fusion 360 gives g*mm^2, URDF wants kg*m^2 - multiply by 1e-9; (4) link origin = negative of the joint position; link COM = CAD COM minus joint position; joint origin = difference of the two joint positions; (5) after URDF -> USD import, open the USD separately and set it instanceable before saving.
Outcome
A URDF whose rotations are all zero and whose inertia tensors are CAD-exact, eliminating an entire class of hand-transcription plant errors.
Mechanism
Keeping one shared origin turns every frame transform into a pure translation computable by subtraction, and leaves inertia tensors in the frame CAD already computed them in - no rotation of inertia tensors, the most error-prone manual step, is ever needed.
Applies when
- building or regenerating URDF/MJCF from CAD
- inertia or frame bugs suspected in the plant model
- importing URDF into Isaac Lab / USD
“导出 STL 时隐藏其他零件但导出整机——这样所有零件共享同一原点,URDF 里所有 rotation 全是 0,惯量矩阵直接等于 CAD 值 / 单位:Fusion 360 给 g·mm²,URDF 要 kg·m²,乘 1e-9 / link origin = 该关节坐标取负 … URDF → USD 导入后必须单独打开 USD 设成 instanceable 再存”
Experience.md § URDF 制作流程 (lines 87-92) The shipped checkpoint was chosen by scanning checkpoints on the full gate - neighbours 100 iterations apart failed 1 and 38 cells, late checkpoints degraded - never by taking the last one, and training stopped on signals, not on a schedule
checkpoint-choice-is-a-full-gate-scanChoose a release checkpoint by running the full acceptance battery over a band of checkpoints (including the transfer axis), stop training on measured signals rather than a fixed iteration count, and expect adjacent checkpoints to differ sharply.
Symptom
Gate results moved sharply and non-monotonically between checkpoints of the same run, and the last checkpoint was often not the best.
Context
One-leg V0r1: the 2,000 neighbourhood was best; from 2,500 on the nominal gates degraded (late overtraining); 2,000 itself had one real micro-hop (17.7 mm over 5 frames); 2,300 was all green and shipped. V0r2: failed cells per checkpoint 2,000:19, 2,100:38, 2,200:1, 2,300:3, 2,400:27, 2,500:12, 3,000:18 - 2,200 shipped. The recovery line learned the same from the other side: stopping v2_6 early at a scheduled point left a policy whose re-fall rate had spiked to 9-22% before consolidation healed it ("stop on signals, not on the schedule"), and a continuation's transfer decayed checkpoint by checkpoint while Isaac stayed perfect.
Change
The acceptance rule "scan checkpoints, do not look only at the last one" is written into the one-leg gates (called the S1 discipline); release candidates are chosen from the scan.
Outcome
Both one-leg releases were mid-run checkpoints (2,300 and 2,200) chosen by the full 40-cell battery.
Mechanism
PPO keeps changing the policy after the gates saturate; with no gradient toward the gate's conditions, later checkpoints wander, so gate quality is a noisy function of iteration.
Applies when
- picking which checkpoint of a run to export and stamp
- a run is stopped at a fixed iteration budget
- final-checkpoint results are worse than mid-run smoke tests
“Isaac 侧 S1 纪律: 验收扫 checkpoint,不是只看最后一个。 … 扫描判决: 2000 邻域最优——2500+ 标称面退化(⑤③② 散挂, 晚期过训), 2000 有一例真微跳(L s100 μ1.2, 17.7mm/5帧), 2300 全绿。”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §6 验收门 / §8 核查单 5 与 7