Training Coach
Doctrine
A report may cite any of these as doctrine-N.
doctrine-1Contract freeze and fingerprint disciplineThe policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.
Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).
Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.
doctrine-2Attribution by resolved training params - never eval-override knobsCapability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.
Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).
Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.
doctrine-3PASS gates become constraints; FAIL gates become objectivesOnce a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.
Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.
Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".
doctrine-4One variable per ladder rung - counted against what the checkpoint sawA rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.
Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.
Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.
doctrine-5Pre-register risks, readings, and stop criteria before the ladderBefore a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.
Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).
Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.
doctrine-6Plant parameters are measured, never inventedEvery plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.
Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).
Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.
doctrine-7Sim2sim gate before sim2real - under deployment conditionsEvery checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).
Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).
Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.
doctrine-8Observation honesty - the actor's inputs are a hardware contractThe actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.
Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).
Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.
doctrine-9Reward economics are audited in realized currencyReward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.
Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).
Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.
doctrine-10The zero-cost option must be the desired behaviorFor every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.
Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).
Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.
doctrine-11Measurement discipline: independent referees, signs, distributionsA disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.
Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).
Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.
doctrine-12The deployment pipeline is plantIrreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.
Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).
Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.
doctrine-13DR budget is finite; its distribution is the measured supportRobustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.
Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).
Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.
doctrine-14Gates measure what hardware feels: posture, margins, stripped assistsAcceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.
Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).
Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.
doctrine-15Fork and root selection: recoverability, maturity, frozen rewardsChoose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.
Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).
Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.
doctrine-16Curricula: verified engagement, lineage counters, disease-phase gatingAutomatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.
Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).
Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.
doctrine-17Probe before training: feasibility first, hypotheses in tablesAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.
Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).
Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).
doctrine-18External advice is recomputed locally; values transfer as ratiosEvery external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.
Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).
Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.
doctrine-19Hardware sessions are scripted experiments, not tuning sessionsReal-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.
Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).
Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.
doctrine-20Close questions in writing; restart when the debt is structuralAudited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.
Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).
Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.
doctrine-21Name the quantity in the space it lives inA goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.
Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).
Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.
doctrine-22Continuation needs a live gradient; a release is chosen by a scanContinue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.
Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).
Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.
Experience cards
79 cards matching “cycle-average-tracking-for-gait-quantities”.
An exponential kernel on instantaneous velocity punishes gait oscillation - track the cycle average
cycle-average-tracking-for-gait-quantitiesReward velocity tracking on gait-cycle averages (or filtered values), not instantaneous samples, whenever the desired behavior oscillates at stride frequency; widening the kernel does not fix a variance penalty.
Symptom
Even while the robot genuinely sidewalked (verified after the metric fix), the Isaac-side tracking reward sat on the ignore-floor: true sidewalk scored 0.178 vs 0.189 for ignoring the command - the reward was mildly punishing the desired behavior.
Context
Sidewalking is inherently oscillatory: per-frame vy std was 0.177 while the tracking kernel width was sigma = 0.15, applied to the instantaneous value. A kernel-width scan showed widening sigma 0.15 -> 0.50 still loses (-0.14 -> -0.07): "指数核惩罚的是方差,而侧步天生带方差" (the exponential kernel penalizes variance, and side-stepping inherently carries variance). Modeling with measured parameters: replacing instantaneous vy with the mean over one gait cycle (0.5 s) flips the margin decisively (true sidewalk 1.888 vs ignore 1.281, +0.607), half-cycle is neutral (+0.006), two cycles adds nothing more. Explicitly flagged as extrapolation pending Isaac-side implementation. This also vindicated a previously dismissed external note (sigma too small) - right conclusion, different mechanism than claimed (variance, not gradient).
Change
Proposed fix recorded: change the tracked quantity from instantaneous vy to a one-gait-cycle running average; widening sigma alone rejected by the scan.
Outcome
Diagnosis complete and quantified; the C4 product shipped via feed-forward before the reward change was implemented, so the cycle-average fix remained a verified-by-model, not-yet-trained change.
Mechanism
E[exp(-(v-c)^2/sigma^2)] decreases with Var(v) even when E[v] = c exactly; a gait's phase-locked oscillation guarantees variance at the stride frequency, so instantaneous tracking rewards structurally prefer standing still at the command mean. Averaging over exactly one cycle removes stride-frequency variance while preserving command-following error.
Conflicts
The cycle-average fix itself is model-extrapolated ("⚠️ 这一条是外推,须在 Isaac 侧实装并复量后才能当结论") - the diagnosis is measured, the remedy untested in training at the time of writing.
Applies when
- tracking rewards for lateral/turn/any oscillation-carrying velocity
- a verified behavior scores below the ignore-floor
- choosing sigma for exp-kernel tracking terms
“侧走时 vy 的逐帧摆幅 std = 0.177,而 track_lin_vel_y_exp 核宽 σ = 0.15,且作用在瞬时值上 … 真侧走(均值 66%,振荡 ±0.18)0.178 | 完全无视指令 0.189 … 真侧走的得分比无视指令还低。… σ 从 0.15 放到 0.50,侧走仍然吃亏 … 把跟踪目标从瞬时 vy 换成一个步态周期(0.5 s)的平均 vy:… 1.888 vs 1.281”
train/C_LADDER_RUN.md § 3n. 二/三 Isaac 训练奖励为何一直坐在「无视底分」/ 修法不是放宽 σ Narrowing the speed range to stop high-speed falls entrenched crouch-shuffling - judge gait quality at the speed that demands a gait
low-speed-commands-reward-draggingSet command ranges to include speeds that physically demand the target behavior, and evaluate behavior-quality gates at those speeds; when a restriction is added to suppress a failure, check what new optimum it creates at the remaining commands.
Symptom
After the command range was narrowed to (0.15, 0.35) m/s (to treat walk_v1's 134% overspeed and 8.3 s fall at 0.5), the policy settled into crouched foot-dragging; tracking rose monotonically with speed (63% at cmd 0.2, 76% at 0.3, 87% at 0.45), showing low speeds were where the degenerate gait was optimal.
Context
The narrowing advice was the author's own and is retracted in the file: it treated the symptom (falls at speed) while reinforcing the root cause (at 0.15-0.35 m/s, shuffling in a crouch is globally optimal - the Froude number is so low that even humans would not lift their feet). A zero-cost experiment confirmed the flip side: at cmd 0.5 the same policy met BOTH tracking (81%) and clearance (23.0/23.2 mm) standards.
Change
Speed range widened back toward (0.15, 0.5) - upper bound deliberately slightly above the mechanically feasible ~0.44 m/s so the policy finds the boundary itself; acceptance re-pointed: gait-quality criteria (tracking, clearance) judged at 0.45-0.5 m/s, low speed kept only as a survival check.
Outcome
v4 -> v6 progression under the widened range delivered 87% tracking with 34 mm clearance; the "low command = drag" account was confirmed by the monotone tracking-vs-speed curve.
Mechanism
Command distribution is part of the reward: physics prices gaits per speed, and at very low speed the energetic optimum is no swing phase at all. Restricting training to that regime makes the degenerate gait the correct answer to the posed problem - and grading a gait at a speed that does not require stepping measures nothing.
Applies when
- a gait degenerates after a command-range restriction
- quality metrics improve monotonically toward the range boundary
- writing acceptance criteria for gait quality vs survival
“现在看那个建议可能起了反作用:0.15~0.35 m/s 下蹲着蹭就是全局最优,抬腿反而亏。收窄治的是"高速摔倒"的症状,却强化了拖地的病根。… 验收标准里的 cmd 0.2 本身就是拖地速度(Froude 数极低,人在那个速度下也不抬脚)。accept_v2 应把速度跟踪与 clearance 的判定点改到 0.45~0.5 m/s”
train/WALK_DIAGNOSIS.md § ② 放宽速度区间 / ① 零成本实验 Exponential tracking kernels go flat exactly when the error is largest - pair them with an L2 term for the far field
exp-kernel-needs-l2-far-fieldNever let an exp/Gaussian kernel be the only tracking pressure on a quantity that can drift far from target: pair it with an unbounded (L2) term sized as the "don't diverge" floor, and check which frame the kernel reads.
Symptom
With only an exp-type yaw tracking term (exp(-err/std^2), std 0.25), a robot whose heading had drifted badly received almost no corrective gradient: at error 0.6 rad/s the term evaluates to exp(-0.36/0.0625) = 0.003 - near zero AND flat.
Context
The exp kernel is excellent for fine tracking near zero error but its gradient vanishes at large error - precisely when correction matters most. Fix: add track_ang_vel_z_err_l2 (-0.5), a plain quadratic on the same quantity: "exp 管精细跟踪、L2 管'别发散', 互补". Both terms deliberately read WORLD-frame wz (matching the exp term's source), because this torso sways enough that body-frame wz means are systematically off (measured -0.039 while actually turning +0.152). The same far-field-gradient argument reappears in the v8 risk list: frozen joints could not climb back because their huge error put them on the exp plateau ("远端梯度消失是冻结自锁的帮凶").
Change
Added the L2 companion term at -0.5 alongside the existing exp term (a term that had been in an earlier draft and was lost in a rewrite - itself worth noticing).
Outcome
Corrective pressure restored across the whole error range; the exp+L2 pairing became the house pattern for tracking terms.
Mechanism
d/de[exp(-e^2/s^2)] -> 0 as e grows: the kernel saturates and cannot distinguish bad from terrible. A quadratic's gradient grows with error, covering the far field; summing the two yields monotone corrective pressure with fine shaping near the target.
Applies when
- tracking rewards use exp/Gaussian kernels alone
- a drifted or frozen state fails to recover during training
- designing tracking terms for quantities with large transient errors
“exp 在误差大时梯度趋零, 恰好在最需要纠正的时候失灵。… 误差 0.6 → exp(-0.36/0.0625) = 0.003, 接近零且平坦。… exp 管精细跟踪、L2 管"别发散", 互补。”
train/WALK_V7_SPEC.md § ② track_ang_vel_z_err_l2 −0.5 —— 补 exp 的梯度洞 Slowing the gait clock at deployment is out-of-distribution and backfires - lower the commanded speed instead, or train the knob
cycle-time-override-is-oodAny deployment override must correspond to a dimension the policy was trained to handle; to make a parameter field-adjustable, randomize it in training and observe it - otherwise use the levers inside the trained envelope (commands) and leave the knob alone.
Symptom
Real-robot feedback "walks very fast and unstable" suggested slowing the gait; a deploy-side --cycle-time override existed, making "just slow the clock" a one-flag temptation.
Context
A sim sweep of the override on walk_v6 @cmd 0.3 showed monotone degradation away from the trained 0.40 s cycle: at 0.50 s tilt jumped 7.9 -> 13.2 deg and landing force 1.52x -> 2.24x; at 0.80 s (half speed) clearance collapsed to 3 mm - dragging again - with 20 deg tilt. Meanwhile the legitimate lever, lowering the commanded speed with the clock untouched, improved everything monotonically: cmd 0.1 gave 104% tracking, 6.7 deg tilt, minimum slip - the most stable operating point. The file distinguishes the two "slows" explicitly: lower command = smaller steps at the same 2.5 Hz rhythm; a slower rhythm itself requires retraining - randomize cycle_time (e.g. 0.40-0.65 s) during training and expose it as an observation, and only then does --cycle-time become a field-adjustable knob.
Change
Deployment guidance: never ship a cycle-time override the policy was not trained under; respond to "too fast/unstable" with lower commands; schedule clock variability as a training-time (contract-level) change if a field knob is wanted.
Outcome
The sweep quantified the trap before hardware paid for it (dragging and 2.2x landing force at slowed clocks); cmd 0.1 documented as the stable demo point.
Mechanism
The policy is a function fitted around the training distribution; a deploy-side override moves an input (phase rate) to values never seen, so behavior degrades unpredictably - the knob LOOKS like a capability because it exists in the code, but capability lives in the training distribution, not the interface.
Applies when
- a deploy tool exposes overrides (clock, scale, gains) beyond the training distribution
- hardware feels "too fast/aggressive" and a quick knob exists
- deciding between a deploy-side tweak and a retrain
“0.80s | 1.25Hz | 0.165 | 3mm(拖地) | 20.0° … 慢一半直接崩 … 策略按 0.40 训练, 别的周期属分布外。… 降指令速度才是有效杠杆 … cmd 0.1 是最稳的工作点。… 要节奏本身变慢必须重训 —— 训练期把 cycle_time 随机化(如 0.40~0.65s)并作为观测的一维, 部署时 --cycle-time 就成了现场可调的旋钮。”
train/WALK_DIAGNOSIS.md § 2026-08-01 追加: 调慢步态时钟(--cycle-time)在仿真里是反效果 A reward on a quantity the actor cannot observe teaches "produce less of it", never "correct it" - closed-loop correction needs an outer loop
reward-observability-limitBefore adding a reward, check the actor can observe (or infer) the quantity: unobservable-error rewards buy only average suppression - route correction tasks to an outer loop whose commands stay in distribution, and do not break a frozen contract to add an observation a deploy-side loop can supply.
Symptom
Heading kept drifting despite world-frame yaw rewards, and a reviewer proposed heading-error rewards - raising the question of what yaw shaping can even teach this actor.
Context
The adopted architectural verdict: the actor's 45-dim base observation cannot see accumulated heading at all - projected_gravity is invariant to rotation about the gravity axis, and omega_z is a rate, not an angle. World-frame yaw-rate rewards are therefore privileged shaping that can only teach "少产生旋转" (generate less rotation), never "偏了以后拉回原线" (pull back to the line after drifting) - the policy cannot represent the error it would need to correct. The S1 gate (<=5 deg / 10 s) demands exactly the former, so the stack is right for its gate; active heading correction is assigned to the deployment outer loop (--heading P-loop converting heading error into in-distribution wz commands) plus small-wz training - and the 215-dim contract is explicitly NOT extended with a heading observation ("契约不加 heading 观测,冻结不动"). The reviewer's companion bias hypothesis was adjudicated with data: drift is bimodal - a basin mechanism decides whether you leave (seeds vary +/-16-46 deg vs -385 to -391 deg), and once out, rotation direction is constant (weight chirality; candidate root: the phase clock always swings left first).
Change
Yaw shaping kept as rate-tracking (three-layer stack); heading correction owned by the deploy outer loop; contract frozen; the "which behaviors need an outer loop" question settled by observability analysis rather than reward tuning.
Outcome
Stopped a contract change and a futile reward direction; drift work split correctly into rate-suppression (trainable) and error correction (outer loop), consistent with the earlier measured 10x drift reduction from the deploy-side loop.
Mechanism
A policy can only condition on its observation sigma-algebra; rewards on functions outside it shift the marginal action distribution (open-loop average effects) but cannot create feedback on the unobserved variable. Whether to add an observation, an outer loop, or accept average-shaping is decided by the task's gate: suppression gates need shaping, correction gates need the variable in some loop's view.
Applies when
- adding rewards on accumulated/世界-frame quantities (heading, position)
- deciding between a new observation, an outer loop, and shaping
- a drift symptom persists across reward-weight changes
“actor 的 45 维基座观测不到累计航向(projected_gravity 对绕重力轴旋转不变,ωz 是速率不是角度)——世界系 yaw 奖励是特权塑形,只能教「少产生旋转」,不能教「偏了以后拉回原线」。… 主动纠偏闭环 = S3 把小 wz 进分布 + deploy --heading 外环 … 215 契约不加 heading 观测,冻结不动。”
train/OMNI_V0_SPEC.md § 3. 评审④判决(2026-08-06,S1.3 开训前) Every power cycle starts with the same read-only pre-flight - read the buses, check the torque limits against 12/17/11, verify the IMU axes, check the ports after any new USB device - and any reassembly re-measures the joint zeros
power-cycle-preflightStart every powered session with a fixed, read-only pre-flight - bus responses, torque limits equal to the simulated ones, IMU axes, device identities - and re-measure joint zeros after any mechanical reassembly before running a policy.
Symptom
Hardware state drifts between sessions in ways no policy can see: a motor that stops answering after a power cycle, a torque limit that differs from the one simulated, an IMU axis flipped, two USB devices swapping identities, a joint zero moved by reassembly.
Context
The runbook's session order before any policy runs: read every motor on both CAN buses without enabling them (the first command after every power cycle); set_torque --check, all twelve motors must read 12/17/11 N*m, and any difference is written back; imu_reader --verify-axes, where the operator tilts the robot forward and to the right and every check must pass before continuing; check_ports after plugging in any new USB device (the IMU and a CAN adapter once collided on USB identity). After re-mounting motors: read the buses, then re-measure the calibration offsets (three repeats, written back) - "skipping it means running everything on the wrong zero". Hanging checklists repeat the torque-limit check (the deploy script also self-checks at start).
Change
A fixed, read-only pre-flight run in the same order every session.
Outcome
The runbook records one earlier hardware check in the same spirit: all 12 motors' implied kp fell within 18.4-22.0 for a commanded 20, inside the kp randomization range used in training.
Mechanism
A policy transfers only if the plant matches the one it was evaluated on; the pre-flight turns silent hardware drift into a failed check before the robot moves.
Applies when
- the first command after powering a robot on
- after swapping adapters, cables or motors
- a policy that worked last session suddenly behaves differently
“python tools/set_torque.py --check # 12 颗应全对 12/17/11, 有 diff 就 --write … 插任何新 USB 设备后都先跑一次 check_ports.py(IMU 和 CANable 的 USB 身份撞过车) … python tools/calib_stance.py --repeat 3 --write # 重标 offset —— 8/9/10 重新装, 机械零位变了”
RL系统/FOLLOW THIS copy 2.md § WALK / STAND 每次开始前 / 换CAN / 装回后必做两件 Borrow reward values from other robots as ratios (to tracking weight, to leg length) - never as absolute numbers
transfer-ratios-not-absolutesWhen importing any numeric from another robot's config or paper, identify its natural normalizer (tracking weight, leg length, sqrt(g*L), body mass) and transfer the dimensionless ratio; sanity-check against a same-scale robot when one exists.
Symptom
Published configs offered tempting absolute values (swing height target 0.06-0.08 m, weight -20) that would have been wrong for a robot with half the leg length and a different tracking weight.
Context
The cross-check normalized before transferring: G1's feet_swing_height weight -20 against tracking +1.0 is a 20x ratio, so with local tracking at 1.5 the equivalent is -30, not -20. G1's 0.06 m target on a ~0.70 m leg scales to ~28 mm on the local 0.325 m leg (Humanoid-Gym converts to ~23 mm), confirming the locally chosen 0.03 m and explicitly rejecting copying 0.05-0.08 absolutes. A same-class robot (Menlo's 16 kg) was used to sanity-check feet_air_time (+0.5 vs the local 2.0, flagged over-high). A nondimensional check of the same kind later validated the sidewalk speed target (v/sqrt(gL) = 0.081 vs hardware-verified 0.101 - 0.104 - inside the envelope, conservative).
Change
All borrowed values converted through ratios (weight/tracking-weight, height/leg-length, dimensionless speed) before entering the config.
Outcome
The scaled values worked (0.03 m target matched both the scaling law and measured 22-23 mm baseline); no cross-robot absolute was ever copied raw.
Mechanism
Reward economies are scale-relative (only ratios to the tracking term matter to the optimum) and kinematic quantities are morphology-relative (clearance scales with leg length, speed with sqrt(g*L)); absolutes encode the source robot's scale, ratios encode the design intent.
Applies when
- copying reward weights/targets from open-source configs or papers
- setting clearance heights, speed targets, or impact thresholds
- comparing your weights to published tables
“G1 的 feet_swing_height 是 tracking 的 20 倍(−20 vs +1.0)。我们 tracking 是 1.5,按同比例应为 −30 … G1 目标 0.06 m / 腿长 ~0.70 m,换算到我们 0.325 m 腿长约 28 mm;Humanoid-Gym 换算约 23 mm。故 target 取 0.03 m 是对的 … 不必抄 0.05~0.08 的绝对值。”
train/WALK_DIAGNOSIS.md § 修正 ②(权重放大) / 修正 ③(目标高度按腿长缩放) Score left and right separately - averages hide chirality breaking that mirror augmentation does not prevent
chirality-scored-separatelyReport every mirrored skill as two numbers with an explicit gap budget; never accept an average, and never assume augmentation guarantees symmetry - measure it per lineage and treat breakage as hard to reverse.
Symptom
Policies developed quantified left/right asymmetry (e.g. C2-700 turned right at 82% but left at 67% - a 15 pp gap; push tolerance 40/40 symmetric on the root vs 17/40 on a deep-trained descendant), and averaged metrics would have reported healthy midpoints.
Context
The repo had policy-level symmetry-breaking evidence strong enough to make separate scoring a battery rule: "左右必须分开打分 … 平均 vy 跟踪会把它掩盖". Notably, chirality broke and never recovered even though mirror augmentation (command-level mirror_prob 0.5) was on the whole time - augmentation reduced but did not prevent asymmetry, and once broken it stayed broken through subsequent rungs. PASS conditions therefore carried explicit symmetry budgets (left/right tracking gap <=10 pp), and sim's predicted asymmetry (700: right faster than left) was flagged for direct real-robot timing confirmation.
Change
Battery rule: every directional skill reports left and right (CW/CCW) as separate rows with a max-gap budget; mirror augmentation treated as mitigation, not proof of symmetry.
Outcome
The 700-vs-A800 asymmetry gap (15 pp vs 7 pp) became a first-class selection criterion; C4 product shipped with a measured 5 pp gap.
Mechanism
Averaging over mirrored conditions cancels antisymmetric error exactly where it matters; and symmetry lost during training is a lineage injury (like plasticity loss) that later rungs do not spontaneously heal, so it must be gated, not assumed.
Applies when
- evaluating turn/sidewalk/push-recovery or any mirrored skill
- relying on mirror/symmetry augmentation
- selecting between checkpoints with similar average scores
“左右必须分开打分(left/right lateral、CW/CCW turn 各自一行)—— 本仓已有 policy-level symmetry breaking 的量化证据,平均 vy 跟踪会把它掩盖。”
train/C_LADDER_RUN.md § 5. 固定验收矩阵 (左右分开打分) A single run's drift direction may be a limit cycle, not a policy bias - check the sign distribution across seeds
multiseed-sign-test-for-driftDistinguish "bias" from "broken symmetry limit cycle" by the sign distribution over many seeds; report drift as (mean, sign split), and never compare single-run drift magnitudes across versions.
Symptom
Net yaw over 15 s appeared to worsen from -41 deg (v2) to -84 deg (v4), inviting the conclusion that the new version drifted more.
Context
The Isaac-side view across 32 environments told a different story: per-env yaw was mixed-sign (20 negative / 12 positive) with mean ~0 - the drift is a limit cycle whose direction depends on initial conditions, not a systematic policy bias. The single MuJoCo run had sampled one draw from that distribution, so its magnitude could not be compared across versions as if it were a property.
Change
Evaluation rule: before classifying drift as systematic, run multiple seeds and examine the sign distribution; single-trajectory drift magnitudes are samples, not properties.
Outcome
The v2-vs-v4 drift "regression" was reclassified as not-established; later drift work (hip_roll l+r bias) used cross-policy, cross-seed evidence instead.
Mechanism
Symmetric dynamical systems can settle into either of two mirrored limit cycles; the selected cycle is decided by noise and initial state. A statistic whose sign is initial-condition-dependent has no meaning as a single sample - only its distribution does.
Applies when
- comparing heading drift or lateral drift across policy versions
- a symmetric-looking behavior shows a consistent direction in one run
- deciding whether to fix "drift" in reward or calibration
“偏航反而变差(−41° → −84°):注意 Isaac 侧 32 env 的逐 env 偏航是正负混合(20/12)、均值 ≈0,说明这是极限环性质(方向随初值)而非策略偏置 —— MuJoCo 单次跑测到的是分布里的一个样本,不能当作系统性偏差。要判断需多种子统计。”
train/WALK_DIAGNOSIS.md § walk_v4 独立验收 读法 (偏航) Fix a too-deep nominal pose before adding any penalties - the default stance defines the basin training starts in
nominal-posture-before-penaltiesBefore tuning penalties on a degenerate gait, audit the nominal pose and height targets against morphology and published ratios; if the default stance encodes the degenerate behavior, fix it first - and recompute dependent quantities (init height) by FK, not by hand.
Symptom
Policy lived in a crouched shuffle; nominal knee angle was 0.5 rad (28.6 deg) - deeper than published configs (Unitree G1 0.3 rad / 17.2 deg, Booster T1 0.4 rad) - so the policy's starting point and its action-space center both sat inside the crouch basin.
Context
Initially ranked "secondary" in the local diagnosis, this was promoted to co-first priority by the cross-check against published reward tables, which states that with nominal knee flexion above ~0.4 rad, fixing the posture must precede adding any penalties ("改这个之前别加任何 惩罚都是白费"). Companion base-height items: walk profile had weakened base_height_l2 to -5.0 (base class -10, field standard -10 to -20, "the second most common cause of death"), and the height target must be the STANDING height (0.384), not the crouch height.
Change
Nominal knee 0.5 -> 0.3 rad with init_base_height recomputed by MuJoCo FK (0.3739 -> 0.3802); base_height_l2 restored to -10 with standing height target; both bundled as first-priority alongside the clearance term.
Outcome
Part of the v5/v6 package that lifted swing height to 34 mm and tracking to 87%; the crouch basin stopped being the default answer.
Mechanism
The nominal pose is the fixed point every regularizer pulls toward and the point where action=0 lands; if that point is itself the degenerate posture, every penalty fights the geometry. Correcting the attractor is prior to shaping the gradient field around it.
Applies when
- policy converges to a crouched or collapsed posture
- nominal joint angles were chosen for stability rather than gait
- base-height reward targets or weights were locally weakened
“研究明确说"nominal 膝屈超过 ~0.4 rad 必须先改,改这个之前别加任何惩罚"。我们是 0.50,超标。… base_height_l2 在 walk profile 里被减到 −5.0(基类是 −10)。研究说这是"第二常见死因"且应 −10 ~ −20。改回 −10。目标高度用站立高 0.384 是对的(研究要求 target 必须是*站立*高度而非蹲姿)。”
train/WALK_DIAGNOSIS.md § 修正 ①(升级优先级) / 修正 ④ Three hardware accounts locked the run design point - and the knee's real speed ceiling is tau_limit/kd, not the firmware limit
feasibility-accounts-lock-design-pointBefore opening a dynamic-gait training line, compute the full account set - tau_limit/kd effective speed ceilings, joint ROM under the intended reference geometry, and thermal RMS at the duty cycle - and let the accounts lock the design point; move only to pre-registered in-table alternates, re-running the accounts first.
Symptom
The run line was believed to require a firmware raise of the RS06 speed limit (10 rad/s) as a hard precondition, and the feasibility script's motor-envelope scan had marked 80/100 mm foot-lift cells "physically feasible".
Context
Three added accounts re-decided everything. (1) Damping tax: in MIT mode tau = kp*(q_des-q) - kd*qd, so sustained rotation is capped at tau_limit/kd = 12/1.5 = 8 rad/s - below the firmware's 10; at peak speeds 6.7-7.9 rad/s the damping term alone eats 10.1-11.9 N*m (84-99% of the torque limit). "提固件 limit_spd 越不过这道税 —— 它是 kd 与限扭的比,不是固件旋钮." (2) Joint ROM: the feasibility script had checked motor envelopes but NOT joint range - the ankle-pitch ROM caps 1:2:1 leg-shortening lift at 62 mm (soft) / 77 mm (hard), so the 80/100 mm "feasible" cells were voided; also firmware-independent. (3) Ankle thermal: duty 0.40 puts ankle RMS at 87% of continuous rating (0.35 -> 93%); long-period big-stride cells hit both ankle torque peak and heat. Verdict: firmware raise DEQUEUED (50 mm design point needs knee 6.7-7.3 < the 8 rad/s effective ceiling < firmware 10); vel_limit stays 10 so sim == robot. The three accounts uniquely lock the design point - 50 mm lift / T 0.60 s / duty 0.40 - "三笔账 唯一锁定,不是调参空间", with pre-registered alternates allowed only inside the table and only after re-running the accounts.
Change
Design point frozen from accounts; hardware precondition reversed by arithmetic rather than by test; reference amplitude (0.84 rad = FK inverse of 50 mm) derived, per-joint action scales sized to the required travel (knee 0.9, hip_pitch 0.6, ankle deliberately NOT amplified - hard limit is adjacent).
Outcome
A firmware work item left the critical path; an infeasible region of the design space was closed before any training; the remaining risk (knee tracking lag from the damping tax) was pre-registered with its own criterion and in-table fallback (duty 0.35) - "这不是'奖励没调好', 是 plant 账".
Mechanism
PD actuators in MIT mode pay kd*velocity out of the same torque budget that tracks position, so the effective speed ceiling is a ratio of configuration constants, invisible to firmware settings; and feasibility is the intersection of ALL constraint families (torque envelope, joint ROM, thermal RMS) - a scan that omits one family certifies impossible cells.
Applies when
- planning running/jumping or any high-rate gait on PD actuators
- a firmware or hardware upgrade is assumed as a training precondition
- a feasibility scan covers motor limits but not ROM or heat
“膝的有效速度顶 = τ_limit/kd = 12/1.5 = 8 rad/s,不是固件的 10。… 提固件 limit_spd 越不过这道税 —— 它是 kd 与限扭的比,不是固件旋钮。… 可行性脚本只查了电机包络没查关节 ROM —— 其 80/100mm 的"物理可行"格作废。… 判决:RS06 提固件对 run v0 不是前置,出队”
train/RUN_V0_SPEC.md § 1. 硬件账判决 / 2. 步态设计点 Every new penalty ships with a pre-registered withdrawal clause - if healthy gait must pay above the cap, the term stands down
calibration-threshold-with-withdrawal-clauseIntroduce every new penalty with: the zero-cost-option audit, a replay-calibrated weight formula (healthy pays a fixed small fraction of tracking), and a pre-registered withdrawal condition - and let the clause fire without argument when the calibration says the term cannot be afforded.
Symptom
Three same-shaped crashes had established a failure archetype: v4's clearance, v8a's landing window (weight off by 58x uncalibrated), and v6a's bare hip_yaw suppression all combined a zero-cost "don't move" option with a fee on any motion - a reverse barrier that pushes policies toward standing still.
Context
The v11 landing-window penalty was therefore introduced under a calibration-threshold protocol: (1) shape chosen with the window tightened (h_gate 0.03 -> 0.02, because 0.03 equaled the clearance target and priced the entire descent); (2) weight from a FORMULA, not judgment: measure the term's raw value on healthy replays (v5/v10b), set w = -(0.10-0.15 x tracking reward) / raw_healthy; (3) withdrawal clause pre-registered: if healthy gait must pay >15% of tracking no matter the tuning, the term is withdrawn to the next version rather than forced in - "不硬上". The companion hip_yaw quieting term ran the same protocol (calibrate on replays, healthy pays <=5%) and was later retired entirely when a structural fix (zero action scale) made its shaping tax unnecessary.
Change
Penalty introduction protocol: shape audit (what is the zero-cost option?), replay-based weight formula, healthy-pay cap with a written stand-down condition - all before training.
Outcome
The landing term was in fact withdrawn under its clause (v12 records "P5 落地窗口罚 已撤 … 维持撤下"), demonstrating the protocol firing as designed instead of the fourth same-type crash.
Mechanism
A penalty's damage mode is mispricing healthy behavior; since the healthy price is measurable in advance on replays, both the weight and the go/no-go decision can be computed rather than discovered by a ruined training run. The withdrawal clause converts "make it work" pressure into a clean deferral.
Applies when
- adding any motion-taxing penalty to a working gait
- a proposed term's weight has no measurement behind it
- a previous same-shaped term crashed training
“权重公式而非拍脑袋:先在 v5/v10b 回放上量 h_gate=0.02 的原始值,w = −(0.10~0.15 × 跟踪奖励) / raw_健康;标定门槛:若健康步态无论如何要付 >15% 跟踪,本项撤下留 v12,不硬上 (v4 clearance/v8a-B/v6a 三次同型翻车的教训:代价为零的"不动"选项 + 一动就收费 = 反向壁垒)。”
train/WALK_V11_SPEC.md § 6. P5 —— 落地窗口罚(三代欠账,标定门槛制) Real robot walked at half the sim clock for two generations - resolved by racing a reward-side and a plant-side evidence line, not by guessing
period-doubling-evidence-raceFor a hardware-only pathology, refuse to guess: pre-register one probe per side of the sim2real boundary (can the reward mechanism change it on hardware? can fitted plant parameters reproduce it in sim?) and let the first positive result direct the next version.
Symptom
The number-one sim2real gap: on hardware v6/v7 stepped at 1.23-1.32 Hz - almost exactly half the 2.50 Hz gait clock they were trained and simulated at; sim never reproduced it, two generations running.
Context
Instead of committing training budget to a guess, v8 pre-registered two mutually controlled evidence lines and kept the clock OUT of the training variables: (a) reward-side - if the v8 saturation fix revives joint_pos_ref (the term that pins the gait to the clock), re-run hardware and see whether frequency returns to 2.5 Hz (hypothesis: v7's frozen actions meant NO reward was pinning the gait to the clock, and the real plant - with armature and friction making high frequencies expensive - slid down to the leg's pendulum natural frequency ~1.1 Hz); (b) plant-side - record suspended joint data (fit_actuator), fit armature/friction, load the fitted values into sim2sim and see whether the 1.25 Hz reproduces IN SIM. Decision rule fixed in advance: "谁先给出阳性结果谁定 v9 的方向 (奖励侧 vs plant 侧)" - whichever line goes positive first sets the next version's direction.
Change
Period-doubling excluded from the v8 change set; both diagnostic lines scheduled in parallel as non-blocking work; frequency reported factually in acceptance with no pass/fail attached ("倍周期是否消失 不设判定,它是 §9 的关键证据").
Outcome
The gap was routed into a decisive-experiment structure rather than a speculative retrain; the plant-side line pointed at exactly the unmodeled armature/friction that were later measured and installed as the plant baseline. Resolution (era-2c full-plant retest): the family had TWO causes - v8's low-speed period-doubling vanished once measured armature+friction were installed (1.30 -> 2.50 Hz, bifurcation-edge machine sensitivity), while v7's stood untouched at 1.20 Hz (saturation-freeze-driven policy property) - both evidence lines paid off, one per case.
Mechanism
A behavior appearing only on hardware has candidate causes on both sides of the sim2real boundary; changing training to fix it tests only one side per expensive cycle. Two cheap parallel probes - one intervening on the reward mechanism, one making sim reproduce the real behavior - localize the cause to a side before any training money is spent, and sim-reproduction of a real pathology is itself the strongest form of plant validation.
Applies when
- a gait pathology appears on hardware but never in any simulator
- deciding whether a sim2real gap is reward-side or plant-side
- tempted to change the gait clock/reward to chase a hardware symptom
“倍周期(真机 1.23~1.32 Hz ≈ 时钟一半,v6/v7 连续两代;sim 从不出现):两条证据线互为对照——(a)… 真机重跑看频率是否回 2.5 Hz(假说:v7 没有任何奖励把步态钉在时钟上,真机 plant 有 armature/摩擦、高频贵,自由滑落到复摆自然频率 ~1.1 Hz);(b)真机吊挂录 fit_actuator.py … 看能否在仿真里复现 1.25 Hz。谁先给出阳性结果谁定 v9 的方向。”
train/WALK_V8_SPEC.md § 9. 平行线 (倍周期) A nonzero response with the same sign for + and - commands is bias, not ability
same-sign-response-is-yaw-biasBefore crediting any directional skill, test both command signs: response must flip sign with the command; a same-signed pair is a bias to subtract, not an ability to report.
Symptom
Root-selection probe showed nonzero wz "tracking percentages" on turn commands, tempting the read that candidates could partially turn.
Context
During C-ladder root selection, s1e-500's measured yaw rate was +0.084 rad/s for cmd +0.3 and +0.093 rad/s for cmd -0.3 - same sign both ways. The same check on the C2 baseline gave wz+0.20 -> -0.13 and wz-0.20 -> +0.12 (again same sign), while the alternative root s2e_pd-1400 gave +0.16 / -0.16 - opposite signs, i.e. a genuine 16% command response.
Change
Reading corrected and written into the execution sheet: percentages on directional commands are meaningless unless the +cmd and -cmd responses have opposite signs; all three candidates were re-classified as "cannot turn, cannot sidewalk - C2/C3/C4 learn from zero". Acceptance criteria thereafter required "tracking >=50% AND left/right opposite-signed".
Outcome
Prevented crediting turn/sidewalk ability that did not exist; the antisymmetry clause became a standing part of every turn and sidewalk PASS condition (C2, C4, C4-redo levels all carry "且左右反号").
Mechanism
A constant yaw (or lateral) bias projects onto any command's sign convention and shows up as fake fractional tracking; only sign-antisymmetry under command reversal distinguishes a feedback response to the command from an open-loop offset.
Applies when
- evaluating turn/sidewalk/any signed-command tracking percentages
- a candidate shows partial tracking on an axis it was never trained on
- writing PASS criteria for a new directional skill
“C2/C3 那些非零的 wz 百分比不是转向能力 —— 转向+ 与 转向− 的实测同号(s1e:cmd +0.3 → +0.084,cmd −0.3 → +0.093 rad/s),那是恒定偏航偏置。… 三个候选都不会转、都不会侧走。”
train/C_LADDER_RUN.md § 0. 读数纠正(重要,别引错) Measure yaw rate by integrating heading, not by averaging body-frame angular velocity - the two differed 15x
heading-integral-not-body-rateFor any secular rate (turn gain, drift), integrate the world-frame angle over the window; never average instantaneous body-frame rates during oscillatory motion - and when code comments warn about a measurement, believe them before re-measuring.
Symptom
Two measurements of the same turn gain disagreed by a factor of ~15: time-averaged body-frame omega_z gave -0.05 while the sim2sim harness's heading-angle integration gave +0.473.
Context
The harness code comment had already documented and predicted the failure: during gait the torso oscillates (body-frame omega_z std up to 0.7); projecting world angular velocity onto a swaying body axis and then averaging biases the estimate systematically - "实测体系均值 −0.04 而实际在以 +0.15 转" (measured body-frame mean -0.04 while actually turning at +0.15). The author's own -0.05 measurement was declared void and the training machine's 1.58/2.45 turn gains confirmed valid.
Change
Measurement doctrine fixed: yaw rate for evaluation = net heading change by integration over the window; instantaneous body-frame rates are unusable for averaged directional statistics during legged gait.
Outcome
Subsequent friction sweeps and turn-gain accounting were all conducted in the heading-integral currency, making cross-simulator comparisons (MuJoCo vs Isaac 1.04/1.02) meaningful.
Mechanism
Averaging a vector quantity expressed in an oscillating frame couples the frame's oscillation into the mean (a rectification bias); the heading integral is computed in the world frame where the gait oscillation integrates to ~zero, leaving the secular component.
Applies when
- measuring turn gain, heading drift, or any secular angular rate
- a body-frame-averaged statistic disagrees with trajectory-level truth
- writing evaluation code for oscillating platforms
“我用体坐标系 ωz 的时间均值测,得 −0.05;sim2sim 用航向角积分,得 +0.473。差 15 倍。… 步态中躯干摇晃(体系 ωz std 可达 0.7),把世界角速度投到摇摆的体轴上再取均值会系统性偏掉 … 结论:偏航率必须用航向积分,体系瞬时角速度取均值不可用。”
train/WALK_DIAGNOSIS.md § ③ 转向增益 —— 我的测法是错的,训练机的 1.58/2.45 成立 Torque caps cannot soften footfalls - impact is falling-mass momentum, only the reward can treat it
landing-impact-not-fixed-by-torque-capsClassify each hardware symptom by the physics that sets it: quantities fixed by ballistic momentum at contact must be treated through the policy's trajectory (reward terms on approach velocity/force), never through actuator caps - and size such penalty weights against your own tracking reward, not a lighter robot's.
Symptom
Footfalls slammed at 1.78x body weight in sim baseline (human walking: 1.2-1.5x); the tempting hardware-side fix was cutting actuator torque limits.
Context
Measured directly: scaling torque limits from x1.0 down to x0.4 left peak landing force essentially unchanged (1.75 -> 1.78x body weight) - the impact force comes from the momentum of the falling mass at touchdown, not from motor effort. The fix has to change the trajectory, i.e. the policy, i.e. the reward: feet_contact_forces penalty above a threshold of 113 N (= 1.2x the 9.58 kg robot's weight), clipped, weight -0.005. The weight was sized locally, not copied: the reference robot's -0.001 would amount to 0.9% of tracking reward on this robot ("策略不会理它" - the policy would ignore it); -0.005 gives 4.4%.
Change
Added threshold-type contact-force penalty (-0.005, threshold 1.2x body weight) as one of v6-minimal's three changes; hardware torque cuts explicitly rejected as a footfall treatment.
Outcome
Landing force 1.72x -> 1.55x by v6 (target <1.5x, missed by 3% - progress booked honestly); the torque-cap dead end was documented so it would not be retried.
Mechanism
At touchdown the ground stops a ballistic mass; the impulse is set by approach velocity and effective inertia, which motors can no longer influence in the final instant. Only earlier trajectory choices (approach velocity, timing) reduce it - and those are selected by the reward, not by actuator limits.
Applies when
- footfall impact or landing noise on hardware
- proposals to derate torque as a softness fix
- importing contact-force penalty weights from another robot
“⚠️ 硬件限扭降不了落脚力 —— 砸地力来自下落质量的动量: 实测 tau ×1.0→×0.4, 落脚力 1.75→1.78× 体重纹丝不动。只有这条奖励能治。… ⚠️ 权重不能用 Pi 的 −0.001 —— 实测在我们身上只占跟踪奖励的 0.9%, 策略不会理它 (Pi 6.94 kg 更轻)。−0.005 给到 4.4%。”
train/WALK_V6_MINIMAL.md § ③ 新增 feet_contact_forces Changing the gait clock silently flipped a hardwired threshold's meaning - write derived constants as expressions
derived-constants-must-track-their-baseBefore changing any base parameter (clock, control rate, scale), enumerate every constant derived from it and every constant that must NOT change; convert derived literals into expressions of the base so the next change cannot silently flip a term's meaning.
Symptom
Slowing the clock 0.40 -> 0.50 s would have silently inverted the feet_air_time threshold's semantics: the 0.25 s threshold was hardwired, so at ct 0.40 the swing window (~0.20 s) sat below it (constant pressure to lengthen strides), while at ct 0.50 the window (~0.25 s) equals it - the term's meaning flips from "push longer" to "neutral" with no code error anywhere.
Context
The clock change audit walked every dependent quantity: most followed automatically (joint_pos_ref / clearance / contact_number cycle_time params, gait_phase observation, deploy/sim2sim/policy_io, export) - wiring confirmed, zero hand edits; the air_time threshold was the one hardwired constant, fixed by preserving the RATIO: 0.25 -> 0.3125 = 0.625 x ct, with the recommendation to commit it as the expression 0.625*ct "一劳永逸" (solved once and forever). The same audit also listed what must NOT follow the clock (50 Hz control rate, physics dt/decimation, 47-dim contract, action_latency absolute seconds, PD/torque limits) - the change's blast radius stated in both directions.
Change
feet_air_time threshold re-expressed as a fraction of cycle_time; auto-following vs must-not-change lists written into the spec for the clock migration.
Outcome
The clock migration (v10, repeated in v11) carried no silent semantic flips; the expression form removed the trap for every future clock change.
Mechanism
Constants derived from a base parameter encode a ratio at their birth; storing the evaluated number severs the dependency, so changing the base leaves stale semantics with no failing test. Expressions preserve the intent; and an explicit both-directions dependency list (follows / must-not-follow) is what makes a base-parameter change reviewable.
Applies when
- changing gait clock, control frequency, or units
- a reward threshold interacts with a phase/window duration
- config audit finds literals that encode ratios
“feet_air_time 阈值 0.25 是写死的,不跟 ct 走——0.40 时摆动窗 ~0.20s<0.25(恒拉长压力),0.50 时摆动窗 ~0.25s≈阈值(语义翻转)。按比例保原压力:0.25 → 0.3125(=0.625×ct;建议直接写成 0.625 * ct 表达式,一劳永逸)。”
train/WALK_V10_SPEC.md § 3. T —— 慢时钟 (训练侧必做一件) Set torque limits per joint from measured gait peaks - a uniform percentage is the wrong shape, and training must use the deployed numbers
torque-limit-shape-by-measured-peaksMeasure per-joint torque peaks in the actual gait and set each limit as measured-peak x margin capped at rating; then propagate the same numbers into training and add an automated deploy-time consistency check - never derate by a uniform percentage, never let training assume torque deployment will not grant.
Symptom
A uniform 50% torque derating (18/8.5/7) had piled safety margin on the joints that never use it while cutting the busiest joint below half its measured demand.
Context
Per-joint gait peaks were measured (walk_v5 at cmd 0.3/0.6): RS06 (hip_pitch/knee) uses 5.5-5.9 N*m = 15-16% of its 36 N*m rating - cutting it to 12 is a free safety win; RS02's ankle_pitch runs at 16.2 N*m = 95% of its 17 N*m rating - "它是速度的硬件瓶颈", no room to cut; RS00 measured 36-44%, capped at 11. The resulting shape 12/17/11 replaced the uniform percentage. Sweeps across several limit sets (rated / 50% / 14-17-11 / 12-17-11) produced identical speed, lift, and landing force - within this range the limits do not shape the gait; what matters is consistency: "关键是训练和硬件必须是同一个数", because the exporter fills effort_limit from tau_limit, and a policy trained at rated 36/17/14 "会假设有三倍力矩可用" while deployed at 12/17/11 (exactly the v5 cross-generation inconsistency later suspected in its wild kicking).
Change
robot.yaml tau_limit set to the measured-shape 12/17/11, firmware written to match, and train/isaac_values.py regenerated so training sees the same limits; the deploy tool self-checks limits against robot.yaml on every run.
Outcome
Free safety margin captured where demand is low, the real bottleneck joint left at rating, and the train/deploy torque worlds unified with an automated consistency check.
Mechanism
Torque demand is grossly unequal across joints in a gait (15% vs 95% of rating here); a uniform percentage misallocates the safety budget by construction. And since the trainer treats effort_limit as a plant truth, any train/deploy mismatch is an invisible plant gap of exactly the mismatch ratio.
Applies when
- choosing safety torque limits for a legged platform
- training-vs-deployment actuator limit audit
- one joint runs near rating while others idle
“曾用统一 50%(18/8.5/7)是错的形状: 把余量堆在用不到的 RS06 上, 却把 ankle_pitch 砍到需求的 52%。… RS02 在 0.6 m/s 已用到额定 95%, 它是速度的硬件瓶颈 … 实测多组限幅 … 完全一致 —— 限幅在这个范围对步态零影响, 关键是训练和硬件必须是同一个数。… 若训练仍按额定 36/17/14, 学出的策略会假设有三倍力矩可用。”
train/WALK_V6_MINIMAL.md § 3. 训练侧必须同步的一件事 Audit which joints your imitation term constrains - a task that needs deviation is fighting the reference
imitation-term-scope-auditList which joints your imitation/deviation terms actually constrain and check the new skill's required motion against that list; for balance-coupled joints deliver references as feed-forward residuals, not absolute-position targets - and never assume "reference = 0" is neutral.
Symptom
Sidewalk would not learn despite a dedicated tracking reward; meanwhile the gait-shaping imitation term (joint_pos_ref) computed its error norm over ALL 12 joints while its reference covered only the 6 sagittal joints - roll/yaw reference was constantly 0.
Context
Two prior generations had shown the forward gait itself was taught by joint_pos_ref, not discovered by PPO (v6 halved the shaping and swing height collapsed 35 mm -> 4 mm). So the reference is load-bearing - but sidewalk requires hip_roll to deviate from nominal, and the all-joints norm punished exactly that deviation: "一边悬赏一边罚过程" (posting a bounty while punishing the process). A follow-up experiment (C4-redo3, free_roll=True releasing the 4 roll joints from the norm) raised the regularization headroom 6x -> 27x yet sidewalk stayed flat and released hip_roll wandered, killing other skills - net negative, withdrawn. A --roll-absolute probe showed the converse failure: pinning roll to a clock-driven absolute trajectory drove tilt 6.9 -> 13.7 deg. Conclusion recorded: absolute-position imitation cannot teach actions that must be superimposed on state feedback.
Change
The audit reframed the problem: neither punishing roll deviation nor freeing roll nor absolute roll tracking works; the reference for a balance-coupled joint must be delivered as feed-forward under the policy's residual control (see feedforward-for-phase-locked-skills).
Outcome
free_roll rung: joint_pos_ref term rose 0.887 -> 1.104 (release confirmed effective) but vy stayed flat; regularization hypothesis eliminated by experiment.
Mechanism
An imitation error norm defines a cage: joints inside it are pulled to the reference in absolute position, so any skill requiring systematic deviation is taxed per step; but joints carrying active balance cannot follow absolute references either, since their correct position depends on state. The scope and the delivery mechanism of the reference are therefore design decisions per joint, not defaults.
Applies when
- adding a skill that moves joints your reference sets to zero/nominal
- an imitation or deviation penalty coexists with a new tracking reward
- considering releasing joints from a shaping term mid-lineage
“前进步态也不是 PPO 自己发现的,是 joint_pos_ref 教出来的(v6 砍半塑形 → 抬脚 35 mm 塌到 4 mm…)。而 ref_joint_offset 原本只写 6 个矢状面关节,roll/yaw 参考恒 0 —— 侧走既没被教,roll 一偏离 nominal 反被 joint_pos_ref 扣分。一边悬赏一边罚过程。”
train/C_LADDER_RUN.md § 3e. 为什么首战 FAIL / 3i. 解锁笼子 Sort external training advice into adopt / already-have / modify / would-trap by recomputing it on your own config
external-advice-audit-against-own-arithmeticNever apply external tuning advice directly: recompute each claim on your own reward table and probe data, classify it adopt / have / modify / trap, and record why - and verify external citations actually exist.
Symptom
External AI/literature advice for the omni ladder arrived plausible-sounding but was written without knowledge of this robot's actual reward table, contract, and history; following it blindly would have broken single-variable discipline and, in one case, made sidewalk unlearnable.
Context
Before the C ladder, every external suggestion was audited: 3 adopted (ellipsoid command sampling; staged wz bands; command-switch acceptance), 3 already present (unified reward; frame history - the frozen 168-dim 5-frame window; per-100-iter acceptance), 2 modified (stand share kept at 20% to avoid a second variable; back share NOT raised because the probe showed backward works untrained 20/20@67%, so oversampling would only crowd out forward), and 1 flagged as a trap: "start vy very small (0.06-0.15)" - on THIS reward table vy was only an L2 tax, so ignoring a vy=0.06 command costs 0.4% of the vx tracking scale, 28-180x cheaper than ignoring forward, with quadratic shrinkage making small commands weaker still. A separate retrieval-reliability note: two search agents returned fabricated verbatim quotes from arXiv PDFs (2 papers, verified fake and discarded); only HTML/abstract/source-verifiable material was used.
Change
Advice classified only after recomputing each claim with local numbers; the "start small" trap was replaced by adding a gated lateral tracking term (the ladder's only true reward surgery) instead of shrinking the command.
Outcome
The adopted items (ellipsoid modes, staged wz, transition acceptance) entered the ladder; the trap was avoided; one external factual error (calling s1g the mainline start - it was falsified 0/20) was caught. Later, one initially-dismissed item (sigma=0.15 too narrow) turned out right for a different reason than claimed - see cycle-average-tracking-for-gait-quantities.
Mechanism
External advice encodes the advisor's reward table and robot, not yours; the transfer-validity test is whether the claim survives recomputation under your own arithmetic (reward margins, probe baselines, contract freeze). Items that survive become experiments; items that don't become documented traps.
Applies when
- incorporating LLM or literature advice into a training plan
- advice conflicts with locally measured baselines
- an external claim depends on reward-table details the advisor cannot know
“其建议 C4「先很小,vy = ±0.06~0.15」—— 在我们这张奖励表下会让侧走学不起来 … 忽略侧走比忽略前进便宜 28~180 倍,且指令越小激励越弱(平方缩放)—— "先很小"在稳定性上对、在梯度上正好把信号缩没了。”
train/C_LADDER_RUN.md § 1. 外部 AI 训练建议的评估(采纳 / 已有 / 要改 / 会踩坑) Gate a new reward term by its command so all old modes score pointwise identical
gate-new-reward-terms-by-commandWhen a reward term must be added mid-lineage, gate it on the condition that defines the new task so every pre-existing situation scores exactly as before - and still watch for value-rescale pathologies inside the new mode.
Symptom
Adding a lateral tracking reward (track_lin_vel_y_exp) ungated would have paid 0-2.0 per step even in modes with cmd_vy = 0 (healthy gait sway of vy ~0.1 already earns 1.28), shifting the whole reward table by a large bias and rescaling the value function - no longer "just adding one mode".
Context
C4 was the C ladder's only true reward surgery. Single-variable discipline required that the change be invisible to every existing mode. The chosen construction: gate_by_cmd=True - the term pays only when |cmd_vy| > 0.02, so for all modes with cmd_vy == 0 the term is pointwise zero, i.e. the reward is pointwise identical to before the change. The same trick appeared earlier in C1: replacing the vy L2 tax with a command-error version that is "对 cmd_vy≡0 逐点同值" (pointwise equal when cmd_vy is 0), explicitly classified as not-a-reward-change.
Change
track_lin_vel_y_exp added with gate_by_cmd=True (weight +2.0, std 0.15); the residual acknowledged honestly - inside the side bucket the values DO change, so the rung still watched the known reward-reshuffle pathology signature (s1c B-arm: scatter -> half-recover -> collapse) as a stop criterion.
Outcome
Old modes provably unaffected (pointwise-equal argument); attribution for any change in old-skill metrics stayed clean through the C4 redo series.
Mechanism
PPO's critic normalizes to the reward scale it sees; an ungated additive term shifts returns in every state and re-scales advantages globally, entangling the new skill with all old ones. Command-gating confines the new term's support to the new mode's state distribution, making "pointwise identical elsewhere" a provable property rather than a hope.
Applies when
- adding a tracking/shaping term for a new command or skill to a lineage that must not regress
- reward change proposed while other skills are still being gated
- reviewing whether a config diff counts as a reward change
“只在 |cmd_vy| > 0.02 时付。不门控的话它对 cmd_vy≡0 的老模式也给 0~2.0 分(健康摇摆 vy≈0.1 → 1.28),等于给整张奖励表加一个大偏置、值函数尺度全变 … 门控后老模式逐点得 0 = 与加项前逐点同值,单变量纪律成立。… 但 side 桶内的值确实变了 —— 这仍是奖励表改版,开级盯 s1c B 臂签名”
train/C_LADDER_RUN.md § 3d. gate_by_cmd=True(重要) Record-high training reward hid a fully-failing DR subgroup - aggregate metrics average over draws, gates must test per condition
aggregate-metrics-mask-subgroup-failureNever gate on metrics aggregated across DR draws: evaluate at fixed representative conditions (especially the deployment-critical stratum), and if a difficulty axis matters, ramp it on measured per-stratum success rather than sampling the full range from iteration zero.
Symptom
omni_s1e trained under constant-wide latency DR (0, 0.06 s) posted the lineage's highest-ever Isaac reward (129) - while the --delay 2 smoke evaluation showed 3/3 falls from iter 1500 onward, persisting to early stop; the usable checkpoint window shrank to iters 500-1000.
Context
Diagnosis written plainly: "聚合奖励掩盖重延迟尾部子群体失败" - the aggregated reward averages over latency draws, so the majority of light-delay environments can mask the total failure of the heavy-delay tail. The remedy for the training side was a survival-gated ratchet curriculum (survival_gated_latency): the sampling cap starts at 0.02 s and rises +0.01 only when a 4096-reset window's survival (time_out share) reaches >=90%, capped at 0.06, ratchet up-only - "增益与延迟耐受一起长,不升到策略撑不住的 地方" (gain and delay tolerance grow together; never raise past what the policy can hold). The detection side was already in place from the noise-crutch episode: the per-condition smoke curve, not the training reward, is the health readout.
Change
Latency exposure made curriculum-gated on measured subgroup survival instead of uniform-from-zero; per-condition (--delay 2) smoke evaluation kept as the authoritative curve; watcher scoring adjusted (survival weighted 3x) so recovery during hard phases is not early-stopped away.
Outcome
The failure mode was caught by the smoke curve within one generation; the follow-up redesign (deterministic staged latency) superseded the ratchet, but the aggregate-masking lesson held through both.
Mechanism
Expected-return training weights each DR draw by probability, so a subgroup can contribute bounded loss while being catastrophically failed; any scalar averaged over the randomization cannot distinguish "uniformly decent" from "great on easy draws, dead on hard ones". Only conditioning the evaluation on the stratum reveals the split, and curricula should raise difficulty on measured stratum success, not on schedule.
Applies when
- training reward hits records while a fixed-condition eval degrades
- wide DR on an axis where deployment sits at one known value
- designing curricula for difficulty axes (delay, push, terrain)
“常量 latency DR (0,0.06) 从零训被证伪——Isaac reward 129 历代最高,但 --delay 2 冒烟 iter1500 起 3/3 全摔持续到早停(聚合奖励掩盖重延迟尾部子群体失败,可用窗口只剩 500/1000)。… 采样上限 0.02 起步 … ≥90% 才 +0.01s,0.06 封顶,棘轮只升不降。”
train/OMNI_V0_SPEC.md § 3. S1.5(s1e 训练塌方复盘) A ratio metric flipped the verdict - spectral share rose while absolute high-frequency energy fell 16%
ratio-metrics-need-absolute-checkNever compare share/percentage/centroid metrics across conditions whose totals differ; pair every ratio with its absolute numerator before issuing a verdict, and log retracted judgments so they are not re-derived.
Symptom
walk_v6 was provisionally judged "more jittery" than v5 because the joint-velocity spectral centroid rose 2.91 -> 3.62 Hz and the >4 Hz energy share rose 12.6% -> 17.8%.
Context
Absolute measures said the opposite: first differences of actions fell 1.54 -> 1.31, second differences 2.55 -> 2.19, and absolute high-frequency energy fell 16% (0.490 -> 0.413). The shares and centroid rose only because low-frequency content fell even more - the denominator shrank. The interim judgment was retracted in writing so it would not be reused.
Change
Metric discipline noted: "占比类指标在总量变化时不能直接比较" - share/ratio metrics are not comparable across conditions when the total changes; verdicts about smoothness must cite absolute energies or difference norms.
Outcome
v6 correctly classified as smoother, not jitterier; the retracted judgment logged under "被推翻的一个中间判断(记下来免得复用)".
Mechanism
A ratio confounds numerator and denominator; any intervention that removes low-frequency content raises every high-frequency share without adding a single joule of jitter. Only absolute quantities support cross-condition comparison when totals move.
Applies when
- comparing smoothness/jitter/spectral metrics across versions
- any percentage-based metric moves after an intervention
- writing an eval report that includes normalized quantities
“我一度说"v6 动作更抖" … 错了: 动作一阶差 1.54→1.31、二阶差 2.55→2.19 都在降 … 谱质心升高只是因为低频成分掉得更多, 绝对高频能量实际下降 16%(0.490→0.413)。占比类指标在总量变化时不能直接比较。”
train/WALK_DIAGNOSIS.md § 过程中被推翻的一个中间判断(记下来免得复用) Cutting swing amplitude 40% raised yaw-momentum demand 53% - the falsified fix is recorded so nobody walks that road again
amplitude-cut-falsified-yaw-fixTest gait fixes against the quantity the ground must actually supply (torque/force rates vs friction ceilings), not against kinematic proxies; record falsified fixes with their mechanism so the search space shrinks permanently.
Symptom
Support-foot yaw slip stayed at ~223 deg (vs 225 deg) after walk_v6 cut joint swing amplitudes by ~40% (hip_pitch 20.6->10.4 deg, knee 31.3->18.8 deg) - the change built on the theory "smaller swing = less yaw momentum to dump into the ground".
Context
Direct measurement inverted the theory: yaw-momentum amplitude ROSE 16% (+/-0.1000 -> +/-0.1159) and its rate of change rose 53% (1.86 -> 2.84 N*m demanded from the ground), pinned exactly at the foot's supply ceiling (2.8-3.3 N*m at mu 0.6-0.7) - so slip could not drop. The extra demand lives in higher harmonics: v6's crisper foot placement (better clearance 34 mm, lower landing force) shortens the momentum- exchange window, concentrating the same exchange into less time. The verdict was written as a closed road: "下一轮不要再走这个方向". An honest residue was also booked: WHY amplitude down but momentum up 16% remained unresolved, with the named next step (per-rigid-body decomposition of H_z, since per-joint RMS sensitivity ignores phase correlations).
Change
The "reduce amplitude to reduce yaw momentum" lever was removed from the planning space; future yaw-slip work redirected toward the supply side (friction) and momentum-rate mechanics.
Outcome
Slip unchanged (225 -> 223 deg); the falsification and its mechanism became a permanent constraint on the fix search space.
Mechanism
Ground yaw torque demand scales with the rate of change of angular momentum, not its amplitude; kinematic amplitude cuts that also sharpen contact timing can raise dH/dt while lowering range. When demand exceeds the friction-limited supply ceiling, slip is set by the ceiling, so demand-side changes below the ceiling do nothing visible.
Applies when
- attacking foot slip or yaw drift via gait shape changes
- a fix targets an amplitude while the constraint is a rate
- documenting a failed intervention after a version comparison
“walk_v6 | ±0.1159 (+16%) | 2.50 Hz | 2.84 N·m (+53%) … 脚的供给上限 2.8~3.3 N·m(μ 0.6~0.7)—— v6 正好顶在天花板上, 所以滑移一点没降。… 结论: "减小摆动幅度以降低偏航动量"这条被 v6 证伪 —— 砍 39% 幅度, 需求反升 53%。下一轮不要再走这个方向。”
train/WALK_DIAGNOSIS.md § ② 未生效的机理: 减小摆动幅度反而让偏航需求上升 A single-signal contact detector lied in both directions - foot height flagged 40% false flight on a walking gait, contact force alone flagged false flight during low-friction slip - so flight became force < 5 N AND sole height > 5 mm
contact-detector-single-signal-liesDefine contact and flight events from two independent signals (force and geometry) in conjunction, validate the detector on a behaviour known not to contain the event before using it as a gate, and match the trainer's threshold when comparing across simulators.
Symptom
The run line needed a flight-fraction gate. The first MuJoCo version, "sole higher than 2 mm", measured 40% flight on a walking policy that never flies. Five weeks later the one-leg gate, using contact force alone, reported support-foot "flight" segments at mu 0.4 for a policy that was not hopping.
Context
During a walking step the toe lifts or the heel strikes with the foot pitched, so the ankle-roll origin rises a few mm while part of the sole still touches - height alone calls that flight. Switching to contact force < 5 N (the same threshold Isaac's contact reward uses) zeroed the false flight on walking. In the one-leg re-test every force-only "flight" segment had a measured sole height of 0.0 mm: the normal force chattered while slip corrections played out on low friction.
Change
Run line: flight = contact force < 5 N ("lift-off must be judged by contact force"). One-leg line (2026-09-16): flight = force < 5 N AND sole height > 5 mm, recorded as the same measurement lesson in the opposite direction; the gate's behavioural meaning was unchanged.
Outcome
With force-based detection the walking policy read 0.0 flight and the run policy's zero flight was confirmed by two plants; with the conjunctive definition the one-leg support-foot gate stopped reporting false hops.
Mechanism
Each signal has its own failure: geometry moves without leaving the ground (foot pitch), and forces drop without leaving the ground (slip chatter); requiring both removes both families of false positives.
Applies when
- writing a flight, lift-off, hop or slip detector for a gate
- a gate reports an event the video does not show
- reusing a detector on a different gait or floor friction
“首版用**足底高度>2mm** 判离地, 在 omni_s1e 走路策略上测出 40% 假腾空 … 改用**接触力 <5N**(与 Isaac feet_contact_number 同源阈值)后 空检归零 (walk 策略 flight_frac 0.0)。**课文: 离地判定必须用接触力, 高度判 会把脚的俯仰当腾空**”
train/README.md § run R1 立项 (2026-08-09): Mac 侧新工具 + 一次空检抓获 Reward fixes come in causal chains - foot height, then landing impact, then foot spacing
reward-chain-foot-height-landing-spacingPlan reward shaping as a chain, not a point fix: when you patch a degenerate gait behavior, pre-register which adjacent behavior the optimizer will exploit next and watch for it.
Symptom
Three problems appeared strictly in sequence: (1) swing feet lifted too low; (2) after fixing that, feet slammed down - "实际比视频里暴力得多" (far more violent in person than on video); (3) after fixing that, feet drifted too close together and collided.
Context
Each reward fix removed one degenerate optimum and exposed the next. The fix for foot spacing (COM lateral randomization +/-5 cm to force leg spread) itself caused base side-to-side sway, requiring a further foot-to-centerline distance penalty. Lucen had just solved its own foot-height problem (19mm -> 40mm swing height) and logged landing impact and foot spacing as the predicted next two problems.
Change
Chain of additions - (1) penalty when swing foot below 5 cm; (2) landing vertical-velocity penalty at touchdown; (3) COM lateral randomization +/-5 cm, then foot-centerline distance penalty to cancel the induced sway.
Outcome
Reference robot progressed through each stage; each individual fix worked and predictably surfaced the successor problem. For Lucen the chain served as a pre-registered roadmap of what breaks next.
Mechanism
Locomotion rewards are coupled through contact dynamics: raising swing height adds potential energy that must go somewhere at touchdown (impact); penalizing impact and forcing robustness to COM shifts changes lateral support strategy (spacing/sway). The optimizer always exploits the cheapest unpenalized channel, so fixing one channel routes the exploit to its neighbor.
Applies when
- adding a foot-height / clearance reward
- feet slam or landing impact grows after a clearance fix
- feet converge toward the centerline or self-collide
- any single-reward fix to a coupled gait behavior
“抬脚太低 → 加惩罚:摆动足低于 5 cm 就扣分 / 加完之后砸脚 → 抬起来了但落地极猛,"实际比视频里暴力得多" → 加落地速度惩罚 … / 两脚太近甚至互撞 → 先试质心横向随机化 ±5 cm … 有效但引发新问题——基座开始左右摇摆 → 再加足-中心线距离惩罚 … 这三条是串联的:每个修复都会暴露下一个问题。”
Experience.md § 三个问题的解法链 (lines 72-77) A constant-value plant rung passed every binary gate with record scores - and shipped 60% thinner posture margins that hardware exposed
constant-value-dr-overfits-marginRandomize deployment-critical axes over a narrow band spanning the measured real support - never a single value, never a fictitious tail - and report graded margin quantities (tilt margin) next to binary gates, because saturated gates rank thin-margin and thick-margin policies identically.
Symptom
s2_lag1 (trained at constant 1-frame latency) posted the strongest sim gate sheet in history (20/20 everywhere) yet was unstable on hardware, while s1e (trained across the full 0-3 frame band) was the every-run-stable SOTA at the same power.
Context
The sim autopsy (new --delay-jitter harness modeling the BusWorker's time-varying phase drift): 18 runs across constant and time-varying delays ALL survived - time variation alone does not kill - but the tilt-margin ordering reproduced hardware exactly: s1e 7.7-9.1 deg (thickest) < s2_lag1 10.7-15.0 < s1c 16.2-18.2. Attribution: constant-value training permits precise specialization to that one value; s1e's band diversity forced cross-value robustness - "恒定 1 帧训练 vs s1e 的 0~3 帧全带——分布多样性逼出跨值鲁棒,恒定值允许精确 特化" - so the constant-rung policy's margins were ~60% thinner, fine in sim's clean world, pushed over the line by real-world disturbances. Tool lesson booked: "存活门二值饱和后掩盖裕度差" - binary survival gates saturate and hide margin differences; graded margin columns (tilt-max) belong in the report. The synthesis with the opposite failure (wide tails cause drag-glide): the proposed resolution was a NARROW uniform band (0.02, 0.04) covering exactly the real 1-2 ticks - diversity inside the measured support, no tail, no single point. s1e's root selection later leaned on the same property: its full-band latency training "预装" the delay rungs and delivered "全工况稳定裕度" that survived power derating.
Change
DR-on-an-axis design refined to a three-way distinction: no wide fictitious tails (drag), no single constant values (thin margins), but a narrow band spanning the measured real support; acceptance reports gained graded margin columns alongside binary gates.
Outcome
The tilt-margin column entered the standard report; the s1e root (band-trained) carried the C ladder while the constant-value branch was archived with its three contributions credited.
Mechanism
Robustness margins are shaped by the diversity of the training distribution, not just its support: a point-mass distribution lets the optimizer trade margin for on-point performance, while a band forces solutions that keep margin across the band - and binary survival metrics cannot see the difference until the margin is spent on hardware.
Conflicts
The narrow-band (0.02,0.04) resolution was a pending recommendation ("裁决建议(待用户)") at the time of writing; the lineage instead moved root to s1e whose full-band training predated the staged ladder - the deterministic-staging card and this card record the two failure modes the final design must avoid simultaneously.
Applies when
- a rung trained at a fixed plant value aces sim but wobbles on hardware
- binary acceptance gates are all saturated across candidates
- choosing between constant, banded, and wide DR on one axis
“18 跑全活,时变性单独不足以击杀;但 tilt_max 裕度排序完整复现真机:s1e 7.7~9.1°(最厚)< s2_lag1 10.7~15.0 … 恒定 1 帧训练 vs s1e 的 0~3 帧全带——分布多样性逼出跨值鲁棒,恒定值允许精确特化 … 存活门二值饱和后掩盖裕度差(s2_lag1 sim 门 20/20 史上最强却真机不稳)”
train/README.md § s2_lag1 真机不稳 × s1e 稳的 sim 对拍(2026-08-07,时变延迟实验) When the eval proxy has a known systematic bias, gate on within-proxy differences, not absolutes
relative-metrics-survive-proxy-biasWhere the evaluator is known-biased, design gates as within-evaluator contrasts (left vs right, A vs B, pre vs post) that cancel common-mode error; reserve absolute thresholds for quantities whose proxy calibration has been checked.
Symptom
MuJoCo systematically overestimated yaw turn gain (1.58/2.45 where Isaac measured 1.04/1.02), and friction alignment recovered only part of the gap (mu 1.0 -> 0.6 pulled it to 1.43/2.27) - so any absolute turn-gain acceptance threshold in MuJoCo would be judging the proxy's bias, not the policy.
Context
The acceptance criterion was rewritten to use only the left/right difference of the turn gain (standard: <20%): both directions pass through the same biased proxy, so the bias largely cancels in the difference while the policy's chirality - the thing being gated - survives. The absolute-gain row was dropped: "转向只判左右差,不判绝对值". A parallel task was still opened to align MuJoCo contact parameters to Isaac's (mu, restitution), since drift proved highly friction-sensitive (straight-line drift -65 deg at mu 1.0 vs +1 deg at mu 0.4) - bias reduction and bias-robust metrics proceeded together.
Change
Gate metric changed from absolute turn gain to left-right gain difference; proxy-alignment work scheduled separately rather than blocking acceptance.
Outcome
Turn acceptance became meaningful across proxy versions (v5 70% -> v6 19% difference measured the real improvement) while the absolute bias question was pursued without holding the ladder hostage.
Mechanism
A systematic multiplicative or additive proxy bias applies to both arms of a mirrored measurement; differencing (or ratioing) mirrored conditions cancels the common-mode bias to first order, leaving the asymmetry signal. Metrics built this way remain valid while the proxy is imperfect - which it always is somewhere.
Applies when
- a sim proxy disagrees with the trainer or hardware in absolute terms
- writing acceptance thresholds for direction-paired skills
- proxy calibration work would otherwise block a ladder
“转向只判左右差,不判绝对值 —— MuJoCo 的偏航增益系统性高估(实测 1.58/2.45 vs Isaac 1.04/1.02),摩擦只能解释一部分(μ 1.0→0.6 仅拉到 1.43/2.27)。差值是相对量。”
train/WALK_V5_SPEC.md § 6. 验收 Decompose the offending quantity by channel first - then penalize the failure event, not the joints
penalize-the-slip-not-the-jointBefore penalizing motion to fix a side effect, measure which channels actually carry the offending quantity; prefer penalties conditioned on the failure event that are exactly zero for healthy behavior - and do not medicate behaviors that measurement shows are not sick.
Symptom
Heading drift with support-foot yaw slip (v5: 212-284 deg accumulated over 15 s); the previous v6 draft had attacked it by penalizing lateral joints (a roll 4.0 / yaw 2.0 "home" group) - which collapsed training into the standing basin.
Context
Before choosing the penalty target, the yaw angular momentum was decomposed by joint group with MuJoCo subtree_angmom weighted by real walking joint velocities: pitch-class joints (hip_pitch + knee) carry 95.3%, hip_roll 3.3%, hip_yaw 1.4%. The failed "home" group had been taxing 2.7/step to manage a 4.7% channel. The replacement, feet_yaw_slip (-0.2, |support-foot yaw rate| while in contact), targets the failure event itself and - decisively - costs a non-slipping gait exactly zero, which "横向回家组做不到". The same rung's do-not-do table applied the complementary principle to foot spacing: measured 196-214 mm, stable, no crossing - "没病不吃药" (no disease, no medicine).
Change
Removed joint-usage penalties for the drift problem; added the event-conditional slip penalty (-0.2, realized tax 0.141/step = 12% of tracking) alongside the existing linear-slip term.
Outcome
Turn-gain left/right difference improved 70% -> 19% and heading 185 -> 60.3 deg by v6 without a standing-basin collapse; the 2.7/step lateral tax never returned.
Mechanism
Penalizing joints taxes every use of a channel including healthy use, and if the channel carries little of the offending quantity the tax buys nothing while pushing the optimum toward immobility. An event-conditional penalty (slip while in contact) prices only the failure, leaving the healthy gait's cost surface untouched - and the channel decomposition tells you in advance whether a joint-side fix can even work.
Applies when
- choosing a penalty target for drift/slip/impact problems
- a proposed penalty taxes joints or motions rather than failure events
- a previous joint-penalty attempt collapsed the gait
“pitch 类 (hip_pitch + knee) 占偏航角动量 95.3% … hip_yaw 1.4% … 压 hip_yaw 是管 1.4% 的通道收 2.7/步 的税 —— 上一轮正是这样把策略推进了站立盆地。滑移项不惩罚走路: 不打滑的步态代价为零, 这是横向"回家"组做不到的。”
train/WALK_V6_MINIMAL.md § ① / ② 新增 feet_yaw_slip Three same-shaped judging errors - task metrics (survival, tracking, displacement) cannot stand in for posture metrics
task-metrics-vs-posture-metricsKeep validated posture-class rows (tilt, per-joint L/R asymmetry, temperature) in every acceptance battery alongside task rows; when operator feel contradicts the gates, suspect the metric class before the operator - and never build a new skill on what is actually an asymmetry defect.
Symptom
The C2 product judged "full pass" on task metrics (A800: turn-gap 7 pp, vx+0.30 19/20) felt WORSE in the operator's hands than the half-pass 700: A800 tilted up to 12.90 deg (700: 6.64), drifted left while standing, showed larger per-joint asymmetries, and ran its hip_rolls 5 degC hotter.
Context
The re-judgment catalogued three same-type metric errors in one campaign: (1) stand judged by SURVIVAL - missed 0.5-1.4 m wandering; (2) stand ranked by DISPLACEMENT - ordering was opposite to real feel (tilt ordering matched); (3) chirality judged by wz-tracking GAP - measured turning symmetry while the robot's actual disease was postural left/right asymmetry, "两个不同的东西,且结论相反". Common pattern named: "我一直用「任务指标」当判据,而真机手感对应的是「姿态 指标」… 任务类指标不能替代它". The fix was already in the data: the per-joint left/right asymmetry table (printed identically by sim2sim and deploy) agreed with hardware in direction on every row - "判据可用、有预测力,我只是没把它写进 PASS 条件". Shipping decision followed the posture read: product reverted to 700 ("又一次「买到 精度、卖掉别的」"), and the C4 root moved to 700 as well, with the sharpest line of the episode: A800's left-drift "like sidewalking" is probably its frontal-plane asymmetry defect, not a capability - "在缺陷上建能力是危险的".
Change
Two posture quantities with demonstrated real-robot predictive power promoted into every PASS battery: tilt-max median and per-joint left/right asymmetry (both sim-computable, deploy-homologous); motor-temperature readout added to session close-out.
Outcome
Deployment flipped to the posture-better checkpoint; the hip_roll temperature table (43-48 degC vs 25-28) confirmed the earlier 90%-of-heat account; the run-line acceptance battery inherited the posture rows from birth ("任务类替代不了姿态类").
Mechanism
Task metrics measure goal attainment under the evaluator's episode definition; posture metrics measure the body state trajectory that operators, motors, and downstream skills actually experience. The two can rank candidates oppositely because task success tolerates postural pathology - so a battery without posture rows is blind to exactly what hardware feel reports first.
Applies when
- hardware feel disagrees with a green acceptance table
- choosing between checkpoints that split task vs posture metrics
- selecting the root for a skill that resembles an existing defect
“共同模式:我一直用「任务指标」(存活 / 跟踪率 / 位移)当判据,而真机手感对应的是「姿态指标」(倾角、逐关节左右不对称)。→ 验收判据里必须有姿态类指标,任务类指标不能替代它。… A800 的「左飘像 side walk」很可能 … 是它更大的额平面不对称的表现 —— 在缺陷上建能力是危险的。”
train/README.md § C2 选点改判 (2026-08-09): 手性判据第三次选错指标 A suspended (no-load) test acquits or convicts the actuator before you blame authority
suspended-test-isolates-actuator-authorityBefore attributing a failure to actuator authority, measure no-load tracking error and steady-state torque fraction; blame authority only if the task fails while the error grows with demanded force - and then fix gains or targets, not training.
Symptom
hip_roll sagged 0.21 rad on the ground and saturation questions loomed over the sidewalk plan - was the roll axis physically too weak (authority), or was something else limiting it?
Context
Before C4, the roll-authority question was settled by measurement triage: suspended test (--suspend, feet off ground) showed hip_roll tracking error 0.0008 rad - actuator acquitted; the entire 0.21 rad ground sag is load-induced. Steady-state torque was 25% of limit - 75% margin remains. Since sidewalk needs lateral force, not exact angles, authority was ruled "not a hard limit", with a pre-registered criterion for when it WOULD become one: sidewalk fails to track AND roll error keeps growing - then the fix is raising hip_roll kp or lowering the vy target, not more training.
Change
Hypothesis "roll authority insufficient" demoted from blocker to a monitored branch with an explicit trigger condition; C4 proceeded.
Outcome
Later open-loop probes confirmed the actuator could produce the behavior (sidewalk feed-forward ran at full amplitude, 5/5 survival), and the eventual C4 failure causes were measurement and reward, never authority.
Mechanism
Suspended vs loaded comparison separates the actuator's closed-loop competence from the load path: tiny no-load tracking error means the motor/controller is fine and any loaded deviation is statics (gravity / stiffness budget, kp trading error for force). Torque-fraction measurement then bounds how much force headroom actually remains.
Applies when
- suspecting an axis is "too weak" for a new skill
- large position sag on a loaded joint
- deciding between hardware fix, gain change, and more training
“吊挂(--suspend)实测 hip_roll 跟踪误差 0.0008 rad → 执行器无罪,地面下垂 0.21 rad 全是负载所致;稳态占限扭 25% → 仍有 75% 扭矩余量。… 判据:若 C4 出现「侧走跟不动且 roll 误差继续变大」,那才是权限账 … 解法是提 hip_roll 的 kp 或降 vy 目标,不是硬训。”
train/C_LADDER_RUN.md § 3d. roll 权限:已部分澄清,不是硬上限 A joint frozen at the action clamp pays zero action_rate forever - penalize pre-clip saturation to make the cheat cost money
saturation-cheating-zero-rate-costWhenever actions are clipped and any smoothness/rate penalty exists, add a pre-clip saturation penalty so living at the clamp costs more than oscillating - and audit for frozen-at-clamp joints (action std ~0, |a| at exactly the clip value) as a standing acceptance row.
Symptom
With action_rate_l2 raised to -0.2, walk_v7's hip_pitch actions froze at exactly +/-1.000 (the clamp), reproduced bit-for-bit on hardware (splits frozen at +/-0.35 rad); the gait-shaping term joint_pos_ref collapsed to 0.026-0.035. The repo had died in the same trap once before (walk_v0: four joints pinned at +/-1.0).
Context
Mechanism: a joint pinned at the clamp has action-rate cost exactly zero and forever zero - under a strong smoothness tax, "push to the clamp and freeze" becomes the dominant optimum. Lowering the weight (-0.2 -> -0.1) only reduces temptation; the frozen state still costs nothing, so the structural fix adds action_saturation = sum(relu( |a_raw| - 0.9)) at weight -1.0, computed on the PRE-clip network output - post-clip, |a|=1.01 and |a|=3 punish identically and the out-of-range gradient dies (v0's old disease: mean |a| 1.71 soaked in saturation). Economics: freezing at |a|=1.0 now pays 0.1/joint/step (two hips = 40% of alive) vs ~0.0004/step for the healthy reference oscillation - the cheat flips from free to ~250x negative. Honest limits were recorded: A1 does not forbid freezing at 0.89 (the anti-freeze pressure must come from the oscillation demand of joint_pos_ref), and the alternative "rate on post-clip target" was rejected as 换汤不换药 - a pinned target also has zero rate.
Change
v8-A: add action_saturation (-1.0, thresh 0.9, pre-clip) AND halve action_rate_l2 (-0.2 -> -0.1, still 3.3x the v5 value); success criterion pre-declared (joint_pos_ref telemetry returns to v6 scale).
Outcome
Booked as the structural repair of the v7 freeze; also fixed a config hygiene trap discovered on the way - action_rate was assigned twice in __post_init__ (v5 comment line then v7 line), merged to one assignment "别再留两处赋值给下次审计埋雷".
Mechanism
Clipping creates a zero-gradient, zero-cost absorbing region in action space; any penalty on action derivatives makes that region strictly optimal once entered. Only a penalty on clamp proximity itself (measured pre-clip so depth of violation is visible) restores a slope out of the absorbing region.
Applies when
- joints sit at exactly the action clip with near-zero variance
- raising a smoothness penalty degrades gait amplitude
- shaped-oscillation terms collapse after a rate-weight increase
“钉死在钳位的关节 action_rate 代价精确为零且永远为零;−0.2 之下"推到钳位冻起来"成了压倒性最优 … 本仓第二次栽在同一坑(walk_v0 死于四关节钉死 ±1.0)。回调权重(−0.2→−0.1)只降低诱惑不消除作弊 … 算在 clip 前的原始网络输出上 … 作弊收支从"白赚"变成"倒贴 ~250 倍"。”
train/WALK_V8_SPEC.md § 1. 改动 A — 治饱和作弊 The latency DR range must cover the measured deployment pipeline - 0-20 ms could not even reach the real 1-2 control steps
latency-dr-covers-measured-pipelineMeasure end-to-end action latency in control steps on your own stack (including cross-process queue boundaries), set the DR range to cover it with margin, and never import a delay count without its control frequency.
Symptom
Action latency was randomized over 0-20 ms (0-1 control step at 50 Hz), but the measured deployment path is 1-2 steps: the deploy process writes the target, an independently running BusWorker picks it up on its NEXT cycle, plus CAN round-trip - the training range could not cover the robot's actual latency at all.
Context
Fix: widen action_latency_s to 0-0.06 (0-3 steps). The external reference's "uniform 6 steps" was explicitly NOT copied - that number depends on his unknown control frequency; locally, a sweep at 0/1/2/3 steps showed walk_v5 survives all with insensitive metrics, so 6 steps "在我们这里没有依据" (has no local basis). The range was set from the measured pipeline with margin, not from a foreign constant.
Change
action_latency_s (0, 0.02) -> (0, 0.06), justified by pipeline analysis (writer/worker cycle boundary + bus time) and bounded by the local latency sweep.
Outcome
The DR band now brackets the true deployment latency; the policy trains against the delay it will actually face instead of a fictional sub-step world.
Mechanism
Latency DR only immunizes against delays inside its support; a range below the physical pipeline guarantees an untrained distribution shift at deployment. The correct range comes from tracing the pipeline's worst case (queueing boundaries + transport), and foreign step-counts are meaningless without the control rate they were measured at.
Applies when
- setting or auditing action-delay randomization
- deployment uses a separate bus/worker process from the policy loop
- importing delay-modeling numbers from other projects
“现行 0~20 ms = 0~1 个 50Hz 控制步, 而实测部署链路是 1~2 步(deploy 写 STATE.target 后, 独立跑的 BusWorker 下一轮才取走下发, 再加 CAN 往返)——现在的区间覆盖不到真机的实际延迟。… 不照抄参考来源的"统一 6 步": 那取决于他的控制频率(未知), 而我们扫过 0/1/2/3 步 … 6 步在我们这里没有依据。”
train/WALK_V7_SPEC.md § ⑤ action_latency_s 0~0.02 → 0~0.06 Calibrate a wall penalty by measuring healthy and sick policies - healthy pays ~0, the disease pays a wall
calibrate-threshold-between-healthy-and-sickCalibrate every threshold penalty by evaluating its exact formula on replays of at least one healthy and one sick policy: place the threshold between their distributions, size the weight so the sick policy pays a decisive fraction of tracking while the healthy one pays ~0, and pre-compute neighboring thresholds for cheap adjustment.
Symptom
Real walk_v7 occasionally clipped its own legs (stance narrowed to 133 mm mean vs nominal 214.5); a foot-distance penalty was needed, but an uncalibrated threshold/weight risked either doing nothing or becoming a reverse barrier.
Context
The term (relu(d_min - lateral foot distance), measured in the base yaw frame because world-frame y is meaningless after turning) was calibrated before training by replaying three known policies through the exact reward formula at the acceptance operating point: healthy v5 (183 mm) pays 0.6% of tracking - effectively free; narrowed v7 (133 mm) pays 22% - effective widening pressure; collapsed v8 (93 mm) pays 55% - a wall. d_min 0.16 was placed deliberately between healthy and sick, with alternative thresholds (0.14/0.18) pre-computed in the tool output for later adjustment. The shape self-check was named as a standing question: "先问'零代价的选项是什么'" - the zero-cost region must be exactly the desired behavior.
Change
feet_lateral_distance added at d_min 0.16 / weight -10, with telemetry expectation pre-registered (should decay toward 0 as stance learns >160 mm; if bow-legged over-widening >214 appears, only then discuss an upper bound).
Outcome
v9_probe on hardware: no leg contact ("没碰腿(N2 兑现)"), stance min 145/126 mm green; the term's zero-cost design left healthy gait untaxed.
Mechanism
A relu threshold penalty defines a free region and a priced region; its correctness is entirely in where the boundary sits relative to the healthy and pathological distributions. Replaying known-good and known-bad policies through the exact formula measures both distributions in the term's own currency, making the threshold and weight a placement decision instead of a guess.
Applies when
- adding any relu/threshold-style wall penalty
- a safety margin (foot distance, joint limit, clearance) needs enforcement without taxing normal behavior
- choosing between candidate thresholds for a new term
“形状自检 (v6 横向组/v8-B 的教训 —— 先问"零代价的选项是什么"): 标称站距付 0, 健康步态付 ~0, 收窄才付费 … v5(健康) 183 mm … 0.6% ≈ 免费 | v7(收窄) 133 mm … 22% —— 有效推宽 | v8(塌陷) 93 mm … 55% —— 墙 … d_min=0.16 恰在 v5(183)与 v7(133)之间”
train/WALK_V9_SPEC.md § 2. N2 —— 脚距惩罚(已完成权重预标定) FK-verify a borrowed reference's structure, then size its amplitude by the reference's job - it pins phase, the policy adds lift
reference-structure-fk-amplitude-divisionWhen borrowing a reference trajectory: verify its structural claim against your own kinematics (an invariant like flat-foot), assign it the phase-pinning job, and size amplitude low enough that the policy contributes the lift - moving toward a proven foreign value in halves, not jumps.
Symptom
walk_v4 had big knee swing (40-46 deg) but only 18-24 mm foot lift - amplitude without hip/knee/ankle phase coordination; later, walk_v5's real-robot swing ballooned to 73.6 deg (sim 55.7) with violent footfalls - amplitude over-driven by the reference.
Context
Structure first: Humanoid-Gym's 1:2:1 hip:knee:ankle reference was verified on the local model before adoption - the ratio exactly satisfies the locally derived flat-foot constraint hip - knee + ankle = 0, FK-tested at multiple amplitudes with sole pitch 0.00 deg throughout. Amplitude second, and here the first reasoning failed honestly: FK said shorter legs need LARGER reference scale (0.30 for 30 mm lift), and the FK was correct - but the premise was wrong ("FK 没错, 但前提错了"): it assumed foot lift must come from the reference. HighTorque Pi, same scale, uses 0.08 with a 0.02 m foot-height target - proof that lift is added by the policy ON TOP of the reference, whose actual job is pinning the phase relationship. Scale 0.30 made the reference the entire gait: over-constrained and over-driven. The correction went to 0.15, deliberately not Pi's 0.08: "一次只走一半, 留退路" (walk half the distance, keep a retreat).
Change
target_joint_pos_scale 0.30 -> 0.15 as one of v6-minimal's three changes, treating both the footfall force and the lateral kicking (yaw momentum scales with leg swing amplitude).
Outcome
v6 improved landing force 1.72x -> 1.55x, suspended tilt 45.9 -> 23.0 deg, turn-gain asymmetry 70% -> 19%; the later v6-halved-shaping experiment (35 mm -> 4 mm collapse) confirmed the reference still carries the gait's existence on this machine - the division of labor is real but machine-specific.
Mechanism
A joint-space reference plays two separable roles: encoding structure (phase relations that keep the foot flat) and injecting amplitude (energy). Structure transfers across robots and is checkable by FK against an invariant; amplitude is a negotiation with the policy, and over-assigning it to the reference removes the policy's freedom to modulate lift with state.
Applies when
- importing a reference gait / imitation target from another codebase
- reference amplitude reasoning based on leg length alone
- real swing amplitude far exceeds sim's under a strong reference
“FK 没错, 但前提错了。我默认抬脚必须由参考轨迹产生。HighTorque Pi 同尺度机器人 … 用 0.08, 而它 target_feet_height = 0.02 m —— 说明抬脚是策略在参考之上加出来的, 参考只负责钉住髋/膝/踝的相位配合。我们取 0.30 等于让参考本身就是整个步态, 过约束 + 过驱动”
train/WALK_V6_MINIMAL.md § ① target_joint_pos_scale 0.30 → 0.15 Before training a one-leg stand, the accounts and a probe showed the default gains could not hold it at all - kp 20 needs 0.39 rad of error to carry the static roll moment, more than the whole adduction range - so per-joint gains came first, and thermal limits set the session length
single-support-gain-authority-probeBefore training a posture that loads one joint statically, compute the steady tracking error load/kp and the series stiffness against m*g*h, and prove with a simple hand-written controller that the posture can be held under the deployment gains - change the gains first if it cannot; then size session length from the thermal account.
Symptom
The one-leg line (standing on one foot, the other folded back, no hopping) had to decide whether the existing gain profile could hold single support before any reward was designed.
Context
Hardware accounts (9.792 kg, COM 0.234 m high, 170 x 80 mm feet, legs 80% of the mass): moving the COM over one foot needs 107 mm of shift and the 20 deg hip-roll adduction range gives 131 mm - geometrically enough. The static frontal moment is 7.8-9 N*m, within RS02's 17 N*m - torque is enough. But at kp 20 carrying 7.8 N*m needs 0.39 rad of tracking error, more than the entire adduction range, and the real robot had already shown it: commanded +0.17, actual -0.04 (0.21 rad droop) under load, 0.0008 rad hanging - load, not the motor. A probe (probe_oneleg.py) then showed open-loop PD cannot hold single support on physics grounds, so the criterion became "an equilibrium exists and a hand-written 4-gain COM feedback can hold it": single-support roll stiffness is hip and ankle in series and must exceed m*g*h_com = 22.5 N*m/rad; ankle kp 12 in series with hip kp 80 gives only 10.4 (open loop 16/16 fell), ankle 60 with hip 80 gives 34.3 (52% margin).
Change
A per-joint gain profile (rl_oneleg: hip_roll kp 80, ankle_roll kp 60, the rest as rl_default) - which needed per-joint gain support in robot.yaml, the bridge, deploy and the trainer's actuator groups - decided before training. Thermal account: single support makes hip_roll the dominant heat load (about 7.8 N*m against a 7 N*m continuous rating), so acceptance and demos run in segments of at most 60 s with a temperature check.
Outcome
Under rl_oneleg the hand-written feedback held six cells cleanly for 6 s (hip_roll steady torque 2.1-3.4 N*m, half the thermal budget); under rl_default the same feedback on the same cells fell 0/4. The trained V0 policy then passed its 40-cell acceptance.
Mechanism
With PD position control, the steady error needed to carry a static load is load/kp; when that error exceeds the joint's range the posture is unreachable whatever the policy does, and series compliance between joints lowers the effective stiffness below the gravity stiffness that single support demands.
Applies when
- single-support, crouched or one-arm-load postures on PD actuators
- a joint "droops" under load on hardware but tracks well when hanging
- deciding whether a new skill needs its own gain profile
“但 kp=20 时撑住 7.8 N·m 需要 **0.39 rad 跟踪误差 > 整个内收行程**。真机已实测: 命令 +0.17 实际 −0.04(droop 0.21 rad),悬挂时 0.0008 rad——是负载不是电机。 … 单支撑滚转是 hip/ankle **串联**刚度,必须 > m·g·h_com = 22.5 N·m/rad;ankle kp12 串 hip80 只有 10.4(开环 16/16 全摔),60 串 80 = 34.3(裕 52%) … **rl_default 同反馈同格 0/4 全摔**(增益档必要性对照)”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §1-1 单脚站: 几何可行,卡点是 hip_roll 增益权限 / §2 A 线增益 / §5 probe 定谳 Foot dragging is an attractor, not a low amplitude - and joint damping is the mode switch, adjustable at deploy time
swing-bistability-damping-switchWhen a quality metric is bimodal, stop treating it as an amplitude to be trained up: map the modes against initial conditions and plant parameters, find the parameter that switches basins, apply it first as a deployment lever, and only then bake it into the training distribution (as a plant-family shift, never as an execution-mapping change).
Symptom
s2e_pd-1400's swing height "median 12.1 mm" hid a perfect bimodal distribution: 20 seeds split into a drag mode (2.6-4.9 mm) and a step mode (19.3-24.0 mm) with NOT ONE seed in between - the median sat in the empty gap, and "swing debt -11 mm" really meant "50% probability of falling into the drag attractor".
Context
Two designed experiments closed the mechanism. Test A (nominal plant, 40 seeds): step 42% / drag 58% / middle 0 - at nominal gains, initial conditions alone pick the mode, both modes 100% survivable. Test B (fixed init, kp x kd grid): kd is the mode SWITCH - at kd 1.3 all surviving cells step (13-22 mm), at kd 0.7 nearly all drag (2.7-4.3), only at kd 1.0 does init get a vote; kp >= 1.2 is dangerous (5/6 falls). Global verification at kd x1.3 (20-seed, delay 2): survival 20/20 at ZERO cost, step share 42 -> 80%, swing median 12.1 -> 18.4 mm, slip record low 334, thicker tilt margin - costs: vx 85 -> 78%, saturation +5 pp. A Pareto sweep then priced the knob: step share 42/72/75/88/82/90 across kd 1.00-1.30 with a linear vx tax of -2.3 pp per 0.1 kd - the basin gain is fully collected at kd 1.20 ("1.30 是 over-damping 纯多付税"). Mechanism: low damping leaves a landing micro-oscillation / ground-slide channel the policy can exploit to drag; damping plugs the channel.
Change
Deployment lever adopted: kd-scale 1.20 (conservative 1.15) as the legitimate successor to the power-0.8 crutch ("前者削幅度保稳,后者堵 拖地通道换步态,且不牺牲存活"); training-side prescription: move the DR band to nominal-1.2 x (0.9,1.1) = [1.08,1.32], deleting the [0.7,1.0) drag-teaching zone - a contract-level change requiring digest re-baselining, gated on measuring the real robot's actual kd dispersion first.
Outcome
The kd surgery rung (s2e_kd) delivered basin 8 -> 11/20, slip 405 -> 331, vx 81 -> 85% with no out-of-band fragility (below-band check 20/20) - "拐杖烧进分布的正确姿势", explicitly contrasted with the failed s1g amplitude version: this one changes the plant family the policy has seen, that one changed the execution mapping the policy would have to relearn.
Mechanism
The gait's swing behavior is a bistable dynamical system whose basin boundaries are set by plant parameters; a policy trained across a kd band that includes the drag basin has learned to inhabit it. Shifting the deployed (and then trained) damping moves the system into the step basin without touching the policy - a plant-side fix for what looked like a training deficiency.
Applies when
- a gait quality metric splits into distinct modes across seeds
- deciding between more training and a gain/damping change
- converting a deployment crutch into a training-distribution change
“20-seed 里拖地模式 2.6~4.9mm 与迈步模式 19.3~24.0mm 各半,中间一个不落 … kd 是模式开关——kd1.3 下 6/6 存活格全迈步 … kd0.7 下几乎全拖地 … swing 债的解(至少大半)在部署端阻尼档,不在训练端 … 机理:低阻尼下落脚微振荡/贴地滑给了策略顺势拖行的通道,加阻尼堵之。”
train/README.md § swing 双稳态定性 + kd 部署杠杆 (2026-08-07, 用户设计 Test A/B) An outer heading P-loop at deploy cut drift 10x because its output stays inside the trained command band - then training was aligned to it
deploy-heading-loop-and-align-trainingFix drift-class problems first with an outer loop whose output provably stays inside the trained command band; when adopting it permanently, align the training command generator to the deployment's actual command mixture (feedback-driven AND constant), matching law, gain, and clip exactly.
Symptom
Persistent heading drift on straight-line walking (v6 net yaw 60.3 deg over 15 s) that reward-side fixes had only partially tamed.
Context
The deploy stack added --heading: an external P loop wz = clip(0.5 * wrap_to_pi(theta0 - theta), +/-0.6), recomputed each frame and fed into the policy's ordinary wz command slot. Measured: net yaw walk_v6 60.3 -> 5.8 deg, walk_v5 17.4 -> 4.2 deg. A run-level audit later corrected the mechanism story: training had heading_command=False since v1 - the policy had NEVER seen heading-error feedback, so the loop works purely because its output lands inside the trained command distribution wz ~ U(+/-0.6): "收益真实,当时的机理解释写错了" (the benefit is real; the mechanism explanation had been wrong). v8 then closed the loop properly: training-side heading command enabled with rel_heading_envs=0.5 - half the envs get heading-error-driven wz, half get explicit constant wz, because deployment feeds wz BOTH ways (straight-line = heading feedback, turning = constant command) and rel=1.0 would have made constant-wz turning out-of-distribution. The law, gain, and clip were aligned item-by-item between trainer and deploy tool.
Change
Deploy-side outer loop first (no retrain needed); then v8-D enabled the matching training-side heading command at rel=0.5 with identical gain (0.5) and clip (+/-0.6), contract unchanged (wz slot carries the computed value).
Outcome
Drift handled at deploy (5.8 deg) generations before training caught up; the alignment removed the residual train/deploy distribution mismatch, with the accepted cost booked (open-loop straight walking becomes more OOD for heading-envs - irrelevant since acceptance and deployment always run the loop).
Mechanism
A learned velocity-tracking policy is a valid inner loop for any outer controller whose commands stay within the trained command distribution - the policy needs no knowledge of the outer objective. Full alignment then requires training on the same mixture of command sources the deployment actually uses, in the observed proportions.
Applies when
- heading/position drift on a velocity-tracking policy
- designing outer loops over learned locomotion controllers
- training command distribution differs from how deployment feeds commands
“审计更正(2026-08-02,run 级 env.yaml):训练侧自 v1 复盘起就是 heading_command=False … 策略从未见过航向误差反馈。--heading 是评估/部署侧外加的航向 P 环(wz=clip(0.5·err,±0.6), 落在训练分布 wz~U(±0.6) 内)。实测净偏航 walk_v6 60.3° → 5.8° … 收益真实,当时的机理解释写错了”
train/WALK_V7_SPEC.md § 0. 本轮之前已经改掉 (航向闭环, 含审计更正) The restart reward table lists a reason for every term AND a lesson for every exclusion - absent terms are removed, not zero-weighted
minimal-reward-table-with-provenanceMaintain the reward table as an evidence ledger: every term cites the episode that justifies it, every excluded term cites the episode that convicted it (including the development stage it is valid at), and retired terms are deleted from the config, never left at weight zero.
Symptom
Seven walk generations had accumulated an entangled reward table where nobody could say which term earned its place; the restart needed a table that could be audited line by line.
Context
The minimal table v2 was built under three written principles: "一项管 一件事、结构性反抬脚的项一个不留、塑形只留一套相位逻辑;不在表里的一律 不加" (one term per job; zero structurally-anti-lift terms; exactly one phase-shaping logic; nothing outside the table gets added). Every row carries its provenance (e.g. base_height target = standing height cites the v4 crouch lesson; split x/y tracking cites the merged-exp gradient hole; world-frame yaw cites the v1/v3 body-frame lesson). Every EXCLUSION carries its same-type precedent: feet_landing_vel out because it is poison while the gait is unformed (drag pays 0, lifting pays - the v4-clearance / v8a-B "reverse threshold" family) though it was fine in v7 when the gait already existed - term validity depends on developmental stage; feet_air_time out with its measured non-lever evidence (v6 had it at 2.0 and still lifted 4 mm); and replaced terms are REMOVED from the config ("置 None,不是权重 0 挂着") so audits see truth, not dormant weight.
Change
Reward table rebuilt as ~20 rows each with weight + provenance column; exclusion list maintained alongside with the falsifying episode for each; dormant terms deleted rather than zeroed.
Outcome
Later revisions (S1.2/S1.3) modified the table by citing and updating specific rows' evidence rather than re-arguing the whole design; the table doubled as the lineage's reward-lesson index.
Mechanism
A reward table is a set of standing hypotheses; attaching each row's evidence makes revisions targeted and reversible, and recording why a term is absent prevents the cycle of re-adding known poisons. Deleting vs zero-weighting matters because config audits and DR interactions see the term either way - a zero-weight term is dormant complexity waiting to be flipped on wrongly.
Applies when
- designing a reward table for a restart or new task
- someone proposes re-adding a previously removed term
- auditing which reward rows still earn their place
“原则:一项管一件事、结构性反抬脚的项一个不留、塑形只留一套相位逻辑;不在表里的一律不加 … feet_landing_vel(评审 #4):拖地时代价恒 0、抬脚才收费——与 v4-clearance/v8a-B 同属「反抬脚门槛」家族,步态未成形时是毒;v7④ 加它时步态已存在。… 已从 cfg 移除(置 None),不是权重 0 挂着。”
train/OMNI_V0_SPEC.md § 3. 最小奖励表 v2 / 明确不带 Knee swing collapsed because it directly trades against the slip penalty - price the conflict explicitly and clamp what reward cannot hold
knee-swing-vs-slip-pricingWhen a behavior collapses as another metric improves, look for the term pair trading them and set their price ratio deliberately (with escalation and reverse tripwires pre-registered); where the policy actively spends action budget to undo your target, stop paying more reward and clamp the target space structurally.
Symptom
Knee peak-to-peak swing collapsed across generations - v5 33 deg, v10 26-30, v10b 7-8, v11 6.5-8.6 - and rolling the clock back did not recover it, acquitting the clock; the collapse tracked the gated slip penalty instead: "屈膝与不打滑在当前奖励里直接对抗" - v10b's excellent 93 deg slip was purchased with knee amplitude.
Context
Reward-side flexion fixes had failed three times: raising reference amplitude backfired twice (v9/v11), and v11's deep-squat default was actively fought by the policy - it spent 0.68 of action budget pulling the squat straight ("被策略花 0.68 动作拉直反杀"). v12's design accepted the conflict as real and attacked on two tracks: (1) ECONOMICS - a direct knee_swing_amplitude reward (+0.3, target 0.55 rad, capped at 0.6/step = 55% of tracking), explicitly opposed to the slip penalty by design ("显式对立——这正是设计:v12 就是这场对抗的定价实验"), with an escalation ladder (K +0.3 -> +0.5, then slip -0.5 -> -0.3, one layer at a time) and a reverse tripwire (slip telemetry back at v10 levels -> slip weight to -0.8, accept ~20 deg knee compromise); (2) STRUCTURE - knee target bounds [0.2, 0.9] rad so full straightening is physically impossible (straightest 11.5 deg) and the 0.68 fighting budget is released. A bonus falsifiable prediction was attached: phase-lock strength tracks amplitude (v9_probe 48 deg locked 2.5 Hz; v11 low-amplitude 1.36 Hz unlocked), so if K works, hardware phase-lock should return - one change, two verdicts.
Change
knee_swing_amplitude reward + knee target clamp + pre-registered escalation/reverse levers; the failed reward-side-only approach retired.
Outcome
The lineage was frozen before v12 trained (strategic reset), but the diagnosis stands as the walk line's clearest example of two reward terms trading a behavior between them, with the pricing experiment and structural clamp fully designed and calibrated.
Mechanism
When two terms price opposite aspects of one motion (swing amplitude creates yaw momentum that becomes slip), the optimizer settles wherever the price ratio puts it - patching one side moves the equilibrium, not the conflict. Explicit pricing makes the trade a designed quantity; structural clamps remove the regions where the policy spends budget fighting the designer.
Conflicts
The pricing experiment (K vs slip) was designed and calibrated but never trained - the 2026-08-05 reset suspended v12; the collapse attribution table and the 0.68-action counterattack are measured, the remedy's效果 is untested.
Applies when
- one gait quality degrades in lockstep with another's improvement
- the policy visibly fights a default pose or reference
- repeated reward-side fixes for the same behavior have failed
“膝摆塌在 v10→v10b,头号嫌疑是门控滑移罚(四代实测膝 p2p:v5 33° / v10 26~30° / v10b 7~8° / v11 6.5~8.6°;退时钟没救回 → 非时钟)——"屈膝"与"不打滑"在当前奖励里直接对抗 … 奖励侧修屈膝已三败 … v11 深蹲 default 被策略花 0.68 动作拉直反杀”
train/WALK_V12_SPEC.md § 0. 定位 / 2. K —— 膝摆经济(与滑移罚的对偶)