Training Coach
Doctrine
A report may cite any of these as doctrine-N.
doctrine-1Contract freeze and fingerprint disciplineThe policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.
Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).
Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.
doctrine-2Attribution by resolved training params - never eval-override knobsCapability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.
Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).
Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.
doctrine-3PASS gates become constraints; FAIL gates become objectivesOnce a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.
Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.
Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".
doctrine-4One variable per ladder rung - counted against what the checkpoint sawA rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.
Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.
Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.
doctrine-5Pre-register risks, readings, and stop criteria before the ladderBefore a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.
Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).
Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.
doctrine-6Plant parameters are measured, never inventedEvery plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.
Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).
Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.
doctrine-7Sim2sim gate before sim2real - under deployment conditionsEvery checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).
Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).
Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.
doctrine-8Observation honesty - the actor's inputs are a hardware contractThe actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.
Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).
Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.
doctrine-9Reward economics are audited in realized currencyReward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.
Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).
Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.
doctrine-10The zero-cost option must be the desired behaviorFor every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.
Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).
Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.
doctrine-11Measurement discipline: independent referees, signs, distributionsA disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.
Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).
Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.
doctrine-12The deployment pipeline is plantIrreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.
Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).
Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.
doctrine-13DR budget is finite; its distribution is the measured supportRobustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.
Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).
Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.
doctrine-14Gates measure what hardware feels: posture, margins, stripped assistsAcceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.
Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).
Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.
doctrine-15Fork and root selection: recoverability, maturity, frozen rewardsChoose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.
Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).
Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.
doctrine-16Curricula: verified engagement, lineage counters, disease-phase gatingAutomatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.
Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).
Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.
doctrine-17Probe before training: feasibility first, hypotheses in tablesAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.
Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).
Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).
doctrine-18External advice is recomputed locally; values transfer as ratiosEvery external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.
Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).
Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.
doctrine-19Hardware sessions are scripted experiments, not tuning sessionsReal-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.
Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).
Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.
doctrine-20Close questions in writing; restart when the debt is structuralAudited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.
Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).
Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.
doctrine-21Name the quantity in the space it lives inA goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.
Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).
Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.
doctrine-22Continuation needs a live gradient; a release is chosen by a scanContinue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.
Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).
Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.
Experience cards
93 cards matching “kd-bandwidth-mu-law-attribution”.
Cutting swing amplitude 40% raised yaw-momentum demand 53% - the falsified fix is recorded so nobody walks that road again
amplitude-cut-falsified-yaw-fixTest gait fixes against the quantity the ground must actually supply (torque/force rates vs friction ceilings), not against kinematic proxies; record falsified fixes with their mechanism so the search space shrinks permanently.
Symptom
Support-foot yaw slip stayed at ~223 deg (vs 225 deg) after walk_v6 cut joint swing amplitudes by ~40% (hip_pitch 20.6->10.4 deg, knee 31.3->18.8 deg) - the change built on the theory "smaller swing = less yaw momentum to dump into the ground".
Context
Direct measurement inverted the theory: yaw-momentum amplitude ROSE 16% (+/-0.1000 -> +/-0.1159) and its rate of change rose 53% (1.86 -> 2.84 N*m demanded from the ground), pinned exactly at the foot's supply ceiling (2.8-3.3 N*m at mu 0.6-0.7) - so slip could not drop. The extra demand lives in higher harmonics: v6's crisper foot placement (better clearance 34 mm, lower landing force) shortens the momentum- exchange window, concentrating the same exchange into less time. The verdict was written as a closed road: "下一轮不要再走这个方向". An honest residue was also booked: WHY amplitude down but momentum up 16% remained unresolved, with the named next step (per-rigid-body decomposition of H_z, since per-joint RMS sensitivity ignores phase correlations).
Change
The "reduce amplitude to reduce yaw momentum" lever was removed from the planning space; future yaw-slip work redirected toward the supply side (friction) and momentum-rate mechanics.
Outcome
Slip unchanged (225 -> 223 deg); the falsification and its mechanism became a permanent constraint on the fix search space.
Mechanism
Ground yaw torque demand scales with the rate of change of angular momentum, not its amplitude; kinematic amplitude cuts that also sharpen contact timing can raise dH/dt while lowering range. When demand exceeds the friction-limited supply ceiling, slip is set by the ceiling, so demand-side changes below the ceiling do nothing visible.
Applies when
- attacking foot slip or yaw drift via gait shape changes
- a fix targets an amplitude while the constraint is a rate
- documenting a failed intervention after a version comparison
“walk_v6 | ±0.1159 (+16%) | 2.50 Hz | 2.84 N·m (+53%) … 脚的供给上限 2.8~3.3 N·m(μ 0.6~0.7)—— v6 正好顶在天花板上, 所以滑移一点没降。… 结论: "减小摆动幅度以降低偏航动量"这条被 v6 证伪 —— 砍 39% 幅度, 需求反升 53%。下一轮不要再走这个方向。”
train/WALK_DIAGNOSIS.md § ② 未生效的机理: 减小摆动幅度反而让偏航需求上升 Before resuming a checkpoint, diff the current cfg against what the checkpoint was trained with
resume-state-dr-audit"One variable per rung" counts variables against what the checkpoint actually experienced: audit the checkpoint's logged training config and align every unintended difference before resuming.
Symptom
Two consecutive rungs (C1 back-mode, C2' forward-turn) failed from the same root with the same full-regression signature despite adding different new modes - so the mode was not the cause.
Context
Both runs resumed s1e-500 with the then-current cfg, which carried PD band (0.8,1.2) plus three DR events (base_com, joint_friction, push_robot) accumulated by later lineages. Verified on the training machine from the source of truth (the run's logged params/env.yaml): s1e-500's actual training state was PD +/-10% (0.9,1.1) and all three DR events None. Resuming it under the new cfg meant eating 4 new plant variables plus a new mode at once - the intended "1 variable" was actually 5. A worse variant (c1_redo from s2e_pd-1400) added push +/-0.3 to a root that had never seen it: near-total collapse within +100 iters.
Change
C2 aligned the cfg to the checkpoint's training state before resuming (PD back to (0.9,1.1), three DR events off) - making the new mode the only true variable. Permanent rule recorded: compare the checkpoint's training-time DR with the current cfg before any resume.
Outcome
C2 trained successfully from the same root that had "failed" twice (wz 20/20 with genuine sign-antisymmetric response by iter 700-800); the A/B falsification ("两个不同模式同签名崩") plus the env.yaml verification closed the attribution.
Mechanism
A resumed policy is instantly evaluated (and its value function trained) under whatever plant distribution the cfg specifies; every DR term the checkpoint never adapted to is a distribution shift applied on day one, compounding with the intended change. Single-variable discipline is therefore a property of (cfg diff) x (checkpoint history), not of the cfg diff alone.
Applies when
- resuming or forking any checkpoint under an evolved config
- a resumed run degrades broadly within the first few hundred iterations
- two different changes from the same root fail with the same signature
“A/B 定谳:两个不同模式同签名崩 → 病因不是模式,是「从 s1e-500 续训」。… s1e-500 训练态 = kp/kd ±10% (0.9,1.1),base_com / joint_friction / push_robot 全 None;而 cfg 里带着 (0.8,1.2) + … 三个 DR —— 从它续训等于一次吃 4 个新 plant 变量 + 新模式 … 永久教训:续训前必须比对 checkpoint 的训练态 DR 与现行 cfg。单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」。”
train/C_LADDER_RUN.md § 3b. 这不是重复实验 —— 前两次的病根已定位并修掉 An auto-curriculum ratchet capped out at iter 248 and never engaged - stage difficulty manually or verify engagement
auto-curriculum-engagement-checkPrefer manually staged difficulty with gated transitions; if you use an automatic curriculum, instrument its internal state and alarm when it stops engaging - a saturated curriculum is constant DR wearing a curriculum's name.
Symptom
A curriculum mechanism intended to grow difficulty adaptively (s1f's ratchet) hit its cap at iteration 248 and never bit again - for 96% of the run its effect was equivalent to constant DR, i.e. the curriculum existed in name only.
Context
When external advice suggested graded wz bands (start ±0.15, then ±0.30), the team agreed with grading but explicitly rejected automatic curriculum, citing the s1f episode. The same logic had already been paid for with push grading: ±0.6 failed twice, ±0.3 was feasible - grading matters, but the grade transitions were made by hand at verified checkpoints.
Change
Ladder policy: difficulty staged manually, one band per rung, each transition gated by the acceptance battery; automatic ratchets not used unless their engagement is monitored and demonstrated.
Outcome
Every C-ladder band change (wz ±0.12-0.25 first, wider later) was an explicit, attributable rung; no silent constant-DR-in-disguise runs recurred.
Mechanism
Adaptive curricula couple their own state machine to noisy training metrics; a ratchet that saturates early stops adapting but keeps its name, so the operator believes difficulty is progressing when it is frozen. Manual staging costs more decisions but each decision is observable and reversible.
Applies when
- choosing between auto-curriculum and staged bands for a new skill
- a curriculum's difficulty parameter plateaus early in training
- post-hoc attribution of what difficulty a lineage actually saw
“C2 的 wz 分级(±0.15 → ±0.30,不要一上来 ±0.6)—— 与我们付过学费的 push 分级同形(±0.6 两轮 FAIL,±0.3 才可行)。但必须手动分级,不做自动课程 —— s1f 的课程化棘轮 iter 248 封顶未咬合,96% 时长等价常量 DR”
train/C_LADDER_RUN.md § 1. 采纳 3 条 (C2 的 wz 分级) Under continuous 3-axis uniform sampling, pure straight-line walking is a zero-measure event the policy never trained
zero-measure-commands-need-mode-samplingEnumerate the exact command points users will actually issue (straight, stop, in-place turn) and give each explicit probability mass via mode sampling with off-axes pinned to zero - never assume a continuous sampler covers its measure-zero subsets.
Symptom
"The robot drifts even in sim when told to walk straight" persisted across reward tunings - because with commands drawn as vx in [0.15,0.5] x vy ~ U(+/-0.2) x wz ~ U(+/-0.6), the event vy=0 AND wz=0 has probability zero: pure straight-line walking was never sampled even once.
Context
Restart evidence item #3: "纯直行是零测度点 … 'sim 里直行就漂'是分布的 必然,不是 reward 没调好" - the drift metric was legitimately drowned by commanded turning (v11's own comment self-documented this). The structural fix is discrete mode sampling: a custom ModeVelocityCommand that first draws a mode by share (stand/forward/back/turn/side/mixed), then draws values only on that mode's axes with all others pinned to exact zero - which is also what preserves single-variable discipline in the C ladder (native 3-axis uniform "采不出'离散模式桶' … 把 C1~C4 的单变量纪律直接毁掉"). The mixed mode later got an ellipsoid constraint rather than a cube for the same reason in reverse - corner combinations of a cube are unrepresentative extremes.
Change
Command generation moved from independent per-axis uniforms to mode-bucket sampling with pinned-zero off-axes (plus 20% rel_standing); acceptance likewise evaluates per mode.
Outcome
Straight-line behavior became a trained, testable mode instead of a measure-zero hope; the C ladder could add one mode per rung with provable isolation.
Mechanism
A policy optimizes expected reward under the command distribution; events of probability zero contribute nothing to the objective, so exact-zero-command behaviors (straight walk, stand, in-place turn) are only learned if the sampler gives them mass. Product-of-uniforms distributions concentrate mass on mixtures and give none to the pure behaviors users actually command.
Applies when
- a "simple" command (straight, stop) underperforms mixtures in sim
- designing command distributions for velocity-tracking tasks
- a ladder needs per-mode isolation for attribution
“纯直行是零测度点:最终 command 为 vx∈[0.15,0.5] × vy∈U(±0.2) × wz∈U(±0.6) 连续均匀,vy=0∧wz=0 从未被专门采样 —— "sim 里直行就漂"是分布的必然,不是 reward 没调好 … Isaac 原生 UniformVelocityCommand 是三轴各自 uniform,采不出"离散模式桶"”
train/OMNI_V0_SPEC.md § 0. 为什么从零 (3) / 三件前置 (1) A frame-history observation under zero DR memorizes the trainer's plant fingerprint - the estimator must see variation to learn estimation
history-obs-needs-plant-variationIf the observation carries history (stacked frames, RNN), keep at least minimal plant variation (gain/latency jitter) on from the first iteration - "nominal first, robust later" is structurally invalid for estimator-bearing contracts.
Symptom
omni_s1 (fresh 215-dim contract with a 5-frame history window, trained with DR fully off): training all green, yet the MuJoCo gate scored 0/20 on all eight doors - falls within 2 s, seven checkpoints, not one transferred.
Context
The history window exists precisely to let the actor implicitly estimate line velocity and actuator dynamics (the actor is denied base_lin_vel by observation honesty). Under a constant plant that implicit estimator has nothing to estimate - it learns the trainer's exact response fingerprint instead, and any other simulator's micro-differences are out-of-distribution: "5 帧窗按设计就是隐式估计器, plant 恒定时它学到 Isaac 精确响应的指纹". The planned "nominal-first-robust-later" staging was declared STRUCTURALLY incompatible with history observations: "估计器要见过变化才学估计, 否则学背诵" (an estimator must see variation to learn estimation, otherwise it learns recitation). Honest confound note kept: this is mixed with "zero DR does not transfer, period" - but both attributions prescribe the same fix, so no control was run.
Change
S1.1: minimum actuator jitter turned on from day one - kp/kd +/-10%, latency 0-1 frame (friction/COM/mass still nominal, no push - those stay for the S2 ladder).
Outcome
Transfer restored: survival 0/20 -> 20/20, speed 19/20, foot distance 20/20 (remaining failures moved to gait quality, a different disease); the staging doctrine was amended - history-carrying contracts never train under a frozen plant.
Mechanism
A recurrent/history channel fits whatever temporal structure minimizes loss; with a deterministic plant the cheapest structure is the plant's own impulse-response signature, yielding features that are simulator-specific rather than physics-general. Plant variation forces the channel to carry state-estimation features that transfer.
Conflicts
Attribution is explicitly confounded with the simpler "zero DR never transfers" reading ("与「零 DR 本身就不迁移」混杂 … 两种归因处方相同, 不做对照") - the source chose not to spend a control run separating them.
Applies when
- adding frame stacking or recurrence to an actor observation
- a nominal-plant policy fails a cross-simulator gate within seconds
- planning DR staging for a new contract
“frame_hist × 零 DR = plant 指纹过拟合——5 帧窗按设计就是隐式估计器, plant 恒定时它学到 Isaac 精确响应的指纹, MuJoCo 的微小差异即 OOD, 2 s 内摔, 七个 checkpoint 无一迁移。「先标称后鲁棒」的分段与历史观测结构性冲突:估计器要见过变化才学估计,否则学背诵。”
train/OMNI_V0_SPEC.md § 3. S1.1 修订记录 ① The walking lines' safety setting, power-scale 0.8, broke the recovery policy's full-range contract - it cut the ends of the joint travel (4/50 could not get up) and left the torque spikes untouched; a kp x 0.9 gain profile inside the trained kp band did the job
power-derating-cuts-full-range-contractA deployment derating knob means something only relative to the action contract: before reusing a line's "safe setting" on a new skill, check what it does to that skill's reachable range and to the term that makes the spikes, prefer a gain change inside the band the policy was randomized over, verify it in simulation, and re-decide when the contract changes.
Symptom
After the violent first real get-up (2026-08-09), the recovery policy needed a gentler setting for its next hardware test, and the walking and omni lines' standard derating - deploying at power-scale 0.8 - was the obvious candidate.
Context
The V0 recovery contract maps actions to absolute targets over the full joint range: a = +/-1 lands exactly on the URDF limits, and standing puts the knee at the clip. Candidates were compared on R3.1 in MuJoCo (5 categories x 10 seeds) on 2026-08-10 before any hardware time was spent.
Change
A new gain profile, rl_kp090 (kp x 0.9, kd unchanged), recorded in robot.yaml as the recovery hardware-test setting, with power-scale 0.8 explicitly banned for recovery.
Outcome
kp x 0.9: 48/50 got up; median torque demand on hip_pitch/knee fell from 120-125% to 100-104% of the deployment limit; leg-leg contact frames 2,152 -> 1,095; the change sits inside the +/-10% kp randomization the policy trained with. power-scale 0.8: 4/50 could not get up, because under the full-range contract it removes the ends of the travel (the deep squat's tucked legs, the straight standing knee), and the torque spikes (kp x error) did not fall at all. When the line moved to the beta-anchored contract, rl_kp090 was declared a V0-era choice that does not fit (beta is calibrated at kp 30) and deployment returned to rl_default; the deploy switch applies power scaling to the walking side only.
Mechanism
A power scale multiplies the action, which under an absolute full-range mapping shrinks the reachable workspace instead of softening the actuator; the spikes come from the proportional term on large errors, which only a gain change reduces - and a gain change inside the trained randomization band stays in distribution.
Conflicts
The undated operator runbook still carries an R3.1 "B comparison" command at power-scale 0.8 beside the rl_default baseline; the sources do not say whether it was written before the ban or was ever run.
Applies when
- reusing a power, torque or action scale from one skill on another
- a policy whose actions map to absolute targets over the full joint range
- choosing a gentler setting for a first or second hardware trial
“kp×0.9 / kd 不动 —— recovery_r3_1 成功 48/50, τ 需求中位 hip_pitch/knee 120~125% -> 100~104% 部署限, 腿-腿接触 2152 -> 1095 帧; ±10% 在训练 kp DR 带内. ⚠️ power-scale 0.8 对 recovery **禁用**: 全 ROM 契约下 0.8 砍的是行程 端点 (深蹲收腿/站直够不到), 实测 4/50 起不来, 且尖峰 (kp·err) 一点不降 —— 它是 walk/omni 的安全档, 不是 recovery 的.”
git:Lucen-recovery@origin/recovery:robot.yaml § gain_profiles 注释: recovery 真机测试安全档 (2026-08-10) / rl_kp090 The advisor's "runtime five-stage state machine + per-stage reference poses + RL residual" appeared in none of the three papers it cited - reading the originals changed the plan and downgraded two widely repeated industry claims
advisor-paraphrase-vs-paperRead the primary source behind any piece of advice before adopting its architecture; record where the paraphrase and the original differ, downgrade claims the originals do not support to speculation, and adopt what the verified sources actually share.
Symptom
After the violent first real run, an advisor proposed re-architecting recovery as a runtime staged state machine with reference poses and an RL residual, citing HoST, HumanUP and StableMimic.
Context
The three papers were read in full on 2026-08-09 and tabulated (deployment form, what the "stages" really are, hard constraints, references). HoST: one end-to-end policy, height-gated rewards in training, action anchored as q + beta*a with a beta curriculum. HumanUP: two training stages with the same observation/action; the vendor's three-stage state machine is the baseline it beats (41.7% vs 78.3%); Stage II tracks an 8x slowed Stage I trajectory (4x too violent, 10x does not converge). StableMimic: a learned soft gate. On 08-10 more sources were checked the same way: the Agility page does not say Digit's self-righting was learned in simulation (only step recovery is stated as RL), so the claim was downgraded to speculation; HoST's support for the 12-DoF armless Mini Pi exists in its code repository, not in the paper text.
Change
Adopted only what the originals share: a hard action bound or anchored action space, strong smoothing including a second-difference term, a slowed hidden reference, heavy DR and real fallen states. The staged runtime state machine was not adopted.
Outcome
The next re-rooting candidates came straight from the verified material, and the beta-anchored action space (HoST, with the Mini Pi configuration as the nearest real-robot precedent) became V2, the lineage that later stood up on hardware.
Mechanism
A paraphrase compresses a paper into the advisor's own architecture; only the original shows what was actually deployed, what was a baseline, and which numbers came with which ablation.
Applies when
- an advisor, agent or summary proposes an architecture with citations
- an industry claim ("X learned it in sim") is about to justify a design
- several papers are cited for one combined recipe
“⇒ **顾问的核心形态"runtime 五阶段状态机 + 每阶段参考姿态 + RL residual"在三篇引文 里均不存在**,其中 HumanUP 还点名 state machine 是局限。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 三篇引文精读判决(2026-08-09 全文核对;顾问转述与原文有出入) Bracket a real-robot A/B with a repeated reference run - battery drain is the confound
battery-bracketed-real-abOrder hardware A/B sessions as A-B-A: repeat the first condition at the end, and void the comparison if the bracket runs disagree - never let battery or venue drift ride on the second condition.
Symptom
In a two-policy teleop A/B on hardware, the second policy is measured on a lower battery voltage than the first - a systematic bias that would be read as a policy difference.
Context
C2 real A/B (checkpoint 700 vs A800, same floor, same day) was scripted as 700 -> A800 -> 700-rerun, with the explicit note that a teleop session drains the pack and the trailing policy "naturally suffers".
Change
Protocol: run the reference policy first AND last; if the two reference runs differ noticeably, declare the whole session battery/floor-polluted and void the A/B ("结论作废重来"). Also log electricity per run.
Outcome
Called out as the round's only systematic confound, closed by one extra command ("这是本轮唯一的系统性混淆源,一条命令就能堵掉").
Mechanism
Battery voltage scales available torque, and torque loss hits behavior asymmetrically (see power-scale-hurts-nonforward-axes), so drain masquerades as policy regression; a head/tail reference pair converts the unobserved drift into a measured control.
Applies when
- comparing two policies or settings on hardware in one session
- any sequential hardware evaluation where the plant drifts (battery, temperature, floor wear)
“为什么要 700 复跑:一次遥控 session 下来电池会掉压,第二枚天然吃亏。头尾各跑一次 700,若两次 700 明显不同,说明这轮 A/B 被电量污染,结论作废重来。这是本轮唯一的系统性混淆源,一条命令就能堵掉。”
train/C_LADDER_RUN.md § 3c. A-3 真机 A/B(同一段地板、同一天、电量记账) mj_objectVelocity returns inertial-principal-axis frame - one API assumption poisoned eval and observations for a whole line
body-frame-velocity-api-auditVerify every frame-sensitive API against a hand-computed truth (rotate raw qvel yourself, or command a known world velocity and check where it lands) before trusting any evaluation or observation built on it - especially when a model's inertial frame is rotated from its body frame.
Symptom
Sidewalk vy read ~0 under every condition; separately, whole-policy performance was mysteriously mediocre in sim2sim while training-side numbers looked fine. Four training rungs were declared FAIL partly on these readings.
Context
base_link's URDF inertial frame is rotated 90 deg about x relative to the body frame (iquat = [0.7071, 0.7071, 0, 0]). mj_objectVelocity(flg_local=1) rotates into ximat - the inertial principal-axis frame - not the body frame, and returns center-of-mass point velocity, not body-origin velocity. Consequences measured: the "vy" column was actually vertical velocity vz (walking at cmd 0.25: old reading +0.0093 vs true -0.0424); the angular velocity fed to the policy in sim2sim was [wx, wz, -wy] - a different quantity than Isaac and the real IMU provide. RMS check over 8 s of walking: y/z axes swapped between v6[:3] and the qvel truth.
Change
Fixed sim2sim and both probes to compute ang_b = qvel[3:6] and lin_b = xmat.T @ qvel[0:3] (identical quantity to Isaac's root_ang_vel_b / root_lin_vel_b), with a standalone reproduction script (frame_bug_repro_0809.py).
Outcome
Re-scoring the "failed" C4 lineage under correct coordinates reversed the verdicts: c4r4 checkpoints showed vy 80-126% tracking (old reading: +/-2%) and vx+0.30 at 91-95% where the old metric said 0/5 - the bad frame both mis-measured vy and, via corrupted policy observations, systematically depressed all measured performance. Final product passed 260/260 cells.
Mechanism
A simulator API's frame convention is part of the observation contract; when the model's inertial frame is rotated relative to the body frame, frame-agnostic use of a "local" velocity silently permutes axes. Feeding a policy an axis-permuted angular velocity is an observation corruption that degrades behavior everywhere, not just on the axis being studied.
Applies when
- building or auditing a cross-simulator evaluation harness
- one measured axis reads near-zero under all conditions
- sim2sim scores are inexplicably worse than training-side metrics
- URDF/MJCF inertial frames are rotated relative to body frames
“base_link 的 iquat = [0.7071, 0.7071, 0, 0] … mj_objectVelocity 用的是这个 … 喂给策略的 base_ang_vel 是 [wx, wz, −wy] —— MuJoCo 侧观测与 Isaac / 真机 IMU 不是同一个量;vy_mean 报的是竖直速度 vz —— 前进 cmd 0.25 时旧读数 +0.0093,真值 −0.0424。”
train/C_LADDER_RUN.md § 3l. ⚠️ mj_objectVelocity 读的是惯性主轴系 / 3m. 一 bug 坐实 The +/-50 mm lateral COM randomization meant to spread the legs coincided with legs pulling IN - rolled back per its own pre-registered contract
com-dr-rollback-on-symptomWhen adopting a DR value that covers no local measurement, write its intent and rollback trigger into the config at adoption time; roll it back as the control arm the moment the symptom contradicts the intent, and promote the symptom's metric into the acceptance battery.
Symptom
After v7 adopted the reference developer's oversized lateral COM randomization (+/-50 mm) explicitly to force leg spread, the real robot's legs narrowed instead - lateral mean 154 mm / closest 107 mm in sim (nominal 214.5), narrower still on hardware with occasional leg contact.
Context
The rollback was clean because the adoption had been honest: the robot.yaml comment recorded the intent AND that the +/-50 value covered no local measurement (only a 16/7 mm measured offset existed; even the prior widening to +/-20 was subjective), plus the reference's own reported side effect (base sway) and the note "这一项要单独跑、 单独归因". When the opposite symptom appeared, v8 returned y to +/-20 mm as the control arm ("要么没起作用、要么帮了倒忙 … 按约定退回做 对照"), kept x/z untouched (a noise-level difference not worth another variable), and named the second suspect: the landing penalty itself, via the reference's own three-link chain (landing penalty -> stance narrows -> spacing penalty needed). A gate lesson was booked in the same table: v7's sim numbers had ALREADY crossed the line (154/107 vs v5's 182/147) - "这个指标本可拦下 v7" - so foot-distance became a standing acceptance row (min >120 mm, zero leg-leg contacts).
Change
base_com_offset_m y: 0.050 -> 0.020 (x/z kept), regenerated through the export tool rather than hand-editing derived files; foot-distance acceptance row added.
Outcome
A borrowed DR lever with no local measurement basis was retired the moment its symptom contradicted its purpose, at single-variable cost; the metric that would have caught it pre-hardware entered the gate.
Mechanism
DR ranges shape behavior through the policy's robustness strategy, which is jointly determined with every reward term; a lever that forces stance width on one robot can be dominated by a stronger narrowing pressure (landing softness) on another. Levers adopted without local measurement must carry their own rollback trigger, because there is no nominal to argue from when they misbehave.
Conflicts
Causality is not fully closed in the source: the narrowing may come from the landing penalty rather than the COM lever ("腿距的第二嫌疑人是 ④ 本身"); the rollback is the pre-agreed control experiment, not a verdict that the lever caused the narrowing.
Applies when
- importing DR ranges or behavioral-forcing randomizations from references
- a DR lever's observed effect contradicts its documented purpose
- a sim metric existed that would have caught a shipped regression
“⑥ 的本意 … 是逼策略把脚分开;真机结果是脚向内收且偶发相碰——要么没起作用、要么帮了倒忙。… 注释当时就写了"这一项要单独跑、单独归因"。现在症状出现了,按约定退回做对照。”
train/WALK_V8_SPEC.md § 3. 改动 C — 质心随机化退回(撤销 v7-⑥ 的 y 项) A curriculum ramp keyed to the process step counter re-fires on every resume - and shipped policies never saw the penalty
curriculum-counter-lineage-stepsKey every curriculum/ramp schedule to lineage-cumulative progress, not per-process counters; and audit where your shipped checkpoints sit relative to every ramp - a penalty that no product ever experienced is not part of your training.
Symptom
vx+0.30 died at a fixed relative time in every resumed run: resume at 500 -> slide at 1100, zero at 1300; resume at 700 -> slide at 1300, zero at 1500 - absolute depths offset by exactly the resume offset, relative timetable identical.
Context
ramp_reward_weight (the saturation penalty ramp) read env.common_step_counter, which restarts at 0 for every run including --resume. So start_step=600 meant "600 iters after THIS resume", not "lineage iteration 600". The A/B arm pair was the clean proof: their env.yaml differed only in log_dir, only resume point distinguished them, and the omni CurriculumManager had exactly one active term - nothing else could produce that timetable. Second consequence: every shipped checkpoint (s1e-500 at +500, C2-700 at +200, A800 at +100) was selected before its run's +600, so the saturation penalty weight was 0.000 for every product ever shipped - explaining saturation 33% and raw |action| 1.9 against clip 1.0 (hip_roll in bang-bang), i.e. half the heat budget.
Change
Two independent recommendations recorded: (1) make the ramp count lineage steps (add the checkpoint's iteration offset on resume) or pin terminal weights in downstream rungs instead of ramping; (2) give the saturation penalty its own rung - never mixed into a skill-learning rung (that would be two variables again).
Outcome
Explained the recurring +600 death of the highest-amplitude command and the persistent actuator saturation of all shipped products with one root cause; honest caveat booked (at +800 the weight is only -0.086, small, but vx+0.30 is the command demanding the largest action amplitude, so it is squeezed first).
Mechanism
Resumable training splits "the lineage" from "the process"; any schedule keyed to process-local counters silently re-applies its transient to every descendant run, and any product-selection habit that picks checkpoints early systematically samples the pre-ramp regime - the curriculum exists in the config but never in any shipped policy.
Applies when
- resumed/forked training with any scheduled reward or DR ramp
- a metric dies at a fixed offset after each resume
- shipped policies show behavior a late-schedule penalty should prevent
“ramp_reward_weight 读的是 env.common_step_counter,它每个 run 从 0 开始,--resume 也不例外。… 原始 run / 臂B | 500 | 1100 = +600 | 1300 = +800;臂A | 700 | 1300 = +600 | 1500 = +800 … 所有出品其实从没见过饱和罚。… 这解释了 sat_max_pct 33%、raw |a| 最大 1.9(clip 是 1.0)—— hip_roll 一直在 bang-bang,而罚它的那一项权重恒 0。热账的一半在这里。”
train/C_LADDER_RUN.md § 3g. 系统性问题:saturation_ramp 每次 resume 归零 An exponential kernel on instantaneous velocity punishes gait oscillation - track the cycle average
cycle-average-tracking-for-gait-quantitiesReward velocity tracking on gait-cycle averages (or filtered values), not instantaneous samples, whenever the desired behavior oscillates at stride frequency; widening the kernel does not fix a variance penalty.
Symptom
Even while the robot genuinely sidewalked (verified after the metric fix), the Isaac-side tracking reward sat on the ignore-floor: true sidewalk scored 0.178 vs 0.189 for ignoring the command - the reward was mildly punishing the desired behavior.
Context
Sidewalking is inherently oscillatory: per-frame vy std was 0.177 while the tracking kernel width was sigma = 0.15, applied to the instantaneous value. A kernel-width scan showed widening sigma 0.15 -> 0.50 still loses (-0.14 -> -0.07): "指数核惩罚的是方差,而侧步天生带方差" (the exponential kernel penalizes variance, and side-stepping inherently carries variance). Modeling with measured parameters: replacing instantaneous vy with the mean over one gait cycle (0.5 s) flips the margin decisively (true sidewalk 1.888 vs ignore 1.281, +0.607), half-cycle is neutral (+0.006), two cycles adds nothing more. Explicitly flagged as extrapolation pending Isaac-side implementation. This also vindicated a previously dismissed external note (sigma too small) - right conclusion, different mechanism than claimed (variance, not gradient).
Change
Proposed fix recorded: change the tracked quantity from instantaneous vy to a one-gait-cycle running average; widening sigma alone rejected by the scan.
Outcome
Diagnosis complete and quantified; the C4 product shipped via feed-forward before the reward change was implemented, so the cycle-average fix remained a verified-by-model, not-yet-trained change.
Mechanism
E[exp(-(v-c)^2/sigma^2)] decreases with Var(v) even when E[v] = c exactly; a gait's phase-locked oscillation guarantees variance at the stride frequency, so instantaneous tracking rewards structurally prefer standing still at the command mean. Averaging over exactly one cycle removes stride-frequency variance while preserving command-following error.
Conflicts
The cycle-average fix itself is model-extrapolated ("⚠️ 这一条是外推,须在 Isaac 侧实装并复量后才能当结论") - the diagnosis is measured, the remedy untested in training at the time of writing.
Applies when
- tracking rewards for lateral/turn/any oscillation-carrying velocity
- a verified behavior scores below the ignore-floor
- choosing sigma for exp-kernel tracking terms
“侧走时 vy 的逐帧摆幅 std = 0.177,而 track_lin_vel_y_exp 核宽 σ = 0.15,且作用在瞬时值上 … 真侧走(均值 66%,振荡 ±0.18)0.178 | 完全无视指令 0.189 … 真侧走的得分比无视指令还低。… σ 从 0.15 放到 0.50,侧走仍然吃亏 … 把跟踪目标从瞬时 vy 换成一个步态周期(0.5 s)的平均 vy:… 1.888 vs 1.281”
train/C_LADDER_RUN.md § 3n. 二/三 Isaac 训练奖励为何一直坐在「无视底分」/ 修法不是放宽 σ Slowing the gait clock at deployment is out-of-distribution and backfires - lower the commanded speed instead, or train the knob
cycle-time-override-is-oodAny deployment override must correspond to a dimension the policy was trained to handle; to make a parameter field-adjustable, randomize it in training and observe it - otherwise use the levers inside the trained envelope (commands) and leave the knob alone.
Symptom
Real-robot feedback "walks very fast and unstable" suggested slowing the gait; a deploy-side --cycle-time override existed, making "just slow the clock" a one-flag temptation.
Context
A sim sweep of the override on walk_v6 @cmd 0.3 showed monotone degradation away from the trained 0.40 s cycle: at 0.50 s tilt jumped 7.9 -> 13.2 deg and landing force 1.52x -> 2.24x; at 0.80 s (half speed) clearance collapsed to 3 mm - dragging again - with 20 deg tilt. Meanwhile the legitimate lever, lowering the commanded speed with the clock untouched, improved everything monotonically: cmd 0.1 gave 104% tracking, 6.7 deg tilt, minimum slip - the most stable operating point. The file distinguishes the two "slows" explicitly: lower command = smaller steps at the same 2.5 Hz rhythm; a slower rhythm itself requires retraining - randomize cycle_time (e.g. 0.40-0.65 s) during training and expose it as an observation, and only then does --cycle-time become a field-adjustable knob.
Change
Deployment guidance: never ship a cycle-time override the policy was not trained under; respond to "too fast/unstable" with lower commands; schedule clock variability as a training-time (contract-level) change if a field knob is wanted.
Outcome
The sweep quantified the trap before hardware paid for it (dragging and 2.2x landing force at slowed clocks); cmd 0.1 documented as the stable demo point.
Mechanism
The policy is a function fitted around the training distribution; a deploy-side override moves an input (phase rate) to values never seen, so behavior degrades unpredictably - the knob LOOKS like a capability because it exists in the code, but capability lives in the training distribution, not the interface.
Applies when
- a deploy tool exposes overrides (clock, scale, gains) beyond the training distribution
- hardware feels "too fast/aggressive" and a quick knob exists
- deciding between a deploy-side tweak and a retrain
“0.80s | 1.25Hz | 0.165 | 3mm(拖地) | 20.0° … 慢一半直接崩 … 策略按 0.40 训练, 别的周期属分布外。… 降指令速度才是有效杠杆 … cmd 0.1 是最稳的工作点。… 要节奏本身变慢必须重训 —— 训练期把 cycle_time 随机化(如 0.40~0.65s)并作为观测的一维, 部署时 --cycle-time 就成了现场可调的旋钮。”
train/WALK_DIAGNOSIS.md § 2026-08-01 追加: 调慢步态时钟(--cycle-time)在仿真里是反效果 A joint-velocity penalty meant to slow the get-up cut joint speed 16% and left the get-up time unchanged - the knob never moved the variable, so the idea it was meant to test stayed untested
dof-vel-penalty-is-not-a-pacing-knobBefore reading a result as a test of an idea, check that the knob actually moved the independent variable; velocity regularizers smooth a schedule they do not set, and a schedule driven by per-step task income moves only when that income's time structure does.
Symptom
The get-up took 0.6-0.9 s in Isaac with large torque demand; the user proposed getting up more slowly so less torque would be needed.
Context
The idea had support in the accounts: the acceptance bound is an upper bound of 5 s (5-8x margin), the quasi-static squat path peaks at 25% of the limits, and rolling over needs no momentum. R3.2 raised dof_vel from -1e-3 to -5e-3 as the single variable.
Change
dof_vel -1e-3 -> -5e-3 (child-run from R3.1).
Outcome
Get-up medians moved +0.02-0.04 s (noise); raw joint velocity -16%; torque demand median got worse (hip_pitch 46-48% -> 63-67%) as the new term competed with torque_headroom on the same joints; MuJoCo 98 -> 96%. Verdict FAIL on the knob, not on the idea, and the rung was not adopted. When pace was later attacked through the income's time structure (V2.5/V2.5b), the MuJoCo get-up moved into the 3.5-4.5 s design band.
Mechanism
The pace was set by base_height_progress paying for every step spent high (stand earlier, earn more); a velocity regularizer only smooths motion along the same schedule and does not change when the robot stands up.
Applies when
- trying to make a skill slower or gentler with smoothness penalties
- an experiment's primary metric did not move and a verdict is being written
- two penalties act on the same joints
“**关键判读:`dof_vel` 罚只把关节速度压了 16%,而起身用时一点没变。** 也就是说**这一级根本没有把"慢下来"这个自变量推动起来** —— 所以它**不构成对 用户假说的检验** … 起身节奏由 `base_height_progress` 的逐步计酬决定(早站起来就多 拿),速度正则只在同一条时间轨迹上把动作抹匀,不改变何时站起来。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §25 R3.2(dof_vel −1e-3→−5e-3,慢一点起身) Record where every failed episode ends - an end-state confusion matrix showed all failures finishing seated and overturned a "cannot roll over" diagnosis that per-category success rates hide by construction
end-state-confusion-matrixFor any multi-category acceptance, report where each failed episode ends, not only which category it started in; it costs a few lines and no extra simulation, and it separates "cannot reach the goal" from "reaches the wrong basin".
Symptom
Prone scored 0% for three generations; the working diagnosis was "prone lacks the roll-over skill", and R0.3 spent a run adding prone-to-side roll-arc start states. It bought nothing: the 45-deg roll band itself only moved from 24.2% to 26.6% after 3,000 iterations.
Context
Acceptance reported success per starting category. A final-state table (lying prone / on the side / supine / seated / standing for every failed episode) was added to accept_recovery.py at R0.3.
Change
The confusion matrix became a permanent part of the acceptance output, and the prone diagnosis was rewritten from it.
Outcome
The prone, side and supine columns were all zero - every failure ended seated - and prone had righted its torso in 159/159 episodes (tilt under 30 deg in 100%). The missing ability was standing up from one specific seated configuration, not rolling over, which redirected the next rungs to foot placement and to a configuration probe.
Mechanism
Per-category success rates collapse "reached the wrong basin" and "never reached anything" into the same zero; the end state separates them.
Applies when
- a category sits at 0% and the diagnosis rests on its label
- recovery, manipulation or navigation tasks with distinct terminal states
- an intervention aimed at the presumed cause shows no effect
“**① 末态混淆矩阵 —— 固化(已在 `accept_recovery.py`)。** 它给出的 "趴/侧躺/仰躺三列全 0、所有失败都终于坐姿"是本线最改变决策的一个事实, 而**逐类成功率按构造看不见它**。 … 成本十来行、零额外仿真。 … **② prone 病因更正(旧诊断作废)。** 旧:"缺翻身"。新:**prone 159/159 全部 把躯干翻正**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §15 R0.3 判决 + 三件事的判断 Sort external training advice into adopt / already-have / modify / would-trap by recomputing it on your own config
external-advice-audit-against-own-arithmeticNever apply external tuning advice directly: recompute each claim on your own reward table and probe data, classify it adopt / have / modify / trap, and record why - and verify external citations actually exist.
Symptom
External AI/literature advice for the omni ladder arrived plausible-sounding but was written without knowledge of this robot's actual reward table, contract, and history; following it blindly would have broken single-variable discipline and, in one case, made sidewalk unlearnable.
Context
Before the C ladder, every external suggestion was audited: 3 adopted (ellipsoid command sampling; staged wz bands; command-switch acceptance), 3 already present (unified reward; frame history - the frozen 168-dim 5-frame window; per-100-iter acceptance), 2 modified (stand share kept at 20% to avoid a second variable; back share NOT raised because the probe showed backward works untrained 20/20@67%, so oversampling would only crowd out forward), and 1 flagged as a trap: "start vy very small (0.06-0.15)" - on THIS reward table vy was only an L2 tax, so ignoring a vy=0.06 command costs 0.4% of the vx tracking scale, 28-180x cheaper than ignoring forward, with quadratic shrinkage making small commands weaker still. A separate retrieval-reliability note: two search agents returned fabricated verbatim quotes from arXiv PDFs (2 papers, verified fake and discarded); only HTML/abstract/source-verifiable material was used.
Change
Advice classified only after recomputing each claim with local numbers; the "start small" trap was replaced by adding a gated lateral tracking term (the ladder's only true reward surgery) instead of shrinking the command.
Outcome
The adopted items (ellipsoid modes, staged wz, transition acceptance) entered the ladder; the trap was avoided; one external factual error (calling s1g the mainline start - it was falsified 0/20) was caught. Later, one initially-dismissed item (sigma=0.15 too narrow) turned out right for a different reason than claimed - see cycle-average-tracking-for-gait-quantities.
Mechanism
External advice encodes the advisor's reward table and robot, not yours; the transfer-validity test is whether the claim survives recomputation under your own arithmetic (reward margins, probe baselines, contract freeze). Items that survive become experiments; items that don't become documented traps.
Applies when
- incorporating LLM or literature advice into a training plan
- advice conflicts with locally measured baselines
- an external claim depends on reward-table details the advisor cannot know
“其建议 C4「先很小,vy = ±0.06~0.15」—— 在我们这张奖励表下会让侧走学不起来 … 忽略侧走比忽略前进便宜 28~180 倍,且指令越小激励越弱(平方缩放)—— "先很小"在稳定性上对、在梯度上正好把信号缩没了。”
train/C_LADDER_RUN.md § 1. 外部 AI 训练建议的评估(采纳 / 已有 / 要改 / 会踩坑) The first real-robot get-up was "very violent, kicking on the floor, dangerous" - a sim-perfect policy with no reason to be slow, unbounded absolute targets, no domain randomization and a rate limiter that filtered nothing; the task was restated as "safe, slow, transferable"
first-real-get-up-violent-stage-one-policyDo not put a get-up policy on hardware until its action is bounded (hard bound or state-anchored targets), smoothed, randomized and tested at the real pipeline's latency, and say explicitly that the task is "safe, slow and transferable" - a simulation-perfect policy optimizes only "gets up".
Symptom
On 2026-08-09 the user ran a V0-lineage recovery policy on the real robot and stopped it: very violent, kicking on the floor, dangerous. The planned next rung (a heavier torque_headroom) was never started.
Context
The spec had pre-registered that R0/R1 products stay in simulation and that the real-robot precondition was the R3 smoothing rungs plus a bridge-slew check plus a hanging protocol; the robustness (DR) rungs had not run. In simulation the policy passed 100% with a get-up of about a second. Which ONNX, which gain profile and whether a torque/joint log existed were left "to be recorded later" and never were.
Change
The V0 ladder was stopped at its best product (R3.1, sim only) and a re-rooting proposal was put to the user. The spec's four-layer account: style (the reward pays for standing early and nothing pays for slowness - HumanUP's "Stage I" get-up, "fast but unsafe ... infeasible for real-world deployment"); impact (full-range absolute targets with no hard bound, raw |a| up to 4.77, action saturation 100%, a single-step change of 0.306 saturating hip_pitch); transfer (zero DR, friction pinned at 1.0, the learned leg bracing); link (the bridge's RL slew equals vel_limit, 0.2-0.66 rad per step, while the real pipeline has 1-2 steps of time-varying latency and acceptance ran at delay 0).
Outcome
The line was re-rooted twice (training-side rate limit, then the beta-anchored action space) and gained a hang protocol before the next real attempt; on 08-11 a beta-anchored policy produced the line's first real get-up.
Mechanism
A task reward that pays for standing early selects the fastest feasible get-up; with absolute full-range targets every large target jump is a torque impulse bounded only by the clip; zero DR and braced-leg solutions do not transfer; and a limiter set at the velocity limit does nothing at 50 Hz.
Conflicts
The four layers are the spec's reconstruction from simulation probes and the literature; the real run's policy file, gain profile and log were never recorded, so no layer was confirmed against hardware data.
Applies when
- a first hardware trial of a high-effort skill is being scheduled
- sim success is high but the policy saturates actions or torques
- pre-registered hardware preconditions are not all met
“用户真机反馈:**非常猛、地上乱踢、危险**,叫停(R3.3 torque_headroom 加档已选型 weight −0.5→−1.5,未启动)。真机细节(哪个 onnx、什么档、有无 τ/q log)**待补记** … 任务从"能起来"变成 **"安全、慢、可迁移"** … **链路层**:桥层 slew RL 档 = vel_limit(10/20/33 rad/s ≈ 每拍 0.2~0.66 rad), 对 recovery 形同虚设;真机 1~2 拍时变延迟,验收默认 delay 0。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 真机叫停与换根判决(2026-08-09) Close a question with an audit, then freeze the wording - later symptoms may not reopen it without new hard evidence
frozen-verdicts-semantic-boundariesWhen an audit closes a hardware-vs-policy question, record the closing evidence, freeze a citable wording for future recurrences, and set the reopening bar explicitly; separate robustness perturbations from plant-truth questions so a DR rung's failure can never silently reopen a closed measurement.
Symptom
Recurring directional bias on the robot kept re-suggesting "maybe the hardware/COM/mechanics are asymmetric", threatening to re-litigate questions that audits had already closed - burning attention each time a descendant policy leaned or drifted.
Context
Two boundary decisions were written as permanent: (1) semantic separation - "S2④ COM ±20mm = 纯鲁棒性扰动,不再承担「解释真机后仰」任务" - if the COM-DR rung degrades, the ONLY allowed conclusion is "policy insufficiently robust to COM uncertainty"; reopening "is the CAD COM wrong" is forbidden because the mass audit was completed and closed (@63f9212). (2) a frozen wording for chirality, to be quoted verbatim whenever left/right bias appears in later rungs: observed directional bias = policy-level spontaneous symmetry breaking; plant asymmetry = no supporting evidence after the mass + model symmetry audit; mitigation candidate pi_sym queued, not blocking. The evidential basis was quantitative: the root policy was perfectly symmetric under +/-6 N*s pushes (40/40) while descendants broke (17/40, 13/40) - "手性是 S2 训练中获得的, 根没有; 机械侧已双 PASS 关案, 不重开".
Change
Closed questions carry (a) the audit commit that closed them, (b) a frozen citable wording for recurrences, and (c) an explicit evidence bar for reopening ("无新硬证据不得重开").
Outcome
Later chirality observations (C2's 15 pp turn gap, hip_roll drift bias) were handled as policy-lineage properties with policy-side mitigations, without a single hardware re-audit cycle.
Mechanism
Symptom classes recur under different guises; without a frozen verdict each recurrence re-runs the same expensive investigation and risks a different (worse-informed) conclusion. Freezing verdict plus wording converts recurring symptoms into citations, while the evidence bar keeps the closure honest rather than dogmatic - the root/descendant symmetry comparison is what makes "it's the training, not the machine" checkable at any time.
Applies when
- a recurring symptom keeps suggesting an already-audited hardware cause
- writing conclusions for a completed calibration/audit
- a DR rung's degradation invites re-measuring the plant
“若 S2④ 退化,结论只能是「当前 policy 对 COM 不确定性不够鲁棒」,不得重开「CAD COM 是不是错了」… 手性冻结表述 … Plant asymmetry: no supporting evidence after mass + model symmetry audit … 无新硬证据不得重开机械不对称”
train/OMNI_V0_SPEC.md § 4. 语义分界与手性冻结表述(2026-08-07 用户定,永久) Removing a hand trim re-exposed the plant offset it had been silently compensating - and a slope scan told bias from sensitivity
hand-trims-hide-plant-offsetsTreat hand-tuned trims as undocumented plant measurements: before deleting one, find what it compensates and re-house that knowledge in the model or the reward budget; diagnose posture errors with a sensitivity sweep to distinguish constant bias from gain problems.
Symptom
After switching from the old hand-trimmed default to the clean geometric zero, the retrained stand policy's only regression was torso lean: 1.8 deg -> 4.1 deg backward.
Context
The old default's ankle-pitch trim (-0.0489/+0.0628) had been pre-compensating a fore-aft COM mismatch; removing the trim removed the hidden compensation, and the posture reward alone was too weak to win it back. A COM sensitivity scan settled what kind of problem this was: sweeping base COM offset -50 to +50 mm gave nearly identical slopes for old and new policies (~0.026 deg/mm) - "不是质心敏感度问题, 是恒定偏置" (not a sensitivity problem, a constant bias). Fix landed in stand_v1b: posture corrected to +0.24 deg while keeping symmetry (<=0.1 deg) and low effort (0.259), disturbance rejection better than both predecessors. Model credibility was checked the honest way: v0's sim prediction at the real COM position (-22 mm) was -2.31 deg lean vs real measured 2.2-3.1 deg - "预测精准命中" - which is what licensed trusting v1b's -0.52 deg prediction. (Side flag from the same file: a sign convention had been documented wrongly in early comments - gravity_base[0] > 0 is forward lean.)
Change
Trims retired in favor of explicit modeling: symmetric geometric default plus a posture-reward budget sized to carry the real COM offset; the offset itself known (real COM ~22 mm behind model).
Outcome
stand_v1b passed acceptance as the standing lineage's final version; the walk-line requirement "加大躯干姿态惩罚权重" was upgraded from suggestion to mandatory, since walking amplifies what standing tolerates (real walk_v1 hit 26 deg lean vs sim 7.4).
Mechanism
Hand trims are plant knowledge stored in the wrong place - invisible, asymmetric, and stale after recalibration; removing them re-exposes the raw plant error. A sensitivity sweep separates the two possible diagnoses (slope change = control problem; parallel offset = constant plant bias), each with a different fix.
Applies when
- cleaning up hand-tuned offsets/trims in defaults or calibration
- a posture bias appears after a default or calibration change
- deciding whether a lean is a COM-sensitivity or constant-offset issue
“两者斜率几乎相同(≈0.026°/mm),v1 只是整体多后仰约 2.4° —— 不是质心敏感度问题,是恒定偏置。成因:旧 default 的踝俯仰 trim(−0.0489/+0.0628)本就预补偿了前后质心偏差,换成零位 default 后这份补偿没了 … v0 在真机质心处(−22 mm)的 sim 预测为 −2.31° 后仰,真机实测 2.2~3.1° 后仰 —— 预测精准命中。”
train/RETRAIN_v2.md § 4b. stand_v1 独立验证结果 / 4c. stand_v1b 验收结果 Measure yaw rate by integrating heading, not by averaging body-frame angular velocity - the two differed 15x
heading-integral-not-body-rateFor any secular rate (turn gain, drift), integrate the world-frame angle over the window; never average instantaneous body-frame rates during oscillatory motion - and when code comments warn about a measurement, believe them before re-measuring.
Symptom
Two measurements of the same turn gain disagreed by a factor of ~15: time-averaged body-frame omega_z gave -0.05 while the sim2sim harness's heading-angle integration gave +0.473.
Context
The harness code comment had already documented and predicted the failure: during gait the torso oscillates (body-frame omega_z std up to 0.7); projecting world angular velocity onto a swaying body axis and then averaging biases the estimate systematically - "实测体系均值 −0.04 而实际在以 +0.15 转" (measured body-frame mean -0.04 while actually turning at +0.15). The author's own -0.05 measurement was declared void and the training machine's 1.58/2.45 turn gains confirmed valid.
Change
Measurement doctrine fixed: yaw rate for evaluation = net heading change by integration over the window; instantaneous body-frame rates are unusable for averaged directional statistics during legged gait.
Outcome
Subsequent friction sweeps and turn-gain accounting were all conducted in the heading-integral currency, making cross-simulator comparisons (MuJoCo vs Isaac 1.04/1.02) meaningful.
Mechanism
Averaging a vector quantity expressed in an oscillating frame couples the frame's oscillation into the mean (a rectification bias); the heading integral is computed in the world frame where the gait oscillation integrates to ~zero, leaving the secular component.
Applies when
- measuring turn gain, heading drift, or any secular angular rate
- a body-frame-averaged statistic disagrees with trajectory-level truth
- writing evaluation code for oscillating platforms
“我用体坐标系 ωz 的时间均值测,得 −0.05;sim2sim 用航向角积分,得 +0.473。差 15 倍。… 步态中躯干摇晃(体系 ωz std 可达 0.7),把世界角速度投到摇摆的体轴上再取均值会系统性偏掉 … 结论:偏航率必须用航向积分,体系瞬时角速度取均值不可用。”
train/WALK_DIAGNOSIS.md § ③ 转向增益 —— 我的测法是错的,训练机的 1.58/2.45 成立 Knee swing collapsed because it directly trades against the slip penalty - price the conflict explicitly and clamp what reward cannot hold
knee-swing-vs-slip-pricingWhen a behavior collapses as another metric improves, look for the term pair trading them and set their price ratio deliberately (with escalation and reverse tripwires pre-registered); where the policy actively spends action budget to undo your target, stop paying more reward and clamp the target space structurally.
Symptom
Knee peak-to-peak swing collapsed across generations - v5 33 deg, v10 26-30, v10b 7-8, v11 6.5-8.6 - and rolling the clock back did not recover it, acquitting the clock; the collapse tracked the gated slip penalty instead: "屈膝与不打滑在当前奖励里直接对抗" - v10b's excellent 93 deg slip was purchased with knee amplitude.
Context
Reward-side flexion fixes had failed three times: raising reference amplitude backfired twice (v9/v11), and v11's deep-squat default was actively fought by the policy - it spent 0.68 of action budget pulling the squat straight ("被策略花 0.68 动作拉直反杀"). v12's design accepted the conflict as real and attacked on two tracks: (1) ECONOMICS - a direct knee_swing_amplitude reward (+0.3, target 0.55 rad, capped at 0.6/step = 55% of tracking), explicitly opposed to the slip penalty by design ("显式对立——这正是设计:v12 就是这场对抗的定价实验"), with an escalation ladder (K +0.3 -> +0.5, then slip -0.5 -> -0.3, one layer at a time) and a reverse tripwire (slip telemetry back at v10 levels -> slip weight to -0.8, accept ~20 deg knee compromise); (2) STRUCTURE - knee target bounds [0.2, 0.9] rad so full straightening is physically impossible (straightest 11.5 deg) and the 0.68 fighting budget is released. A bonus falsifiable prediction was attached: phase-lock strength tracks amplitude (v9_probe 48 deg locked 2.5 Hz; v11 low-amplitude 1.36 Hz unlocked), so if K works, hardware phase-lock should return - one change, two verdicts.
Change
knee_swing_amplitude reward + knee target clamp + pre-registered escalation/reverse levers; the failed reward-side-only approach retired.
Outcome
The lineage was frozen before v12 trained (strategic reset), but the diagnosis stands as the walk line's clearest example of two reward terms trading a behavior between them, with the pricing experiment and structural clamp fully designed and calibrated.
Mechanism
When two terms price opposite aspects of one motion (swing amplitude creates yaw momentum that becomes slip), the optimizer settles wherever the price ratio puts it - patching one side moves the equilibrium, not the conflict. Explicit pricing makes the trade a designed quantity; structural clamps remove the regions where the policy spends budget fighting the designer.
Conflicts
The pricing experiment (K vs slip) was designed and calibrated but never trained - the 2026-08-05 reset suspended v12; the collapse attribution table and the 0.68-action counterattack are measured, the remedy's效果 is untested.
Applies when
- one gait quality degrades in lockstep with another's improvement
- the policy visibly fights a default pose or reference
- repeated reward-side fixes for the same behavior have failed
“膝摆塌在 v10→v10b,头号嫌疑是门控滑移罚(四代实测膝 p2p:v5 33° / v10 26~30° / v10b 7~8° / v11 6.5~8.6°;退时钟没救回 → 非时钟)——"屈膝"与"不打滑"在当前奖励里直接对抗 … 奖励侧修屈膝已三败 … v11 深蹲 default 被策略花 0.68 动作拉直反杀”
train/WALK_V12_SPEC.md § 0. 定位 / 2. K —— 膝摆经济(与滑移罚的对偶) A single run's drift direction may be a limit cycle, not a policy bias - check the sign distribution across seeds
multiseed-sign-test-for-driftDistinguish "bias" from "broken symmetry limit cycle" by the sign distribution over many seeds; report drift as (mean, sign split), and never compare single-run drift magnitudes across versions.
Symptom
Net yaw over 15 s appeared to worsen from -41 deg (v2) to -84 deg (v4), inviting the conclusion that the new version drifted more.
Context
The Isaac-side view across 32 environments told a different story: per-env yaw was mixed-sign (20 negative / 12 positive) with mean ~0 - the drift is a limit cycle whose direction depends on initial conditions, not a systematic policy bias. The single MuJoCo run had sampled one draw from that distribution, so its magnitude could not be compared across versions as if it were a property.
Change
Evaluation rule: before classifying drift as systematic, run multiple seeds and examine the sign distribution; single-trajectory drift magnitudes are samples, not properties.
Outcome
The v2-vs-v4 drift "regression" was reclassified as not-established; later drift work (hip_roll l+r bias) used cross-policy, cross-seed evidence instead.
Mechanism
Symmetric dynamical systems can settle into either of two mirrored limit cycles; the selected cycle is decided by noise and initial state. A statistic whose sign is initial-condition-dependent has no meaning as a single sample - only its distribution does.
Applies when
- comparing heading drift or lateral drift across policy versions
- a symmetric-looking behavior shows a consistent direction in one run
- deciding whether to fix "drift" in reward or calibration
“偏航反而变差(−41° → −84°):注意 Isaac 侧 32 env 的逐 env 偏航是正负混合(20/12)、均值 ≈0,说明这是极限环性质(方向随初值)而非策略偏置 —— MuJoCo 单次跑测到的是分布里的一个样本,不能当作系统性偏差。要判断需多种子统计。”
train/WALK_DIAGNOSIS.md § walk_v4 独立验收 读法 (偏航) No parameter tuning on the floor - a failing config retries once, then it is out; anomalies go back to sim
no-field-tuning-protocolHardware time is for executing and measuring the pre-registered matrix, never for tuning: failing configs get one retry then elimination, anomalies get recorded and reproduced in sim, and contract-check bypass flags stay unused.
Symptom
Hardware sessions create pressure to fix problems live - nudge a gain, tweak a scale - which destroys attribution and risks the robot.
Context
The anomaly-handling section of the acceptance sheet is three fixed plays: (1) falls at start -> retry once at the same settings; falls again -> that configuration is eliminated, "不现场调参" (no on-site parameter tuning); (2) limit cycle or motor screech -> stop immediately, record the gain level and the joint, reproduce in sim before any discussion; (3) systematic disagreement with sim -> record it as a finding (hardware outranks sim) rather than adjusting anything to force agreement. Related guardrails elsewhere in the sheet: never pass --allow-unstamped / --allow-plant-drift to bypass manifest checks - if it errors, something real is wrong, stop and look.
Change
Field sessions restricted to executing the pre-written matrix; every fix path routed through sim reproduction and the normal config/rung process.
Outcome
Sessions stayed interpretable (each run matched a documented config) and safety overrides never became habit; anomalies arrived back in sim as reproducible cases instead of half-remembered floor stories.
Mechanism
Field-tuned values are measured under adrenaline on one floor with no logging or baselines - they contaminate the config lineage and are unattributable afterwards; and every bypass flag that skips a contract check converts a designed safety property into an operator promise.
Applies when
- a config fails or oscillates during a hardware session
- someone reaches for a live gain tweak or a bypass flag
- writing the anomaly-handling section of a deployment runbook
“起步即摔 → 换档重试一次, 仍摔则该档出局, 不现场调参。出现极限环/啸叫 → 立刻停, 记录档位与关节, 回 sim 复现再议。… 不要给 --allow-unstamped / --allow-plant-drift —— 三枚 ONNX 都已盖章 … 真要报错说明有别的问题, 停下来看。”
train/REAL_RUN_S2.md § 4. 异常处置 Decompose the offending quantity by channel first - then penalize the failure event, not the joints
penalize-the-slip-not-the-jointBefore penalizing motion to fix a side effect, measure which channels actually carry the offending quantity; prefer penalties conditioned on the failure event that are exactly zero for healthy behavior - and do not medicate behaviors that measurement shows are not sick.
Symptom
Heading drift with support-foot yaw slip (v5: 212-284 deg accumulated over 15 s); the previous v6 draft had attacked it by penalizing lateral joints (a roll 4.0 / yaw 2.0 "home" group) - which collapsed training into the standing basin.
Context
Before choosing the penalty target, the yaw angular momentum was decomposed by joint group with MuJoCo subtree_angmom weighted by real walking joint velocities: pitch-class joints (hip_pitch + knee) carry 95.3%, hip_roll 3.3%, hip_yaw 1.4%. The failed "home" group had been taxing 2.7/step to manage a 4.7% channel. The replacement, feet_yaw_slip (-0.2, |support-foot yaw rate| while in contact), targets the failure event itself and - decisively - costs a non-slipping gait exactly zero, which "横向回家组做不到". The same rung's do-not-do table applied the complementary principle to foot spacing: measured 196-214 mm, stable, no crossing - "没病不吃药" (no disease, no medicine).
Change
Removed joint-usage penalties for the drift problem; added the event-conditional slip penalty (-0.2, realized tax 0.141/step = 12% of tracking) alongside the existing linear-slip term.
Outcome
Turn-gain left/right difference improved 70% -> 19% and heading 185 -> 60.3 deg by v6 without a standing-basin collapse; the 2.7/step lateral tax never returned.
Mechanism
Penalizing joints taxes every use of a channel including healthy use, and if the channel carries little of the offending quantity the tax buys nothing while pushing the optimum toward immobility. An event-conditional penalty (slip while in contact) prices only the failure, leaving the healthy gait's cost surface untouched - and the channel decomposition tells you in advance whether a joint-side fix can even work.
Applies when
- choosing a penalty target for drift/slip/impact problems
- a proposed penalty taxes joints or motions rather than failure events
- a previous joint-penalty attempt collapsed the gait
“pitch 类 (hip_pitch + knee) 占偏航角动量 95.3% … hip_yaw 1.4% … 压 hip_yaw 是管 1.4% 的通道收 2.7/步 的税 —— 上一轮正是这样把策略推进了站立盆地。滑移项不惩罚走路: 不打滑的步态代价为零, 这是横向"回家"组做不到的。”
train/WALK_V6_MINIMAL.md § ① / ② 新增 feet_yaw_slip The real robot's right-leg kicking was over-trained-delay times loop gain - irreducible pipeline latency is plant, model it fully from day one
pipeline-latency-is-plant-not-drMeasure the end-to-end action pipeline delay and build it into the nominal plant and every acceptance gate from day one; treat power/scale deratings that "fix" oscillation as gain-reduction crutches flagging an unmodeled delay, and expect higher-feedback-gain policies to be MORE delay-fragile.
Symptom
On hardware, s1c/s1d at action scale 1.0 always kicked wildly with the right leg (s1c only ran as SOTA at power 0.8; s1d only at 0.7) - while sim showed nothing under default evaluation.
Context
Sim reproduced the incident item by item once the real pipeline delay was injected: s1d@1.0 with --delay 1 fell at 10.2 s, --delay 2 at 5.2 s; s1c@1.0 stressed (r_hip_roll saturation 5 -> 16%; "右脚" = the policy's chirality makes the right leg its high-gain leg); and the combos that worked on hardware (s1c@0.8+delay2, s1d@0.7+delay2) all survived in sim. Mechanism: the real pipeline is ~1-2 ticks (BusWorker next-cycle pickup + CAN round trip) but S1.1 trained only to 1 tick - "超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定" (delay beyond training x full loop gain = oscillation; the power derating had been buying stability by compressing loop gain). s1d was MORE fragile than s1c because its yaw 3-layer stack had learned higher feedback gain - higher gain, lower delay tolerance. Three changes: latency DR widened to cover reality; acceptance gates and smoke runs moved permanently to --delay 2 ("门必须在真机条件下预测 真机"); and the doctrine written twice-paid: "不可约的管线属性(延迟、 限速)不是'随机化选项',是 plant 本体,第一天就该全额建模" - S1's nominal-then-robust staging falsified by hardware for the second time. The later s1e hardware run at power 1.0 (no kicking, normal force) closed the loop: "0.8 = 旧代拐杖" - the derating had been a crutch for the under-modeled delay, not a real requirement.
Change
Latency modeled as plant from day one of any lineage (measured 1-2 ticks covered, bridge-layer rate limits likewise modeled by default); every gate and smoke evaluation issued under --delay 2.
Outcome
Kicking reproduced, explained, and eliminated in the s1e generation at full scale and full power; the deploy-side crutches (0.7/0.8) retired for the new lineage.
Mechanism
Feedback oscillation onset is a product of loop gain and phase lag; a policy trained below the real delay learns gains that sit past the real stability margin, and any output derating masks it by scaling gain down. Since pipeline delay is deterministic hardware property - not an uncertainty - it belongs in the nominal plant, and every evaluation must include it or the gate predicts a robot that does not exist.
Applies when
- hardware oscillation/kicking that sim only reproduces with added delay
- a policy only runs on hardware at reduced power/scale
- defining what belongs in the nominal plant vs the DR list
“真实链路延迟 ~1~2 拍 … S1.1 只训到 1 拍——超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定。s1d 比 s1c 更脆 = yaw 三层栈学出更高反馈增益,增益越高延迟容忍越低。… 教训入账:S1「先标称后鲁棒」第二次被真机证伪——不可约的管线属性(延迟、限速)不是"随机化选项",是 plant 本体,第一天就该全额建模。”
train/OMNI_V0_SPEC.md § 3. S1.4(真机右脚乱踢事故强制) After installing the measured plant, re-run all generations paired on old and new plant - identity metrics carry verdicts, physics metrics re-baseline
plant-swap-invariants-vs-shiftsTreat every plant upgrade as an era boundary: re-run the retained policy set paired (same seeds/flags) on both plants, carry forward only verdicts whose metrics proved plant-invariant, re-baseline the rest - and mine the systematic shifts as measurements of the old plant's biases.
Symptom
With the plant finally fully measured (weighed masses 9.792 kg, bench-identified armature, in-situ friction), no historical sim number was comparable to new runs - "历史 sim 数字跨纪元不可比" - and it was unknown which historical verdicts still held.
Context
The era-2c cross-test ran all seven walk generations on the complete plant under one harness (14/14 survived), then re-ran the same 14 configurations on the OLD plant retrieved from git, same flags and seeds, with a self-check (one historical record reproduced digit-for-digit). The split was clean. Policy-identity metrics moved essentially zero across the plant swap - dominant frequency (v7's period-doubling 1.20 -> 1.20), knee amplitude (9.9 -> 10.0), foot distance (+/-2 mm), saturation (100 -> 100, 0 -> 0) - so all seven cross-generation verdicts (freeze signature, saturation-line closure, knee-collapse location, slip-penalty accounting, N2 lineage, v6's balance, v11's triad) were re-confirmed on the honest plant. Plant-physics metrics shifted systematically with ordering preserved: slip down 10-25% (measured friction makes ground-twisting costlier), landing vertical velocity down 15-50% ("旧 plant 高估落地 凶度"第二次独立证实), landing force mixed (mass up 2.2% vs friction braking the swing - two effects fighting). Exactly ONE behavior-level change: v8's low-speed period-doubling vanished (1.30 -> 2.50, lift normalizing) - confirming it had been machine-dependent bifurcation-edge behavior that armature+friction push off the knife edge, while v7's period-doubling stood untouched: saturation-freeze-driven, a policy property, not a numerical accident.
Change
Era re-baselining protocol: after any plant upgrade, one paired same-seed sweep of all retained generations on old and new plant; verdicts keyed to identity metrics carry over, thresholds re-read against the new-plant table, and the differences themselves become plant-physics findings.
Outcome
Seven verdicts survived with evidence rather than assumption; two causes of the period-doubling family were separated with plant-side proof; and the sim's landing-violence overestimate was independently confirmed a second time.
Mechanism
A policy's structural properties (frequencies, amplitudes, frozen joints) are functions of its weights and survive plant changes; contact-mediated quantities are joint properties of policy and plant and shift when the plant becomes honest. Pairing seeds across plants isolates the plant's contribution exactly, so the sweep both validates history and measures what the old plant had been lying about.
Applies when
- installing measured masses/armature/friction into the sim
- historical thresholds are cited across a plant change
- a hardware-only behavior might be bifurcation-edge sensitivity
“策略身份指标逐位不动:主频(v7 1.20→1.20)、膝摆 … 这些是策略属性,plant 换代携带无损,历史定论因此全部成立。… 唯一行为级变化:v8@0.15 的倍周期消失(主频 1.30→2.50…)——印证当时"分岔边缘、机器相关"的判定:armature+摩擦把 v8 推离刀锋;v7 的倍周期纹丝不动(1.20→1.20),它是饱和冻结驱动的深层属性,不是数值巧合。”
train/README.md § 纪元 2c 全代同机横测 (2026-08-04): 完全体 plant 上历史结论全部存活 Deployment power derating damages non-forward axes far more than forward - sweep it in sim before deploying
power-scale-hurts-nonforward-axesTreat deployment power/torque scaling as a plant parameter: evaluate the policy in sim at the exact deployment scale, expect non-dominant axes to degrade first under derating, and either deploy at the training power or train with power randomization.
Symptom
Policies deployed at power-scale 0.8 (a safety derating of commanded torque) looked fine walking forward but were weak at backward and turning, inviting the wrong diagnosis "the skill was not trained well".
Context
Measured repeatedly: on s1e, going 1.0 -> 0.8 cost forward 18% but backward 58%; on C4-ff800, turn tracking was +25%/+40% at pw0.8 vs +75%/+58% at pw1.0, backward 51-52% vs 97-103%, while forward stayed 96-98% at both. Sim evaluation numbers in the plan were all pw1.0, but the robot was being run at 0.8.
Change
Pre-deploy protocol added: sweep the exported policy across power in sim (for PW in 0.8 0.9 1.0: eval_c_matrix --power $PW --seeds 20) and deploy at the first level where both turn directions reach >=50%. For C4 the recommendation was raise the robot to pw1.0 - the sweep showed it nearly free (saturation 47%->33%, left foot-clipping danger zone 25%->6%, cost only tilt 6.7->8.3 deg).
Outcome
Turning "weakness" resolved without any retraining; the sim sweep correctly predicted the real-robot signature at both power levels.
Mechanism
Forward walking is the reward-dominant, torque-cheapest skill with the most margin; backward/turn/sidewalk live closer to the torque envelope, so a uniform torque derating consumes their margin first. Training ran at power 1.0 (the trainer does no power scaling), so deploying at 0.8 is a systematic underactuation the policy never experienced.
Applies when
- deploying with any torque/power derating or safety scale
- secondary skills (backward, turn, lateral) underperform on hardware while forward walking looks fine
- choosing the deployment power level for a new policy
“power 衰减对非前进轴的伤害远大于前进轴(s1e:前进 1.0→0.8 掉 18%,后退掉 58%)。转向是非前进轴,0.8 下很可能明显跟不动。”
train/C_LADDER_RUN.md § 3c. A-2 上机前先定部署力度档 / 3p. 二 Prone get-up sat at 0/159 until two gated hinge terms moved the seated feet - first sideways (561 -> 360 mm), then fore-aft (-168 -> +56 mm) - and success went to 158/159 with nothing else changed
prone-dead-end-is-foot-placementWhen a stuck state and a successful state differ geometrically, penalize the discriminating quantity with a gated hinge that is exactly zero in the state the policy actually reaches (measure it - not the nominal), then re-probe: flattening one axis can move the discriminant to another.
Symptom
Prone falls always righted and then sat with the feet splayed wide or tucked behind the hips, from where the policy never stood (0% for four generations).
Context
Four lines of evidence pointed at foot position: the configuration probe (ankles 215 mm apart stood 52.3%, 561 mm apart 0.0%); FK showing the action contract's nominal (a = 0) is itself a 465 mm straddle, so the action_rate and still terms were pulling toward the splits; biomechanics (feet tucked under the body cut peak hip-extension torque 148.8 -> 32.7 N*m, -78%); and HoST's foot-displacement term, which this reward table lacked. The earlier "not a reward hole" reading was corrected to "a gradient hole, not a level hole": at the dead point the heaviest term (upright) was saturated with zero gradient, still paid for not moving, and the one live gradient (base_height) pointed at the thigh-horizontal torque barrier. A prone ROM scan had already ruled out pushing up from prone.
Change
R0.4: feet_spread_excess = clamp(ankle distance - 0.215, 0, inf) x upright gate, weight -2.0, plus a height-decay factor added after measuring that the policy's real standing stance was 406 mm, not the 215 mm nominal (the plain version would have taxed every successful stand 0.38/s). R0.5: the same shape on the fore-aft axis, feet_fore_seated = |fore-aft offset - 0.05| x upright gate x height decay, target +50 mm (the measured natural offset of standing postures). One variable per rung.
Outcome
R0.4: seated ankle distance 561 -> 360 mm, supine/side exactly unchanged, prone 0 -> 1.9%, mid 45.9 -> 62.2%; a probe then showed the discriminant had moved to the fore-aft axis (standing starts +42 to +51 mm, the prone seat -168 mm). R0.5: supine 99.4, prone 99.4, side 100, mid 100%, re-falls 0%; both geometry terms collapsed to ~0 near iteration 13,100 as base_height rose, and the prone fore-aft offset went -168 -> +56 mm - the term's own target, closing the causal chain. The cost, unmeasured at the time: action jitter rose 33% (sum |da|^2 6.82 -> 9.06).
Mechanism
An upright-gated hinge is inert while the robot rolls and exactly zero in the achieved stance, so it adds gradient only inside the stuck basin; a seated robot with its feet behind or outside its COM must make a kinematically unfavourable transition to stand, and moving the feet under the body removes it.
Applies when
- a get-up or transition skill fails from one start category only
- successful and failed episodes differ in a measurable geometric quantity
- a shaping term might tax the posture successful episodes already use
“`recovery_r0_5`,唯一变量 = 追加 `feet_fore_seated`(与 R0.4 同形状,只换测量轴)。 … 对照 R0.4 的 prone(3.1%,360 mm,**−168 mm**):前后偏移从 −168 走到 +56, 正是这一项的目标量,**判别量被消掉后成功率随之到顶** —— 因果链完整。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §20 R0.5(前后向脚位置项):成功率门全过 —— prone 0/159 → 158/159 A ratio metric flipped the verdict - spectral share rose while absolute high-frequency energy fell 16%
ratio-metrics-need-absolute-checkNever compare share/percentage/centroid metrics across conditions whose totals differ; pair every ratio with its absolute numerator before issuing a verdict, and log retracted judgments so they are not re-derived.
Symptom
walk_v6 was provisionally judged "more jittery" than v5 because the joint-velocity spectral centroid rose 2.91 -> 3.62 Hz and the >4 Hz energy share rose 12.6% -> 17.8%.
Context
Absolute measures said the opposite: first differences of actions fell 1.54 -> 1.31, second differences 2.55 -> 2.19, and absolute high-frequency energy fell 16% (0.490 -> 0.413). The shares and centroid rose only because low-frequency content fell even more - the denominator shrank. The interim judgment was retracted in writing so it would not be reused.
Change
Metric discipline noted: "占比类指标在总量变化时不能直接比较" - share/ratio metrics are not comparable across conditions when the total changes; verdicts about smoothness must cite absolute energies or difference norms.
Outcome
v6 correctly classified as smoother, not jitterier; the retracted judgment logged under "被推翻的一个中间判断(记下来免得复用)".
Mechanism
A ratio confounds numerator and denominator; any intervention that removes low-frequency content raises every high-frequency share without adding a single joule of jitter. Only absolute quantities support cross-condition comparison when totals move.
Applies when
- comparing smoothness/jitter/spectral metrics across versions
- any percentage-based metric moves after an intervention
- writing an eval report that includes normalized quantities
“我一度说"v6 动作更抖" … 错了: 动作一阶差 1.54→1.31、二阶差 2.55→2.19 都在降 … 谱质心升高只是因为低频成分掉得更多, 绝对高频能量实际下降 16%(0.490→0.413)。占比类指标在总量变化时不能直接比较。”
train/WALK_DIAGNOSIS.md § 过程中被推翻的一个中间判断(记下来免得复用) Audit rewards by realized contribution (weight x achieved value) - a weight of 2.0 was really paying 0.04
realized-contribution-auditEvaluate a reward table by each term's realized per-step contribution under the current policy, never by its weight column; if a term's realized value is ~0, escalating its weight is a no-op - change the term's structure instead.
Symptom
Foot dragging persisted through repeated weight escalation: feet_air_time had been raised 0.25 -> 1.0 -> 2.0 across versions with no behavioral change, and the training-side comment even recorded the fact ("几乎没有 单支撑相, 是拖着脚蹭") without the fix changing form.
Context
Computing realized per-term contributions in the trained state exposed the economy: feet_air_time contributed weight 2.0 x achieved 0.019 = 0.038 per step, against tracking's +1.20 - lifting the leg earned 3% of what tracking earned, so dragging was the rational optimum no matter the weight escalation. The same table acquitted the energy penalties (total -0.30 negative vs +1.74 positive) that a naive read of weights (-5.0 orientation!) would have blamed.
Change
Fix redirected from "raise the weight again" to "add a term whose realized contribution changes the optimum": a clearance penalty sized so its realized magnitude (~0.018/foot when dragging) is comparable to feet_air_time's, enough to flip the optimum without drowning tracking.
Outcome
With the term economy corrected (plus posture/range fixes), swing height reached 34 mm and tracking 87% by v6; weight escalation of the old term was abandoned.
Mechanism
A reward weight is only a multiplier on whatever the policy currently achieves on that term; when the achieved value is near zero (behavior absent), escalating the weight multiplies near-zero. Optimizer behavior is governed by realized per-step magnitudes, so audits must be conducted in that currency.
Applies when
- a behavior persists despite repeated weight increases
- auditing whether penalties are "too strong" or rewards "too weak"
- sizing a new reward term against existing ones
“把 feet_air_time 权重从 0.25 → 1.0 → 2.0 一路加,但没有加高度项。量级算下来:feet_air_time 权重 2.0 × 实得 0.019 = 0.038,而跟踪奖励是 1.2。抬腿的边际收益只有跟踪的 3%,拖地当然是最优解。… 正项 +1.74,负项 −0.30。能量惩罚不是瓶颈,抬腿没收益才是。”
train/WALK_DIAGNOSIS.md § 决定性证据(#1) / 各项奖励的实际量级 A field experiment is allowed when it is pre-scripted, single-variable, and self-reversing (RAM-only writes)
reversible-single-variable-field-experimentsPermit hardware-side experiments only when scripted in advance with one variable, a log, and automatic reversion (volatile writes, git restore); keep every config mirror (yaml vs firmware) changed and restored as a unit.
Symptom
A hypothesis needed a hardware test - v5's "wild kicking" might trace to the RS00 torque cap being deployed at 11 N*m vs its trained 14 (-21%), on exactly the ankle-roll/hip-yaw joints doing lateral-yaw work - but changing limits in the field is the classic way to lose track of robot state.
Context
The experiment was written to be safe by construction: change exactly one number (robot.yaml RS00 tau_limit 11.0 -> 14.0), write to firmware RAM only (--write without --save, so a power cycle automatically rolls back to the saved 12/17/11), run the single logged trial, then restore both yaml (git checkout) and RAM immediately. The consistency requirement is explicit: the deploy tool's torque self-check compares against robot.yaml, so yaml and firmware must change and restore together; and the hypothesis scoping is itself single-variable - RS06's 67% cut was measured irrelevant (gait uses only 16% of rated) and hip torque was left alone for safety.
Change
Field experimentation policy refined: not "never touch hardware settings" but "only pre-scripted, one-variable, logged, auto-reverting changes with config/firmware kept consistent".
Outcome
The sub-experiment could answer the torque-cap hypothesis without any risk of the robot persisting in an undocumented state - forgetting to restore costs nothing but a git checkout.
Mechanism
The danger of field changes is state divergence (robot config drifting from the repo's record), not the change itself; volatile (RAM-only) writes bound the divergence lifetime to one power cycle, and single-variable scoping preserves attributability even in a field setting.
Applies when
- a hypothesis requires changing firmware limits or gains on the robot
- field debugging tempts persistent config writes
- designing safe escape hatches for deployment tooling
“程序(--write 不带 --save = 只写 RAM,断电自动回滚)… deploy 的限扭自检是对 robot.yaml 比对的,所以 yaml 和固件必须同改同还原;忘了还原也没事,断电重启即回 12/17/11(上次 --save 的值),但 yaml 要 git checkout。”
train/REAL_SWEEP_V5_V8.md § 4. 限扭子实验(可选二期,只对 v5,单变量) Training-log reward values and fixed-command eval values live on different distributions - comparing them once claimed a 44% improvement that was really 6-10%
same-distribution-reward-comparisonQuote reward-term values only with their distribution attached (command range, DR on/off, environment), and compare across runs only when those match; re-measure in a common environment before claiming any improvement percentage.
Symptom
A v6-era analysis concluded slip had dropped 44% by comparing the training log's Episode_Reward against values calibrated in a fixed-command play environment; a same-condition re-measurement showed the true improvement was 6-10%.
Context
The training log's reward is an expectation over the training command distribution (vx 0.15-0.5, yaw +/-0.6, with pushes and domain randomization); play-environment calibrations are taken at a single fixed command with DR off. Subtracting one from the other compares apples to oranges - the warning was written into the v7 pre-flight: "奖励数值只能在同一指令分布下比较 … 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次)".
Change
Rule adopted: any before/after reward-term comparison must hold the command distribution, DR state, and evaluation environment fixed; training-log values compare only against training-log values of runs with identical command/DR configs.
Outcome
The phantom 44% improvement was retracted; later term-level accounting (e.g. the C4 ignore-floor work) consistently specified its distribution before quoting numbers.
Mechanism
A reward term's expectation depends on the visited-state distribution as much as on the policy; changing the command distribution or DR moves every term's baseline. Cross-distribution differences therefore measure the distributions, not the policy change.
Applies when
- comparing reward telemetry across training runs or vs play evals
- claiming improvement percentages from training logs
- term-level reward accounting for diagnosis
“奖励数值只能在同一指令分布下比较。训练日志的 Episode_Reward 是在训练指令分布上算的(vx 0.15~0.5 / 偏航 ±0.6 / 带推力与域随机化), 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次: 据此以为滑移降了 44%, 同条件对拍只有 6~10%)。”
train/WALK_V7_SPEC.md § 3. 开训自查 ⚠️ A nonzero response with the same sign for + and - commands is bias, not ability
same-sign-response-is-yaw-biasBefore crediting any directional skill, test both command signs: response must flip sign with the command; a same-signed pair is a bias to subtract, not an ability to report.
Symptom
Root-selection probe showed nonzero wz "tracking percentages" on turn commands, tempting the read that candidates could partially turn.
Context
During C-ladder root selection, s1e-500's measured yaw rate was +0.084 rad/s for cmd +0.3 and +0.093 rad/s for cmd -0.3 - same sign both ways. The same check on the C2 baseline gave wz+0.20 -> -0.13 and wz-0.20 -> +0.12 (again same sign), while the alternative root s2e_pd-1400 gave +0.16 / -0.16 - opposite signs, i.e. a genuine 16% command response.
Change
Reading corrected and written into the execution sheet: percentages on directional commands are meaningless unless the +cmd and -cmd responses have opposite signs; all three candidates were re-classified as "cannot turn, cannot sidewalk - C2/C3/C4 learn from zero". Acceptance criteria thereafter required "tracking >=50% AND left/right opposite-signed".
Outcome
Prevented crediting turn/sidewalk ability that did not exist; the antisymmetry clause became a standing part of every turn and sidewalk PASS condition (C2, C4, C4-redo levels all carry "且左右反号").
Mechanism
A constant yaw (or lateral) bias projects onto any command's sign convention and shows up as fake fractional tracking; only sign-antisymmetry under command reversal distinguishes a feedback response to the command from an open-loop offset.
Applies when
- evaluating turn/sidewalk/any signed-command tracking percentages
- a candidate shows partial tracking on an axis it was never trained on
- writing PASS criteria for a new directional skill
“C2/C3 那些非零的 wz 百分比不是转向能力 —— 转向+ 与 转向− 的实测同号(s1e:cmd +0.3 → +0.084,cmd −0.3 → +0.093 rad/s),那是恒定偏航偏置。… 三个候选都不会转、都不会侧走。”
train/C_LADDER_RUN.md § 0. 读数纠正(重要,别引错) A get-up policy righted itself and sat - three terms paid the seated pose 84% of the return, and the only shaping term that could tell sitting from standing was an exp kernel outputting 5e-5
seated-basin-dead-exp-kernelWhen a policy parks in a degenerate posture, tabulate what each reward term pays that posture against the target (watch contact terms that reward touching rather than bearing load) and evaluate every exp kernel at the error actually observed; a kernel narrower than the real error is switched off, and widening it is a one-variable repair that adds nothing new.
Symptom
R0 converged by iteration 700 and gained 1.7% over the next 2,300; success 0.0% in all four fall categories. Three of four success conditions passed (tilt median 1.0 deg, both feet in contact 99.6%, angular rate low); height passed 0.6% (median 0.204 m against 0.326). The robot knelt in a W-sit: hip yaw +/-47 deg, knees folded to 92% of the hard limit, shins flat, pelvis on the ground, torso vertical.
Context
Minimal reward table: upright (1-g_z)/2 +2.0, base_height linear progress +1.5, stand_pose exp(-||q-q_stand||^2/std^2) x upright gate +1.0 with std 1.0, still +0.5 and feet_on_ground +0.5 both x the upright gate, plus regularizers. The upright gate is a hinge that opens below 30 deg of tilt. The pre-registered fallbacks were then checked against the measured state: tightening the tilt gate was falsified (tilt was already 1.0 deg); a success bonus contradicted the spec's own no-cliff-bounty rule; narrowing the categories was useless (all four converged to the same pose); raising init noise was too weak for a basin this deep. Only a half-rise intermediate state addressed it, and a cheaper repair existed.
Change
R0.1 (user decision, single variable): stand_pose std 1.0 -> 3.0. Not a new term and not a bounty - repairing a declared term that was numerically dead. The two runs' logged env.yaml differ in log_dir and std only.
Outcome
R0.1 58.2% overall (R0 0.0%): supine 91.5%, side 82.2%, mid 55.9%, prone 0/156; knees fully straight; ||q-q_stand||^2 9.99 -> 0.91 and the stand_pose term 4.6e-5 -> 0.90; get-up ~1 s, no re-falls, the curve still rising at the 3,000-iteration cap. Prone stayed at zero and needed a different fix (see prone-dead-end-is-foot-placement).
Mechanism
Sitting earned upright 1.98/2.0, still 0.43/0.5 and feet_on_ground 0.46/0.5 - 3.0 of a 3.57 per-second return - because feet_on_ground asked for contact, not load. The only terms separating sitting from standing were base_height (+0.70/s for standing) and stand_pose, whose kernel at the real 9.99 rad^2 error (75% of it in the two knees) was exp(-9.99) = 4.6e-5 with a gradient near 1e-4. Standing up meant risking 3.0/s to gain 0.70/s while unfolding knees at 92% of their limit under load. With std 3 the same term is exp(-9.99/9) = 0.33 - a live gradient, three quarters of it on the folded knees.
Applies when
- a policy converges early to an upright but low, seated or kneeling pose
- a posture-matching exp term reads ~0 in the training logs
- contact-based rewards saturate while the task metric does not move
“**关键:`feet_on_ground` 只问"触地"不问"承重", 跪坐时双脚确实贴地,照样满分。** 三项 3.0/s = 总回报 3.57/s 的 84%。 … **exp(−9.99) = 4.6e-5** —— 权重 1.0 的项实际输出 5e-5、梯度 ~1e-4, **不是"还没学会",是数值上根本不存在**。 … **R0.1 决定(用户 2026-08-09 定,单变量)**:`stand_pose` 的 `std` **1.0 → 3.0**。 不是加新奖励、不是悬崖悬赏,而是**修复一个已声明但数值失效的项**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §11 R0 首跑(recovery_r0, 2026-08-09):FAIL —— 翻正了但坐着 A sim veto needs real confirmation too - the worst sim cell was scheduled as the most informative hardware run
sim-veto-needs-real-confirmationNever let sim alone both condemn a purpose-built configuration and escape audit: spend one cheap, safeguarded hardware run on the condemned cell, pre-registering what agreement and disagreement would each imply about the proxy.
Symptom
The fric-2400@kd1.0 combination was sim's worst cell across the board (survival 17/20 - the only miss, mu0.4 1/20, push 103/160, zero-cmd 2/20), yet it was the only product specifically trained for the kd1.0 deployment gain - discarding it on sim evidence alone would leave the sim's own validity untested exactly where it mattered.
Context
The team had been burned in the other direction before ("Isaac 指标三次 零预警" - training-side metrics gave zero warning three times), so the symmetric rule was written: sim's rejection also needs hardware confirmation ("sim 判被支配 ≠ 真机被支配 … sim 的否决也要真机确认"). The run was pre-registered with a dual reading: real matches sim -> the S2f ladder closes and the fork root is settled; real clearly better than sim -> the MuJoCo proxy has a systematic bias in the kd1.0/low-margin region, "那比选型本身重要得多" - and every S2f sim acceptance would need re-scoring.
Change
The condemned configuration was kept on the hardware roster (last, spotted, minimal exposure) explicitly as a proxy-validation probe, not as a deployment candidate.
Outcome
Session design captured either result as progress: selection confirmed, or a proxy bias discovered that would re-price the whole ladder's verdicts.
Mechanism
Every sim verdict is a joint statement about the policy AND the proxy; cells where a policy was purpose-trained for the exact condition sim condemns are where proxy error is most likely and most costly. Testing the veto converts a selection decision into a calibration measurement of the evaluator itself.
Applies when
- sim rejects the configuration that targets the actual deployment condition
- the eval proxy's calibration has never been checked in that regime
- deciding which hardware runs are worth their risk
“但它也是唯一为 kd1.0 部署档专门训的产物 —— sim 判被支配 ≠ 真机被支配, 「Isaac 指标三次零预警」的教训反过来同样成立: sim 的否决也要真机确认。… 若真机明显好于 sim → MuJoCo 代理在 kd1.0/低裕度区有系统性偏差, 那比选型本身重要得多。”
train/REAL_RUN_S2.md § 上机名单 note / 2. sim 侧预注册预期 ⑤ Single-impulse push recovery is a binary chaotic quantity - cross-machine floating-point divergence can flip the outcome
single-impulse-recovery-is-chaoticNever gate or compare single-event recovery outcomes across machines or domains: evaluate disturbances as survival distributions over phases and seeds, compare longitudinally on one machine, and treat any single-point cliff as unconfirmed until it survives the statistical protocol.
Symptom
Mac evaluation found a hard "0.8 N*s cliff" (0/3 survival) that the training machine flatly contradicted: the identical protocol (0.8 impulse at 8 s, cmd 0.2) survived 3/3 there, and a 0.6/0.8/2/4 cross sweep survived everything.
Context
The verdict became a named lesson ("跨机混沌课文"): whether one specific push at one specific phase is survived depends on a trajectory that diverges across machines from floating-point differences alone - "单次冲量恢复是二值混沌量, 跨机浮点发散可翻结局". The boundary was drawn precisely: the 20-seed statistical gates DO agree across machines (established precedent), but that agreement cannot be extrapolated to single-point recovery tests. Protocol amended: disturbance evaluation uses multiple push phases (8/10/12 s), >=10 seeds, and only same-machine longitudinal comparisons; the Mac-side recommendation built on the unreproducible cliff was not adopted, while its directionally-consistent small-impulse data was kept.
Change
Push evaluation redefined from single-event pass/fail to multi-phase multi-seed statistics, with cross-machine comparison banned for event-level results and allowed for distribution-level ones.
Outcome
A false hardware-relevant "cliff" was prevented from steering the ladder (the s2e push rung decisions were made on same-machine statistics); the chaos lesson was cited again when real push tests were restricted to qualitative cross-domain use.
Mechanism
Perturbation recovery near the viability boundary has sensitive dependence on initial conditions; different BLAS/GPU reduction orders yield different trajectories from identical configs, so a binary outcome at one phase is machine-specific noise. Averaging over phases and seeds restores a quantity whose expectation is machine-stable.
Applies when
- a push/disturbance result differs between machines or sim and real
- designing push-recovery acceptance tests
- a sharp pass/fail cliff appears in a chaotic-regime evaluation
“训练机上 Mac 原协议 (0.8 @8s cmd0.2) 3/3 全活 … 与 Mac 的 +0.8 0/3 直接矛盾。定性: 单次冲量恢复是二值混沌量, 跨机浮点发散可翻结局;统计门 (20-seed 八门) 跨机吻合的先例不能外推到单点恢复测试。协议改判: 抗推评测多相位 (push 时刻 8/10/12s) + ≥10 seed + 只做同机纵向比”
train/README.md § s2e 支线终章 (跨机混沌课文) A real-robot verdict is (policy x deployment stack) - when the stack changes materially, old verdicts expire
stale-verdicts-under-old-stackDate every hardware verdict with the deployment-stack version it was measured under; after any material stack change, re-test before trusting old condemnations or old praises - with the interpretation of each possible result written down first.
Symptom
walk_v5 stood condemned as "kicks wildly" and walk_v6 as "cannot walk unassisted" - but those verdicts were issued under an earlier deployment stack (--heading did not exist yet, several fixes had just landed); only v7 had ever run under the current unified stack.
Context
New sim evidence sharpened the doubt: a same-harness four-version sweep showed v6 was the HEALTHIEST archive at the real operating point (cmd 0.15: dominant frequency locked at 2.50, steps 35:35 perfectly symmetric, foot distance 203/190 mm best of four, mean tilt 4.2 deg, saturation 15%). A full re-test under the unified stack was scheduled with per-version questions and a pre-filled interpretation table ("判读表(预填假设,回来对号)"): e.g. v6 walks + frequency ~2.5 -> old verdict was the stack's fault, v6 becomes the comparison champion; v6 walks but at ~1.25 -> period-doubling on hardware = confirmed plant gap, actuator fitting promoted to mainline; v5 no longer kicks -> the kicking was an old-stack artifact.
Change
All four versions re-queued on hardware under one stack (same torque limits, slew profile, heading loop, logging), with the version-specific legacy profile pinned; verdicts held provisional until re-issued.
Outcome
The re-test design separated policy properties from stack artifacts before any policy was permanently written off - and turned each outcome into a specific conclusion via the pre-filled table.
Mechanism
A deployed behavior is produced by the policy plus everything between it and the motors (heading loop, slew limits, torque caps, clock); verdicts implicitly condition on that whole stack. Fixing the stack invalidates the conditioning, so old failures may be stack artifacts and old successes may not survive either.
Applies when
- deployment tooling (limits, filters, loops) changed since a policy was last judged
- deciding which historical policy is the rightful baseline
- a sim sweep contradicts an old hardware verdict
“只有 v7 在完整的今日部署栈下上过真机 … v5"左右乱踢"、v6"未能自主"的判决全部来自更早的栈(--heading 尚不存在, 部分修复刚落地)——判决已过期。且 2026-08-02 四代同机仿真横测翻出了新证据:v6 在真机工况(cmd 0.15)下是四代里最健康的仿真档案”
train/REAL_SWEEP_V5_V8.md § 0. 为什么重测 A suspended (no-load) test acquits or convicts the actuator before you blame authority
suspended-test-isolates-actuator-authorityBefore attributing a failure to actuator authority, measure no-load tracking error and steady-state torque fraction; blame authority only if the task fails while the error grows with demanded force - and then fix gains or targets, not training.
Symptom
hip_roll sagged 0.21 rad on the ground and saturation questions loomed over the sidewalk plan - was the roll axis physically too weak (authority), or was something else limiting it?
Context
Before C4, the roll-authority question was settled by measurement triage: suspended test (--suspend, feet off ground) showed hip_roll tracking error 0.0008 rad - actuator acquitted; the entire 0.21 rad ground sag is load-induced. Steady-state torque was 25% of limit - 75% margin remains. Since sidewalk needs lateral force, not exact angles, authority was ruled "not a hard limit", with a pre-registered criterion for when it WOULD become one: sidewalk fails to track AND roll error keeps growing - then the fix is raising hip_roll kp or lowering the vy target, not more training.
Change
Hypothesis "roll authority insufficient" demoted from blocker to a monitored branch with an explicit trigger condition; C4 proceeded.
Outcome
Later open-loop probes confirmed the actuator could produce the behavior (sidewalk feed-forward ran at full amplitude, 5/5 survival), and the eventual C4 failure causes were measurement and reward, never authority.
Mechanism
Suspended vs loaded comparison separates the actuator's closed-loop competence from the load path: tiny no-load tracking error means the motor/controller is fine and any loaded deviation is statics (gravity / stiffness budget, kp trading error for force). Torque-fraction measurement then bounds how much force headroom actually remains.
Applies when
- suspecting an axis is "too weak" for a new skill
- large position sag on a loaded joint
- deciding between hardware fix, gain change, and more training
“吊挂(--suspend)实测 hip_roll 跟踪误差 0.0008 rad → 执行器无罪,地面下垂 0.21 rad 全是负载所致;稳态占限扭 25% → 仍有 75% 扭矩余量。… 判据:若 C4 出现「侧走跟不动且 roll 误差继续变大」,那才是权限账 … 解法是提 hip_roll 的 kp 或降 vy 目标,不是硬训。”
train/C_LADDER_RUN.md § 3d. roll 权限:已部分澄清,不是硬上限 Three same-shaped judging errors - task metrics (survival, tracking, displacement) cannot stand in for posture metrics
task-metrics-vs-posture-metricsKeep validated posture-class rows (tilt, per-joint L/R asymmetry, temperature) in every acceptance battery alongside task rows; when operator feel contradicts the gates, suspect the metric class before the operator - and never build a new skill on what is actually an asymmetry defect.
Symptom
The C2 product judged "full pass" on task metrics (A800: turn-gap 7 pp, vx+0.30 19/20) felt WORSE in the operator's hands than the half-pass 700: A800 tilted up to 12.90 deg (700: 6.64), drifted left while standing, showed larger per-joint asymmetries, and ran its hip_rolls 5 degC hotter.
Context
The re-judgment catalogued three same-type metric errors in one campaign: (1) stand judged by SURVIVAL - missed 0.5-1.4 m wandering; (2) stand ranked by DISPLACEMENT - ordering was opposite to real feel (tilt ordering matched); (3) chirality judged by wz-tracking GAP - measured turning symmetry while the robot's actual disease was postural left/right asymmetry, "两个不同的东西,且结论相反". Common pattern named: "我一直用「任务指标」当判据,而真机手感对应的是「姿态 指标」… 任务类指标不能替代它". The fix was already in the data: the per-joint left/right asymmetry table (printed identically by sim2sim and deploy) agreed with hardware in direction on every row - "判据可用、有预测力,我只是没把它写进 PASS 条件". Shipping decision followed the posture read: product reverted to 700 ("又一次「买到 精度、卖掉别的」"), and the C4 root moved to 700 as well, with the sharpest line of the episode: A800's left-drift "like sidewalking" is probably its frontal-plane asymmetry defect, not a capability - "在缺陷上建能力是危险的".
Change
Two posture quantities with demonstrated real-robot predictive power promoted into every PASS battery: tilt-max median and per-joint left/right asymmetry (both sim-computable, deploy-homologous); motor-temperature readout added to session close-out.
Outcome
Deployment flipped to the posture-better checkpoint; the hip_roll temperature table (43-48 degC vs 25-28) confirmed the earlier 90%-of-heat account; the run-line acceptance battery inherited the posture rows from birth ("任务类替代不了姿态类").
Mechanism
Task metrics measure goal attainment under the evaluator's episode definition; posture metrics measure the body state trajectory that operators, motors, and downstream skills actually experience. The two can rank candidates oppositely because task success tolerates postural pathology - so a battery without posture rows is blind to exactly what hardware feel reports first.
Applies when
- hardware feel disagrees with a green acceptance table
- choosing between checkpoints that split task vs posture metrics
- selecting the root for a skill that resembles an existing defect
“共同模式:我一直用「任务指标」(存活 / 跟踪率 / 位移)当判据,而真机手感对应的是「姿态指标」(倾角、逐关节左右不对称)。→ 验收判据里必须有姿态类指标,任务类指标不能替代它。… A800 的「左飘像 side walk」很可能 … 是它更大的额平面不对称的表现 —— 在缺陷上建能力是危险的。”
train/README.md § C2 选点改判 (2026-08-09): 手性判据第三次选错指标 Teleop fed the sidewalk axis a command beyond its training band - feet clipped; give each axis its own speed setting
teleop-command-band-per-axisGive every command axis its own teleop scale, clamped to that axis's training band, and reproduce any hardware incident in sim with the exact deployed command values before touching training.
Symptom
Robot stepped on its own foot when sidewalking left under teleop - and only when going left.
Context
The teleop tool used one speed setting for all axes: --teleop-speed 0.20 applied to A/D sent cmd_vy = 0.20, above the training band's top (0.08-0.18) where foot-spacing margin is thinnest. Sim reproduction of the incident (product policy, pw0.8, 5 seeds x 20 s, true collision threshold = single foot width 104 mm): at vy 0.20 the minimum foot distance was 111-115 mm - 7-11 mm from self-collision - vs 147 mm at vy 0.10. Left was 4x more dangerous than right (25% vs 6% of time inside the 160 mm soft wall at vy 0.10), matching the left-only symptom; the margin did not degrade over time (pressing more just lengthened exposure).
Change
deploy_policy gained --teleop-side (default 0.10), separating the lateral speed from the forward speed so each axis's teleop command sits inside its own trained band.
Outcome
Command now inside the band with 43 mm margin at default; the incident became a quantified, reproduced, closed account rather than a mystery.
Mechanism
The policy's competence envelope is the training command distribution per axis; teleop mappings that share one scalar across axes silently command out-of-band inputs on the weakest axis. Asymmetric risk (left vs right) came from the policy's own chirality bias, so a symmetric command produced an asymmetric hazard.
Applies when
- wiring a joystick/teleop layer over a learned policy
- a hardware incident occurs on one command direction only
- training bands differ across command axes
“A/D 一直与 W/S 共用速度档,所以按 A 下发的是 vy = 0.20 —— 既超训练带(0.08~0.18)上沿 … 0.20(遥控实际值)| 111~115 mm | 7~11 mm … 且左比右危险 4 倍 … 处置:deploy_policy 新增 --teleop-side(默认 0.10),侧移与前进档分开。”
train/C_LADDER_RUN.md § 3p. 一 向左走踩到自己 → --teleop-speed 0.20 同时喂给了 vy