Training Coach
Doctrine
A report may cite any of these as doctrine-N.
doctrine-1Contract freeze and fingerprint disciplineThe policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.
Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).
Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.
doctrine-2Attribution by resolved training params - never eval-override knobsCapability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.
Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).
Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.
doctrine-3PASS gates become constraints; FAIL gates become objectivesOnce a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.
Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.
Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".
doctrine-4One variable per ladder rung - counted against what the checkpoint sawA rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.
Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.
Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.
doctrine-5Pre-register risks, readings, and stop criteria before the ladderBefore a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.
Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).
Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.
doctrine-6Plant parameters are measured, never inventedEvery plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.
Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).
Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.
doctrine-7Sim2sim gate before sim2real - under deployment conditionsEvery checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).
Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).
Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.
doctrine-8Observation honesty - the actor's inputs are a hardware contractThe actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.
Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).
Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.
doctrine-9Reward economics are audited in realized currencyReward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.
Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).
Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.
doctrine-10The zero-cost option must be the desired behaviorFor every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.
Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).
Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.
doctrine-11Measurement discipline: independent referees, signs, distributionsA disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.
Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).
Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.
doctrine-12The deployment pipeline is plantIrreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.
Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).
Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.
doctrine-13DR budget is finite; its distribution is the measured supportRobustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.
Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).
Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.
doctrine-14Gates measure what hardware feels: posture, margins, stripped assistsAcceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.
Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).
Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.
doctrine-15Fork and root selection: recoverability, maturity, frozen rewardsChoose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.
Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).
Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.
doctrine-16Curricula: verified engagement, lineage counters, disease-phase gatingAutomatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.
Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).
Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.
doctrine-17Probe before training: feasibility first, hypotheses in tablesAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.
Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).
Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).
doctrine-18External advice is recomputed locally; values transfer as ratiosEvery external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.
Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).
Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.
doctrine-19Hardware sessions are scripted experiments, not tuning sessionsReal-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.
Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).
Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.
doctrine-20Close questions in writing; restart when the debt is structuralAudited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.
Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).
Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.
doctrine-21Name the quantity in the space it lives inA goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.
Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).
Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.
doctrine-22Continuation needs a live gradient; a release is chosen by a scanContinue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.
Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).
Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.
Experience cards
171 cards matching “curriculum-history-is-part-of-the-product”.
Isaac splits Coulomb friction into static and dynamic columns - wiring only static means zero loss during motion, silently discarding the identified value
sim-api-friction-columnsWhen installing identified actuator parameters, map each measured quantity to the simulator's exact API column for the operative regime (dynamic for moving loss, viscous for damping), verify per joint after landing, and audit how randomization intervals fall on each column's nominal.
Symptom
The hardware-identified Coulomb friction (tau_c) was about to be installed into the trainer through the friction= field alone - which in Isaac 5 populates only STATIC friction, so during motion the joints would lose no torque at all: "只给 static 则运动中不损耗, 辨识的 τ_c 走路时等于没接" (the identified tau_c would effectively not be connected while walking).
Context
The v12 integration wired all three columns deliberately: armature= and friction= from the 2026-08-04 hardware identification, PLUS dynamic_friction= (Isaac 5 splits Coulomb into static/dynamic; the moving-loss column is dynamic) and viscous_friction= (= the measured joint damping 0.02, aligned to MJCF's damping). Each value was re-checked per joint after landing. A DR interaction was audited and booked rather than hidden: randomize_joint_parameters jitters ALL friction columns with ONE interval - the [-0.05, +0.10] band was calibrated against the Coulomb nominal, and landing on the viscous nominal 0.02 it becomes [0, 0.12], "偏宽但保守" (wide but conservative), accepted with the note that pre-viscous behavior was already [0, 0.10] on a base of 0.
Change
Measured actuator parameters installed across all applicable API columns (armature, static, dynamic, viscous), with the DR side effect on shared randomization intervals audited and recorded.
Outcome
The first generation where the identified plant actually acts during motion in the trainer; the silent-column failure mode documented before it cost a training run.
Mechanism
Physics engines decompose "friction" differently (single coefficient vs static/dynamic/viscous columns); a measured parameter is only installed when it reaches the column the solver reads in the regime that matters (motion, not stiction). Randomizers that share one interval across columns rescale the band by each column's nominal - a hidden unit change.
Applies when
- installing identified friction/armature into any trainer
- porting plant parameters between simulators or engine versions
- joint losses in sim do not match bench measurements during motion
“并额外传 dynamic_friction=(Isaac 5 把库仑拆 static/dynamic 两列,只给 static 则运动中不损耗,辨识的 τ_c 走路时等于没接)与 viscous_friction=(= joint_damping 0.02,对齐 MJCF damping)。… randomize_joint_parameters 用同一个 friction 区间抖三列, [-0.05,+0.10] 是按库仑标称标的,落到粘滞标称 0.02 上成了 [0,0.12]”
train/WALK_V12_SPEC.md § 7. 核查单 (Isaac 接 V.ACTUATORS 新字段) Close a question with an audit, then freeze the wording - later symptoms may not reopen it without new hard evidence
frozen-verdicts-semantic-boundariesWhen an audit closes a hardware-vs-policy question, record the closing evidence, freeze a citable wording for future recurrences, and set the reopening bar explicitly; separate robustness perturbations from plant-truth questions so a DR rung's failure can never silently reopen a closed measurement.
Symptom
Recurring directional bias on the robot kept re-suggesting "maybe the hardware/COM/mechanics are asymmetric", threatening to re-litigate questions that audits had already closed - burning attention each time a descendant policy leaned or drifted.
Context
Two boundary decisions were written as permanent: (1) semantic separation - "S2④ COM ±20mm = 纯鲁棒性扰动,不再承担「解释真机后仰」任务" - if the COM-DR rung degrades, the ONLY allowed conclusion is "policy insufficiently robust to COM uncertainty"; reopening "is the CAD COM wrong" is forbidden because the mass audit was completed and closed (@63f9212). (2) a frozen wording for chirality, to be quoted verbatim whenever left/right bias appears in later rungs: observed directional bias = policy-level spontaneous symmetry breaking; plant asymmetry = no supporting evidence after the mass + model symmetry audit; mitigation candidate pi_sym queued, not blocking. The evidential basis was quantitative: the root policy was perfectly symmetric under +/-6 N*s pushes (40/40) while descendants broke (17/40, 13/40) - "手性是 S2 训练中获得的, 根没有; 机械侧已双 PASS 关案, 不重开".
Change
Closed questions carry (a) the audit commit that closed them, (b) a frozen citable wording for recurrences, and (c) an explicit evidence bar for reopening ("无新硬证据不得重开").
Outcome
Later chirality observations (C2's 15 pp turn gap, hip_roll drift bias) were handled as policy-lineage properties with policy-side mitigations, without a single hardware re-audit cycle.
Mechanism
Symptom classes recur under different guises; without a frozen verdict each recurrence re-runs the same expensive investigation and risks a different (worse-informed) conclusion. Freezing verdict plus wording converts recurring symptoms into citations, while the evidence bar keeps the closure honest rather than dogmatic - the root/descendant symmetry comparison is what makes "it's the training, not the machine" checkable at any time.
Applies when
- a recurring symptom keeps suggesting an already-audited hardware cause
- writing conclusions for a completed calibration/audit
- a DR rung's degradation invites re-measuring the plant
“若 S2④ 退化,结论只能是「当前 policy 对 COM 不确定性不够鲁棒」,不得重开「CAD COM 是不是错了」… 手性冻结表述 … Plant asymmetry: no supporting evidence after mass + model symmetry audit … 无新硬证据不得重开机械不对称”
train/OMNI_V0_SPEC.md § 4. 语义分界与手性冻结表述(2026-08-07 用户定,永久) Run acceptance under measured contact parameters - honest condim/torsional-friction flipped a false PASS into a real-matching FAIL
eval-plant-honesty-contact-paramsPin the evaluation plant's contact model to measured values (contact dimension, torsional/rolling friction, mu) before trusting any gate that involves slip, impact, or drift - a gate can only fail a policy for physics its simulator contains.
Symptom
walk_v5 passed the old acceptance battery yet failed on the real robot (footfall force, drift, kicking) - the evaluation plant was flattering the policy.
Context
The battery was re-run under "honest contact parameters" - condim 4 (adding torsional contact) with measured torsional friction 0.035 - and v5 then FAILED exactly the rows corresponding to its real problems: heading 185 deg (limit 30), support-foot yaw slip 284 deg (limit 80), landing force 1.72x (limit 1.5x), suspended tilt 45.9 deg (limit 10). The slip physics depends on torsional friction, which the default contact model (condim 3) does not even simulate - a slip problem is invisible to an evaluator that cannot represent yaw friction at the foot. Term-sizing measurements for the new rewards were likewise taken under the same honest parameters (cmd 0.45, skipping the 5 s start transient).
Change
Acceptance harness pinned to condim 4 / torsion 0.035 (measured); verdicts issued under defaults declared non-citable for these rows.
Outcome
Sim acceptance verdicts began agreeing with hardware ("现在失败, 与真机一致"); the v6 fixes could be developed and validated against an evaluator that could actually see the disease.
Mechanism
An evaluator is a plant model too: contact dimensionality and friction values decide which failure modes exist in the simulation at all. Evaluating under default contact parameters tests the policy in a world where its real failure is physically impossible, producing structurally false PASSes.
Applies when
- sim acceptance passes policies that fail on hardware
- slip/drift/impact gates run under default simulator contact settings
- setting up a cross-simulator evaluation harness
“accept_v2.py 已加三条判据, walk_v5 在诚实的接触参数下(--condim 4 --torsion 0.035)现在失败, 与真机一致:直行 15s 航向累计 <30° | 185° ✗ … 落脚力峰值 <1.5× 体重 | 1.72× ✗”
train/WALK_V6_MINIMAL.md § 5. 验收 Borrow reward values from other robots as ratios (to tracking weight, to leg length) - never as absolute numbers
transfer-ratios-not-absolutesWhen importing any numeric from another robot's config or paper, identify its natural normalizer (tracking weight, leg length, sqrt(g*L), body mass) and transfer the dimensionless ratio; sanity-check against a same-scale robot when one exists.
Symptom
Published configs offered tempting absolute values (swing height target 0.06-0.08 m, weight -20) that would have been wrong for a robot with half the leg length and a different tracking weight.
Context
The cross-check normalized before transferring: G1's feet_swing_height weight -20 against tracking +1.0 is a 20x ratio, so with local tracking at 1.5 the equivalent is -30, not -20. G1's 0.06 m target on a ~0.70 m leg scales to ~28 mm on the local 0.325 m leg (Humanoid-Gym converts to ~23 mm), confirming the locally chosen 0.03 m and explicitly rejecting copying 0.05-0.08 absolutes. A same-class robot (Menlo's 16 kg) was used to sanity-check feet_air_time (+0.5 vs the local 2.0, flagged over-high). A nondimensional check of the same kind later validated the sidewalk speed target (v/sqrt(gL) = 0.081 vs hardware-verified 0.101 - 0.104 - inside the envelope, conservative).
Change
All borrowed values converted through ratios (weight/tracking-weight, height/leg-length, dimensionless speed) before entering the config.
Outcome
The scaled values worked (0.03 m target matched both the scaling law and measured 22-23 mm baseline); no cross-robot absolute was ever copied raw.
Mechanism
Reward economies are scale-relative (only ratios to the tracking term matter to the optimum) and kinematic quantities are morphology-relative (clearance scales with leg length, speed with sqrt(g*L)); absolutes encode the source robot's scale, ratios encode the design intent.
Applies when
- copying reward weights/targets from open-source configs or papers
- setting clearance heights, speed targets, or impact thresholds
- comparing your weights to published tables
“G1 的 feet_swing_height 是 tracking 的 20 倍(−20 vs +1.0)。我们 tracking 是 1.5,按同比例应为 −30 … G1 目标 0.06 m / 腿长 ~0.70 m,换算到我们 0.325 m 腿长约 28 mm;Humanoid-Gym 换算约 23 mm。故 target 取 0.03 m 是对的 … 不必抄 0.05~0.08 的绝对值。”
train/WALK_DIAGNOSIS.md § 修正 ②(权重放大) / 修正 ③(目标高度按腿长缩放) Training-log reward values and fixed-command eval values live on different distributions - comparing them once claimed a 44% improvement that was really 6-10%
same-distribution-reward-comparisonQuote reward-term values only with their distribution attached (command range, DR on/off, environment), and compare across runs only when those match; re-measure in a common environment before claiming any improvement percentage.
Symptom
A v6-era analysis concluded slip had dropped 44% by comparing the training log's Episode_Reward against values calibrated in a fixed-command play environment; a same-condition re-measurement showed the true improvement was 6-10%.
Context
The training log's reward is an expectation over the training command distribution (vx 0.15-0.5, yaw +/-0.6, with pushes and domain randomization); play-environment calibrations are taken at a single fixed command with DR off. Subtracting one from the other compares apples to oranges - the warning was written into the v7 pre-flight: "奖励数值只能在同一指令分布下比较 … 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次)".
Change
Rule adopted: any before/after reward-term comparison must hold the command distribution, DR state, and evaluation environment fixed; training-log values compare only against training-log values of runs with identical command/DR configs.
Outcome
The phantom 44% improvement was retracted; later term-level accounting (e.g. the C4 ignore-floor work) consistently specified its distribution before quoting numbers.
Mechanism
A reward term's expectation depends on the visited-state distribution as much as on the policy; changing the command distribution or DR moves every term's baseline. Cross-distribution differences therefore measure the distributions, not the policy change.
Applies when
- comparing reward telemetry across training runs or vs play evals
- claiming improvement percentages from training logs
- term-level reward accounting for diagnosis
“奖励数值只能在同一指令分布下比较。训练日志的 Episode_Reward 是在训练指令分布上算的(vx 0.15~0.5 / 偏航 ±0.6 / 带推力与域随机化), 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次: 据此以为滑移降了 44%, 同条件对拍只有 6~10%)。”
train/WALK_V7_SPEC.md § 3. 开训自查 ⚠️ The smoke watcher panicked at the wrong operating point and its score goes blind when gates saturate - treat it as a survival sentinel, not a judge
smoke-watcher-operating-pointConfigure every automated evaluator at the lineage's declared deployment operating point, give it graded metrics that cannot saturate, and until then scope its authority to catastrophe-detection - never let a mis-configured watcher stop or rank a rung on its own.
Symptom
Two watcher misfires in one ladder: (1) during the friction rung the watcher (evaluating at kd 1.0) reported panic-level 1/3 survival from iter 3300 - falsified by the official kd 1.2 scan, because the lineage's design operating point was kd 1.2 and the watcher lacked the --kd-scale passthrough; (2) during the PD rung the watcher's early-stop score froze at iter 1050 despite ongoing drift improvements, because with all eight gates passing (constant 0/3 failures) the score has no gradient left - "八门全过恒 0/3 时 score 对漂移改善盲, s1f 课文三现".
Context
Both are the same category: the in-training smoke loop is an instrument with its own configuration (operating point, score design), and its verdicts are only as aligned as that configuration. The booked doctrine: "冒烟只当存活哨兵" - until the watcher evaluates at the deployment operating point with graded metrics, its role is detecting catastrophes, not ranking checkpoints; ranking belongs to the full battery at the design operating point (and the watcher's scoring was separately patched to weight survival 3x so recoveries during hard phases are not early-stopped away).
Change
Watcher debt booked (--kd-scale passthrough); score saturation acknowledged with graded columns planned; selection authority kept with 20-seed batteries at the declared operating point.
Outcome
A false panic did not abort a rung that was actually passing at its design point; a frozen score did not hide real drift gains; the instrument's authority was scoped to what its configuration can actually see.
Mechanism
An evaluator is itself configured (gain profile, delay, metrics); evaluating a policy away from its design operating point measures a counterfactual robot, and bounded scores saturate once binary gates pass, losing all sensitivity. Instruments need the same operating-point discipline as deployments and graded outputs to retain gradient.
Applies when
- an automated smoke loop contradicts the official battery
- early-stop scores freeze while graded metrics still improve
- lineages with non-default deployment gain/delay profiles
“watcher (kd1.0 口径) 3300 起 1/3 恐慌被 kd1.2 正式扫描证伪为考纲外假象 —— 工作点评测口径教训: watch_ckpt 缺 --kd-scale 透传 (待补), 冒烟只当存活哨兵。… watcher score 饱和误停 @1050(八门全过恒 0/3 时 score 对漂移改善盲, s1f 课文三现)”
train/README.md § omni_s2e_fric (watcher 恐慌被证伪) / omni_s2e_pd (500 臂) The +/-50 mm lateral COM randomization meant to spread the legs coincided with legs pulling IN - rolled back per its own pre-registered contract
com-dr-rollback-on-symptomWhen adopting a DR value that covers no local measurement, write its intent and rollback trigger into the config at adoption time; roll it back as the control arm the moment the symptom contradicts the intent, and promote the symptom's metric into the acceptance battery.
Symptom
After v7 adopted the reference developer's oversized lateral COM randomization (+/-50 mm) explicitly to force leg spread, the real robot's legs narrowed instead - lateral mean 154 mm / closest 107 mm in sim (nominal 214.5), narrower still on hardware with occasional leg contact.
Context
The rollback was clean because the adoption had been honest: the robot.yaml comment recorded the intent AND that the +/-50 value covered no local measurement (only a 16/7 mm measured offset existed; even the prior widening to +/-20 was subjective), plus the reference's own reported side effect (base sway) and the note "这一项要单独跑、 单独归因". When the opposite symptom appeared, v8 returned y to +/-20 mm as the control arm ("要么没起作用、要么帮了倒忙 … 按约定退回做 对照"), kept x/z untouched (a noise-level difference not worth another variable), and named the second suspect: the landing penalty itself, via the reference's own three-link chain (landing penalty -> stance narrows -> spacing penalty needed). A gate lesson was booked in the same table: v7's sim numbers had ALREADY crossed the line (154/107 vs v5's 182/147) - "这个指标本可拦下 v7" - so foot-distance became a standing acceptance row (min >120 mm, zero leg-leg contacts).
Change
base_com_offset_m y: 0.050 -> 0.020 (x/z kept), regenerated through the export tool rather than hand-editing derived files; foot-distance acceptance row added.
Outcome
A borrowed DR lever with no local measurement basis was retired the moment its symptom contradicted its purpose, at single-variable cost; the metric that would have caught it pre-hardware entered the gate.
Mechanism
DR ranges shape behavior through the policy's robustness strategy, which is jointly determined with every reward term; a lever that forces stance width on one robot can be dominated by a stronger narrowing pressure (landing softness) on another. Levers adopted without local measurement must carry their own rollback trigger, because there is no nominal to argue from when they misbehave.
Conflicts
Causality is not fully closed in the source: the narrowing may come from the landing penalty rather than the COM lever ("腿距的第二嫌疑人是 ④ 本身"); the rollback is the pre-agreed control experiment, not a verdict that the lever caused the narrowing.
Applies when
- importing DR ranges or behavioral-forcing randomizations from references
- a DR lever's observed effect contradicts its documented purpose
- a sim metric existed that would have caught a shipped regression
“⑥ 的本意 … 是逼策略把脚分开;真机结果是脚向内收且偶发相碰——要么没起作用、要么帮了倒忙。… 注释当时就写了"这一项要单独跑、单独归因"。现在症状出现了,按约定退回做对照。”
train/WALK_V8_SPEC.md § 3. 改动 C — 质心随机化退回(撤销 v7-⑥ 的 y 项) A torque-tail penalty was paid for by bracing the legs against each other - the second simulator's leg-contact count caught it, and the first explanation ("the trainer can't see self-collision") was retracted from the run's own config
torque-penalty-bought-by-leg-bracingWhen a penalty lowers a demand metric, look for what the policy traded to get there - keep self-contact frames and foot spacing as standing sim2sim readouts - and check any "the trainer cannot see X" explanation against the run's resolved config before it enters the record.
Symptom
After R3.1's torque_headroom term collapsed the demand tail, MuJoCo success fell 100 -> 98% and leg-leg contact frames at mu 1.0 rose 750 -> 2,190 (worst rollout 177 -> 450). The one failure (prone seed 2) had the legs crossed, one foot on the other leg, trapped at 0.067 m - visible on video.
Context
Across the ten prone seeds, foot spacing and leg-leg contact frames were monotonically anti-correlated, and the failure was the extreme of the series. Pulling the legs toward the midline shortens the hip_roll lever arm and lowers torque demand. At the time the spec explained it as Isaac training without self-collisions ("a free lunch in a simulator without self-collision").
Change
Leg-leg contact frames and foot spacing were tracked in every MuJoCo gate; R3.2's candidates were "train with self-collision on" or "a minimum leg spacing term" - not stacked.
Outcome
The next rung's joint-velocity penalty incidentally erased the dependency (2,190 -> 86 frames). On 08-10 the runs' logged env.yaml showed enabled_self_collisions true in both r3_1 and v2_2 (inherited from walk v10): the tangle was physically learned bracing, visible to both simulators, and the Isaac/MuJoCo contact-count gap was mesh and contact fidelity. The "self-collision debt" narrative was withdrawn for the whole line.
Mechanism
A penalty on demand rewards any configuration that lowers demand; legs pressed together act as a mutual support that fails when contact geometry shifts slightly.
Conflicts
§24 attributes the dependency to self-collisions being disabled in training; §36 retracts that from the runs' env.yaml ("§24's mechanism explanation was wrong") and keeps the older sections unedited as history.
Applies when
- a torque, impact or energy penalty improves its metric and cross-sim success drops
- legs or links approach each other after a regularization change
- an explanation relies on a simulator setting nobody checked in the run config
“prone 十个 seed 逐条看,脚距与腿-腿接触帧数单调反相关, 而唯一失败的那条正是最极端的一条 … 机制上说得通:把腿收到身体中线附近能缩短 `hip_roll` 力臂、降低力矩需求”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §24 代价:MuJoCo 成功率 100% → 98%,病因是两腿卡住 Push-test protocol - positive side first, fragile side spotted, axes aligned in the log, and cross-machine push counts stay qualitative
push-test-chirality-protocolOrder disturbance tests from the robust side to the fragile side with protection scaled to sim-measured asymmetry, align and log frame conventions before testing, and treat cross-domain disturbance counts as qualitative evidence only.
Symptom
Hand-push testing on hardware risked falls on a side sim had already flagged as fragile, and push counts invited apples-to-oranges comparison with sim numbers.
Context
Sim chirality was explicit: descendants were far more fragile in -y (fric-3000@kd1.2: +6 N*s survived 15/20 vs -6 N*s only 3-9/20) while the s1e control was perfectly symmetric (40/40). The protocol therefore: push the positive direction first, keep a spotter for the negative side; before any push, record which real-robot side corresponds to sim's +y in the log ("上机前对一次坐标"); and - citing the chaos lesson ("混沌课文") - real push results are used only as qualitative corroboration, never compared numerically with sim survival counts across machines.
Change
Push testing became a scripted, chirality-aware protocol with frame alignment as a logged precondition and an explicit epistemic limit on cross-domain count comparison.
Outcome
The fragile side was tested with protection informed by sim's quantified asymmetry; logs stayed interpretable because the frame correspondence was recorded before the first push.
Mechanism
Disturbance-response chirality is a real, quantifiable lineage property, so test order should follow measured fragility; and perturbation outcomes are chaotic in the details (divergent trajectories from tiny differences), so counts do not transfer across domains even when qualitative rankings do.
Applies when
- planning push/disturbance tests on hardware
- sim shows directional asymmetry in disturbance survival
- someone proposes comparing real push counts to sim counts
“先正向后负向, 负向留人扶 —— sim 手性明确: 后代在负 y 向显著更脆 (fric-3000 @kd1.2: +6 N·s 15/20 vs −6 N·s 3~9/20), 而 s1e@0.8 两向 40/40 完全对称。上机前对一次坐标 … 跨机不做二值结论 (混沌课文): 真机推力只作定性对照, 不与 sim 计数对比。”
train/REAL_RUN_S2.md § 3. 抗推 (可选, 人手推; 做则按此协议) Oversized lateral COM randomization (+/-5 cm) deliberately forces leg spread
com-randomization-forces-leg-spreadDR ranges can be behavior-shaping tools, not just robustness padding: oversize a randomization axis to force a strategy the reward struggles to express - and expect a compensating behavior to appear as the cost.
Symptom
Feet drift toward the centerline and even collide; policy has no incentive to keep a lateral support base.
Context
COM randomization ranges were chosen asymmetrically by axis: lateral +/-5 cm ("比常规大,故意的" - larger than usual, on purpose), fore-aft +/-2 cm, vertical +/-2 cm. The oversized lateral range is not robustness padding but a behavioral forcing function. Lucen logged it as directly relevant to its own roll-channel / sideways leg-kick symptom.
Change
Set COM randomization to lateral +/-5 cm, fore-aft +/-2 cm, vertical +/-2 cm, with the lateral band intentionally oversized to make narrow stances fail during training.
Outcome
Effective at separating the feet on the reference robot; side effect - the base began swaying left-right, which then required a foot-centerline distance penalty (see reward-chain-foot-height-landing-spacing).
Mechanism
Randomizing COM laterally makes narrow-stance policies fall for some draws, so PPO discovers wide stances as the only strategy robust across the band - DR used as an implicit reward. The sway side effect appears because the policy hedges against unknown COM by active lateral correction.
Applies when
- feet too close / self-collision in a learned gait
- roll-axis instability suspected to come from narrow stance
- choosing COM or mass-offset DR ranges
“两脚太近甚至互撞 → 先试质心横向随机化 ±5 cm,逼迫策略把脚分开;有效但引发新问题——基座开始左右摇摆 … 横向 ±5 cm(比常规大,故意的,用来逼出分腿)/ 前后 ±2 cm / 垂直 ±2 cm”
Experience.md § 质心随机化范围 (lines 75, 84-86) Every power cycle starts with the same read-only pre-flight - read the buses, check the torque limits against 12/17/11, verify the IMU axes, check the ports after any new USB device - and any reassembly re-measures the joint zeros
power-cycle-preflightStart every powered session with a fixed, read-only pre-flight - bus responses, torque limits equal to the simulated ones, IMU axes, device identities - and re-measure joint zeros after any mechanical reassembly before running a policy.
Symptom
Hardware state drifts between sessions in ways no policy can see: a motor that stops answering after a power cycle, a torque limit that differs from the one simulated, an IMU axis flipped, two USB devices swapping identities, a joint zero moved by reassembly.
Context
The runbook's session order before any policy runs: read every motor on both CAN buses without enabling them (the first command after every power cycle); set_torque --check, all twelve motors must read 12/17/11 N*m, and any difference is written back; imu_reader --verify-axes, where the operator tilts the robot forward and to the right and every check must pass before continuing; check_ports after plugging in any new USB device (the IMU and a CAN adapter once collided on USB identity). After re-mounting motors: read the buses, then re-measure the calibration offsets (three repeats, written back) - "skipping it means running everything on the wrong zero". Hanging checklists repeat the torque-limit check (the deploy script also self-checks at start).
Change
A fixed, read-only pre-flight run in the same order every session.
Outcome
The runbook records one earlier hardware check in the same spirit: all 12 motors' implied kp fell within 18.4-22.0 for a commanded 20, inside the kp randomization range used in training.
Mechanism
A policy transfers only if the plant matches the one it was evaluated on; the pre-flight turns silent hardware drift into a failed check before the robot moves.
Applies when
- the first command after powering a robot on
- after swapping adapters, cables or motors
- a policy that worked last session suddenly behaves differently
“python tools/set_torque.py --check # 12 颗应全对 12/17/11, 有 diff 就 --write … 插任何新 USB 设备后都先跑一次 check_ports.py(IMU 和 CANable 的 USB 身份撞过车) … python tools/calib_stance.py --repeat 3 --write # 重标 offset —— 8/9/10 重新装, 机械零位变了”
RL系统/FOLLOW THIS copy 2.md § WALK / STAND 每次开始前 / 换CAN / 装回后必做两件