Skip to content

Training Coach

The coach reads a training run and returns a diagnosis, in-whitelist proposals and an experiment plan — and every claim cites one of the cards below. This is that corpus: 22 methodology rules distilled from a real sim2real programme, and 171 episode cards behind them, each with the sentence in the war history it came from.

Doctrine

A report may cite any of these as doctrine-N.

  1. doctrine-1Contract freeze and fingerprint discipline

    The policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.

    Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).

    Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.

    contract-freeze-and-checkerlegacy-profile-pinningderived-asset-staleness-checkgain-profile-belongs-in-the-stamprecovery-two-policies-and-a-state-machinewalk-recovery-fsm-handoff

  2. doctrine-2Attribution by resolved training params - never eval-override knobs

    Capability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.

    Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).

    Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.

    kd-bandwidth-mu-law-attributiondeploy-knob-attribution-before-retrainingcycle-time-override-is-oodresolved-config-is-source-of-truth

  3. doctrine-3PASS gates become constraints; FAIL gates become objectives

    Once a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.

    Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.

    Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".

    fixed-acceptance-matrix-per-rungpreregistered-stop-criteria-per-rung

  4. doctrine-4One variable per ladder rung - counted against what the checkpoint saw

    A rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.

    Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.

    Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.

    resume-state-dr-auditorthogonal-batch-with-ablation-orderfill-the-missing-factorial-cell

  5. doctrine-5Pre-register risks, readings, and stop criteria before the ladder

    Before a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.

    Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).

    Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.

    preregister-risks-and-fork-readingspreregistered-stop-criteria-per-rungpreregistered-real-expectationsfeasibility-accounts-lock-design-point

  6. doctrine-6Plant parameters are measured, never invented

    Every plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.

    Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).

    Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.

    friction-measured-not-guessedarmature-n2-rotor-inertiatorque-limit-shape-by-measured-peakslatency-lerp-reverse-extrapolationfeasibility-accounts-lock-design-pointsim-api-friction-columnsget-up-feasibility-accounts-before-trainingsingle-support-gain-authority-probe

  7. doctrine-7Sim2sim gate before sim2real - under deployment conditions

    Every checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).

    Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).

    Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.

    sim2sim-gate-before-sim2realeval-plant-honesty-contact-paramspipeline-latency-is-plant-not-drbody-frame-velocity-api-audittorque-penalty-bought-by-leg-bracingtorque-disagreement-between-simulators-unresolved

  8. doctrine-8Observation honesty - the actor's inputs are a hardware contract

    The actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.

    Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).

    Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.

    observation-honesty-critic-onlyhistory-obs-needs-plant-variationreward-observability-limitdeploy-heading-loop-and-align-training

  9. doctrine-9Reward economics are audited in realized currency

    Reward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.

    Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).

    Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.

    realized-contribution-auditreward-cost-of-ignoring-auditgate-new-reward-terms-by-commandignore-floor-diagnosiscalibrate-threshold-between-healthy-and-sickcalibration-threshold-with-withdrawal-clauseinert-reward-term-auditseated-basin-dead-exp-kerneltail-torque-needs-hinge-on-computed-demand

  10. doctrine-10The zero-cost option must be the desired behavior

    For every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.

    Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).

    Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.

    penalize-the-slip-not-the-jointsaturation-cheating-zero-rate-costmoving-gate-42x-stand-taxtermination-closes-degenerate-basinpenalty-gate-is-an-escape-hatchsoft-limit-penalty-charges-nominal-poseunpriced-foot-attitude-is-a-free-variableenumerate-cheapest-cheats-before-trainingbinary-band-reward-fake-touchdown

  11. doctrine-11Measurement discipline: independent referees, signs, distributions

    A disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.

    Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).

    Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.

    independent-referee-for-metric-disputesbody-frame-velocity-api-auditsame-sign-response-is-yaw-biasmedian-hides-bimodal-distributionratio-metrics-need-absolute-checkheading-integral-not-body-ratesame-distribution-reward-comparisonmultiseed-sign-test-for-driftsingle-impulse-recovery-is-chaotic

  12. doctrine-12The deployment pipeline is plant

    Irreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.

    Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).

    Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.

    pipeline-latency-is-plant-not-drpower-scale-hurts-nonforward-axesdeploy-scaling-not-training-equivalentteleop-command-band-per-axislatency-dr-covers-measured-pipelinedeploy-rate-limiter-windupslew-anchor-is-an-integratorbeta-anchored-action-targetpower-derating-cuts-full-range-contract

  13. doctrine-13DR budget is finite; its distribution is the measured support

    Robustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.

    Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).

    Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.

    push-dr-conditional-budget-conservationdr-tail-plant-continuationconstant-value-dr-overfits-margintask-shaping-before-plant-hardeningcom-randomization-forces-leg-spreadcom-dr-rollback-on-symptomthin-dr-judged-by-channel-coveragefriction-priority-re-measured-after-plant-change

  14. doctrine-14Gates measure what hardware feels: posture, margins, stripped assists

    Acceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.

    Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).

    Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.

    task-metrics-vs-posture-metricsstand-gate-posture-not-survivalconstant-value-dr-overfits-marginsuspension-probe-removes-free-stabilizerhip-roll-sum-predicts-lateral-driftchirality-scored-separatelylow-speed-commands-reward-draggingend-state-confusion-matrixfrozen-acceptance-distribution-and-pinned-seedvideo-as-acceptance-recordepisode-length-bounds-what-a-gate-seesremoved-wall-returns-on-hardware

  15. doctrine-15Fork and root selection: recoverability, maturity, frozen rewards

    Choose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.

    Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).

    Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.

    fork-root-recoverable-shortfallroot-maturity-vs-product-qualityfine-tune-reward-change-falsified

  16. doctrine-16Curricula: verified engagement, lineage counters, disease-phase gating

    Automatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.

    Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).

    Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.

    auto-curriculum-engagement-checkcurriculum-counter-lineage-stepsgate-penalties-to-the-disease-phaseaggregate-metrics-mask-subgroup-failurebucket-share-is-not-a-gradient-levercurriculum-criterion-conditioned-on-lagging-categoryper-step-income-drives-speed-time-gatetime-gate-vs-wide-stance-retire-the-fix

  17. doctrine-17Probe before training: feasibility first, hypotheses in tables

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.

    Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).

    Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).

    open-loop-probe-before-reward-tuninghypothesis-table-code-auditperiod-doubling-evidence-racesuspended-test-isolates-actuator-authorityconfiguration-probe-wall-not-slopeprone-dead-end-is-foot-placementdof-vel-penalty-is-not-a-pacing-knobamplitude-cut-falsified-yaw-fix

  18. doctrine-18External advice is recomputed locally; values transfer as ratios

    Every external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.

    Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).

    Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.

    external-advice-audit-against-own-arithmetictransfer-ratios-not-absoluteslatency-dr-covers-measured-pipelinereference-structure-fk-amplitude-divisioncycle-average-tracking-for-gait-quantitiesadvisor-paraphrase-vs-paper

  19. doctrine-19Hardware sessions are scripted experiments, not tuning sessions

    Real-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.

    Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).

    Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.

    risk-ordered-real-deploymentbattery-bracketed-real-abknow-zero-command-behaviorpush-test-chirality-protocolno-field-tuning-protocolreversible-single-variable-field-experimentssim-veto-needs-real-confirmationfirst-real-get-up-violent-stage-one-policystaged-hang-mat-floor-for-get-uppower-cycle-preflighttwo-machine-config-disciplinefall-guard-becomes-a-statehardware-log-is-the-attribution-input

  20. doctrine-20Close questions in writing; restart when the debt is structural

    Audited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.

    Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).

    Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.

    frozen-verdicts-semantic-boundariesstale-verdicts-under-old-stackzero-offset-calibration-shifts-envelopeplant-swap-invariants-vs-shiftsfreeze-lineage-fix-structure-restartminimal-reward-table-with-provenancewrite-hardware-verdicts-back

  21. doctrine-21Name the quantity in the space it lives in

    A goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.

    Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).

    Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.

    joint-space-proxy-for-task-space-quantitycontact-detector-single-signal-lieszero-partial-credit-is-not-an-iteration-problemheading-integral-not-body-rate

  22. doctrine-22Continuation needs a live gradient; a release is chosen by a scan

    Continue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.

    Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).

    Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.

    converged-continuation-is-poisoncheckpoint-choice-is-a-full-gate-scanstance-decided-by-get-up-pathstop-stacking-roll-back-and-auditcurriculum-history-is-part-of-the-productcontinuation-budget-not-from-zeroroot-maturity-vs-product-quality

Experience cards

37 cards tagged gate-battery.

  • Record-high training reward hid a fully-failing DR subgroup - aggregate metrics average over draws, gates must test per conditionaggregate-metrics-mask-subgroup-failure
    Mechanism understoodomnitraining-rundomain-randomizationgate-batterycurriculum

    Never gate on metrics aggregated across DR draws: evaluate at fixed representative conditions (especially the deployment-critical stratum), and if a difficulty axis matters, ramp it on measured per-stratum success rather than sampling the full range from iteration zero.

    Symptom

    omni_s1e trained under constant-wide latency DR (0, 0.06 s) posted the lineage's highest-ever Isaac reward (129) - while the --delay 2 smoke evaluation showed 3/3 falls from iter 1500 onward, persisting to early stop; the usable checkpoint window shrank to iters 500-1000.

    Context

    Diagnosis written plainly: "聚合奖励掩盖重延迟尾部子群体失败" - the aggregated reward averages over latency draws, so the majority of light-delay environments can mask the total failure of the heavy-delay tail. The remedy for the training side was a survival-gated ratchet curriculum (survival_gated_latency): the sampling cap starts at 0.02 s and rises +0.01 only when a 4096-reset window's survival (time_out share) reaches >=90%, capped at 0.06, ratchet up-only - "增益与延迟耐受一起长,不升到策略撑不住的 地方" (gain and delay tolerance grow together; never raise past what the policy can hold). The detection side was already in place from the noise-crutch episode: the per-condition smoke curve, not the training reward, is the health readout.

    Change

    Latency exposure made curriculum-gated on measured subgroup survival instead of uniform-from-zero; per-condition (--delay 2) smoke evaluation kept as the authoritative curve; watcher scoring adjusted (survival weighted 3x) so recovery during hard phases is not early-stopped away.

    Outcome

    The failure mode was caught by the smoke curve within one generation; the follow-up redesign (deterministic staged latency) superseded the ratchet, but the aggregate-masking lesson held through both.

    Mechanism

    Expected-return training weights each DR draw by probability, so a subgroup can contribute bounded loss while being catastrophically failed; any scalar averaged over the randomization cannot distinguish "uniformly decent" from "great on easy draws, dead on hard ones". Only conditioning the evaluation on the stratum reveals the split, and curricula should raise difficulty on measured stratum success, not on schedule.

    Applies when

    • training reward hits records while a fixed-condition eval degrades
    • wide DR on an axis where deployment sits at one known value
    • designing curricula for difficulty axes (delay, push, terrain)
    “常量 latency DR (0,0.06) 从零训被证伪——Isaac reward 129 历代最高,但 --delay 2 冒烟 iter1500 起 3/3 全摔持续到早停(聚合奖励掩盖重延迟尾部子群体失败,可用窗口只剩 500/1000)。… 采样上限 0.02 起步 … ≥90% 才 +0.01s,0.06 封顶,棘轮只升不降。”
    train/OMNI_V0_SPEC.md § 3. S1.5(s1e 训练塌方复盘)
  • Calibrate a wall penalty by measuring healthy and sick policies - healthy pays ~0, the disease pays a wallcalibrate-threshold-between-healthy-and-sick
    Mechanism understoodwalkreward-shapingreward-shapinggate-battery

    Calibrate every threshold penalty by evaluating its exact formula on replays of at least one healthy and one sick policy: place the threshold between their distributions, size the weight so the sick policy pays a decisive fraction of tracking while the healthy one pays ~0, and pre-compute neighboring thresholds for cheap adjustment.

    Symptom

    Real walk_v7 occasionally clipped its own legs (stance narrowed to 133 mm mean vs nominal 214.5); a foot-distance penalty was needed, but an uncalibrated threshold/weight risked either doing nothing or becoming a reverse barrier.

    Context

    The term (relu(d_min - lateral foot distance), measured in the base yaw frame because world-frame y is meaningless after turning) was calibrated before training by replaying three known policies through the exact reward formula at the acceptance operating point: healthy v5 (183 mm) pays 0.6% of tracking - effectively free; narrowed v7 (133 mm) pays 22% - effective widening pressure; collapsed v8 (93 mm) pays 55% - a wall. d_min 0.16 was placed deliberately between healthy and sick, with alternative thresholds (0.14/0.18) pre-computed in the tool output for later adjustment. The shape self-check was named as a standing question: "先问'零代价的选项是什么'" - the zero-cost region must be exactly the desired behavior.

    Change

    feet_lateral_distance added at d_min 0.16 / weight -10, with telemetry expectation pre-registered (should decay toward 0 as stance learns >160 mm; if bow-legged over-widening >214 appears, only then discuss an upper bound).

    Outcome

    v9_probe on hardware: no leg contact ("没碰腿(N2 兑现)"), stance min 145/126 mm green; the term's zero-cost design left healthy gait untaxed.

    Mechanism

    A relu threshold penalty defines a free region and a priced region; its correctness is entirely in where the boundary sits relative to the healthy and pathological distributions. Replaying known-good and known-bad policies through the exact formula measures both distributions in the term's own currency, making the threshold and weight a placement decision instead of a guess.

    Applies when

    • adding any relu/threshold-style wall penalty
    • a safety margin (foot distance, joint limit, clearance) needs enforcement without taxing normal behavior
    • choosing between candidate thresholds for a new term
    “形状自检 (v6 横向组/v8-B 的教训 —— 先问"零代价的选项是什么"): 标称站距付 0, 健康步态付 ~0, 收窄才付费 … v5(健康) 183 mm … 0.6% ≈ 免费 | v7(收窄) 133 mm … 22% —— 有效推宽 | v8(塌陷) 93 mm … 55% —— 墙 … d_min=0.16 恰在 v5(183)与 v7(133)之间”
    train/WALK_V9_SPEC.md § 2. N2 —— 脚距惩罚(已完成权重预标定)
  • The shipped checkpoint was chosen by scanning checkpoints on the full gate - neighbours 100 iterations apart failed 1 and 38 cells, late checkpoints degraded - never by taking the last one, and training stopped on signals, not on a schedulecheckpoint-choice-is-a-full-gate-scan
    Replicatedonelegsim-evalfork-selectiongate-battery

    Choose a release checkpoint by running the full acceptance battery over a band of checkpoints (including the transfer axis), stop training on measured signals rather than a fixed iteration count, and expect adjacent checkpoints to differ sharply.

    Symptom

    Gate results moved sharply and non-monotonically between checkpoints of the same run, and the last checkpoint was often not the best.

    Context

    One-leg V0r1: the 2,000 neighbourhood was best; from 2,500 on the nominal gates degraded (late overtraining); 2,000 itself had one real micro-hop (17.7 mm over 5 frames); 2,300 was all green and shipped. V0r2: failed cells per checkpoint 2,000:19, 2,100:38, 2,200:1, 2,300:3, 2,400:27, 2,500:12, 3,000:18 - 2,200 shipped. The recovery line learned the same from the other side: stopping v2_6 early at a scheduled point left a policy whose re-fall rate had spiked to 9-22% before consolidation healed it ("stop on signals, not on the schedule"), and a continuation's transfer decayed checkpoint by checkpoint while Isaac stayed perfect.

    Change

    The acceptance rule "scan checkpoints, do not look only at the last one" is written into the one-leg gates (called the S1 discipline); release candidates are chosen from the scan.

    Outcome

    Both one-leg releases were mid-run checkpoints (2,300 and 2,200) chosen by the full 40-cell battery.

    Mechanism

    PPO keeps changing the policy after the gates saturate; with no gradient toward the gate's conditions, later checkpoints wander, so gate quality is a noisy function of iteration.

    Applies when

    • picking which checkpoint of a run to export and stamp
    • a run is stopped at a fixed iteration budget
    • final-checkpoint results are worse than mid-run smoke tests
    “Isaac 侧 S1 纪律: 验收扫 checkpoint,不是只看最后一个。 … 扫描判决: 2000 邻域最优——2500+ 标称面退化(⑤③② 散挂, 晚期过训), 2000 有一例真微跳(L s100 μ1.2, 17.7mm/5帧), 2300 全绿。”
    git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §6 验收门 / §8 核查单 5 与 7
  • Score left and right separately - averages hide chirality breaking that mirror augmentation does not preventchirality-scored-separately
    Replicatedomnigate-batterygate-batteryreal-acceptance

    Report every mirrored skill as two numbers with an explicit gap budget; never accept an average, and never assume augmentation guarantees symmetry - measure it per lineage and treat breakage as hard to reverse.

    Symptom

    Policies developed quantified left/right asymmetry (e.g. C2-700 turned right at 82% but left at 67% - a 15 pp gap; push tolerance 40/40 symmetric on the root vs 17/40 on a deep-trained descendant), and averaged metrics would have reported healthy midpoints.

    Context

    The repo had policy-level symmetry-breaking evidence strong enough to make separate scoring a battery rule: "左右必须分开打分 … 平均 vy 跟踪会把它掩盖". Notably, chirality broke and never recovered even though mirror augmentation (command-level mirror_prob 0.5) was on the whole time - augmentation reduced but did not prevent asymmetry, and once broken it stayed broken through subsequent rungs. PASS conditions therefore carried explicit symmetry budgets (left/right tracking gap <=10 pp), and sim's predicted asymmetry (700: right faster than left) was flagged for direct real-robot timing confirmation.

    Change

    Battery rule: every directional skill reports left and right (CW/CCW) as separate rows with a max-gap budget; mirror augmentation treated as mitigation, not proof of symmetry.

    Outcome

    The 700-vs-A800 asymmetry gap (15 pp vs 7 pp) became a first-class selection criterion; C4 product shipped with a measured 5 pp gap.

    Mechanism

    Averaging over mirrored conditions cancels antisymmetric error exactly where it matters; and symmetry lost during training is a lineage injury (like plasticity loss) that later rungs do not spontaneously heal, so it must be gated, not assumed.

    Applies when

    • evaluating turn/sidewalk/push-recovery or any mirrored skill
    • relying on mirror/symmetry augmentation
    • selecting between checkpoints with similar average scores
    “左右必须分开打分(left/right lateral、CW/CCW turn 各自一行)—— 本仓已有 policy-level symmetry breaking 的量化证据,平均 vy 跟踪会把它掩盖。”
    train/C_LADDER_RUN.md § 5. 固定验收矩阵 (左右分开打分)
  • The +/-50 mm lateral COM randomization meant to spread the legs coincided with legs pulling IN - rolled back per its own pre-registered contractcom-dr-rollback-on-symptom
    Observed oncewalkdr-tuningdomain-randomizationattributiongate-battery

    When adopting a DR value that covers no local measurement, write its intent and rollback trigger into the config at adoption time; roll it back as the control arm the moment the symptom contradicts the intent, and promote the symptom's metric into the acceptance battery.

    Symptom

    After v7 adopted the reference developer's oversized lateral COM randomization (+/-50 mm) explicitly to force leg spread, the real robot's legs narrowed instead - lateral mean 154 mm / closest 107 mm in sim (nominal 214.5), narrower still on hardware with occasional leg contact.

    Context

    The rollback was clean because the adoption had been honest: the robot.yaml comment recorded the intent AND that the +/-50 value covered no local measurement (only a 16/7 mm measured offset existed; even the prior widening to +/-20 was subjective), plus the reference's own reported side effect (base sway) and the note "这一项要单独跑、 单独归因". When the opposite symptom appeared, v8 returned y to +/-20 mm as the control arm ("要么没起作用、要么帮了倒忙 … 按约定退回做 对照"), kept x/z untouched (a noise-level difference not worth another variable), and named the second suspect: the landing penalty itself, via the reference's own three-link chain (landing penalty -> stance narrows -> spacing penalty needed). A gate lesson was booked in the same table: v7's sim numbers had ALREADY crossed the line (154/107 vs v5's 182/147) - "这个指标本可拦下 v7" - so foot-distance became a standing acceptance row (min >120 mm, zero leg-leg contacts).

    Change

    base_com_offset_m y: 0.050 -> 0.020 (x/z kept), regenerated through the export tool rather than hand-editing derived files; foot-distance acceptance row added.

    Outcome

    A borrowed DR lever with no local measurement basis was retired the moment its symptom contradicted its purpose, at single-variable cost; the metric that would have caught it pre-hardware entered the gate.

    Mechanism

    DR ranges shape behavior through the policy's robustness strategy, which is jointly determined with every reward term; a lever that forces stance width on one robot can be dominated by a stronger narrowing pressure (landing softness) on another. Levers adopted without local measurement must carry their own rollback trigger, because there is no nominal to argue from when they misbehave.

    Conflicts

    Causality is not fully closed in the source: the narrowing may come from the landing penalty rather than the COM lever ("腿距的第二嫌疑人是 ④ 本身"); the rollback is the pre-agreed control experiment, not a verdict that the lever caused the narrowing.

    Applies when

    • importing DR ranges or behavioral-forcing randomizations from references
    • a DR lever's observed effect contradicts its documented purpose
    • a sim metric existed that would have caught a shipped regression
    “⑥ 的本意 … 是逼策略把脚分开;真机结果是脚向内收且偶发相碰——要么没起作用、要么帮了倒忙。… 注释当时就写了"这一项要单独跑、单独归因"。现在症状出现了,按约定退回做对照。”
    train/WALK_V8_SPEC.md § 3. 改动 C — 质心随机化退回(撤销 v7-⑥ 的 y 项)
  • A constant-value plant rung passed every binary gate with record scores - and shipped 60% thinner posture margins that hardware exposedconstant-value-dr-overfits-margin
    Mechanism understoodomnidr-tuningdomain-randomizationreal-acceptancegate-battery

    Randomize deployment-critical axes over a narrow band spanning the measured real support - never a single value, never a fictitious tail - and report graded margin quantities (tilt margin) next to binary gates, because saturated gates rank thin-margin and thick-margin policies identically.

    Symptom

    s2_lag1 (trained at constant 1-frame latency) posted the strongest sim gate sheet in history (20/20 everywhere) yet was unstable on hardware, while s1e (trained across the full 0-3 frame band) was the every-run-stable SOTA at the same power.

    Context

    The sim autopsy (new --delay-jitter harness modeling the BusWorker's time-varying phase drift): 18 runs across constant and time-varying delays ALL survived - time variation alone does not kill - but the tilt-margin ordering reproduced hardware exactly: s1e 7.7-9.1 deg (thickest) < s2_lag1 10.7-15.0 < s1c 16.2-18.2. Attribution: constant-value training permits precise specialization to that one value; s1e's band diversity forced cross-value robustness - "恒定 1 帧训练 vs s1e 的 0~3 帧全带——分布多样性逼出跨值鲁棒,恒定值允许精确 特化" - so the constant-rung policy's margins were ~60% thinner, fine in sim's clean world, pushed over the line by real-world disturbances. Tool lesson booked: "存活门二值饱和后掩盖裕度差" - binary survival gates saturate and hide margin differences; graded margin columns (tilt-max) belong in the report. The synthesis with the opposite failure (wide tails cause drag-glide): the proposed resolution was a NARROW uniform band (0.02, 0.04) covering exactly the real 1-2 ticks - diversity inside the measured support, no tail, no single point. s1e's root selection later leaned on the same property: its full-band latency training "预装" the delay rungs and delivered "全工况稳定裕度" that survived power derating.

    Change

    DR-on-an-axis design refined to a three-way distinction: no wide fictitious tails (drag), no single constant values (thin margins), but a narrow band spanning the measured real support; acceptance reports gained graded margin columns alongside binary gates.

    Outcome

    The tilt-margin column entered the standard report; the s1e root (band-trained) carried the C ladder while the constant-value branch was archived with its three contributions credited.

    Mechanism

    Robustness margins are shaped by the diversity of the training distribution, not just its support: a point-mass distribution lets the optimizer trade margin for on-point performance, while a band forces solutions that keep margin across the band - and binary survival metrics cannot see the difference until the margin is spent on hardware.

    Conflicts

    The narrow-band (0.02,0.04) resolution was a pending recommendation ("裁决建议(待用户)") at the time of writing; the lineage instead moved root to s1e whose full-band training predated the staged ladder - the deterministic-staging card and this card record the two failure modes the final design must avoid simultaneously.

    Applies when

    • a rung trained at a fixed plant value aces sim but wobbles on hardware
    • binary acceptance gates are all saturated across candidates
    • choosing between constant, banded, and wide DR on one axis
    “18 跑全活,时变性单独不足以击杀;但 tilt_max 裕度排序完整复现真机:s1e 7.7~9.1°(最厚)< s2_lag1 10.7~15.0 … 恒定 1 帧训练 vs s1e 的 0~3 帧全带——分布多样性逼出跨值鲁棒,恒定值允许精确特化 … 存活门二值饱和后掩盖裕度差(s2_lag1 sim 门 20/20 史上最强却真机不稳)”
    train/README.md § s2_lag1 真机不稳 × s1e 稳的 sim 对拍(2026-08-07,时变延迟实验)
  • A single-signal contact detector lied in both directions - foot height flagged 40% false flight on a walking gait, contact force alone flagged false flight during low-friction slip - so flight became force < 5 N AND sole height > 5 mmcontact-detector-single-signal-lies
    Replicatedrunsim-evalmeasurementgate-battery

    Define contact and flight events from two independent signals (force and geometry) in conjunction, validate the detector on a behaviour known not to contain the event before using it as a gate, and match the trainer's threshold when comparing across simulators.

    Symptom

    The run line needed a flight-fraction gate. The first MuJoCo version, "sole higher than 2 mm", measured 40% flight on a walking policy that never flies. Five weeks later the one-leg gate, using contact force alone, reported support-foot "flight" segments at mu 0.4 for a policy that was not hopping.

    Context

    During a walking step the toe lifts or the heel strikes with the foot pitched, so the ankle-roll origin rises a few mm while part of the sole still touches - height alone calls that flight. Switching to contact force < 5 N (the same threshold Isaac's contact reward uses) zeroed the false flight on walking. In the one-leg re-test every force-only "flight" segment had a measured sole height of 0.0 mm: the normal force chattered while slip corrections played out on low friction.

    Change

    Run line: flight = contact force < 5 N ("lift-off must be judged by contact force"). One-leg line (2026-09-16): flight = force < 5 N AND sole height > 5 mm, recorded as the same measurement lesson in the opposite direction; the gate's behavioural meaning was unchanged.

    Outcome

    With force-based detection the walking policy read 0.0 flight and the run policy's zero flight was confirmed by two plants; with the conjunctive definition the one-leg support-foot gate stopped reporting false hops.

    Mechanism

    Each signal has its own failure: geometry moves without leaving the ground (foot pitch), and forces drop without leaving the ground (slip chatter); requiring both removes both families of false positives.

    Applies when

    • writing a flight, lift-off, hop or slip detector for a gate
    • a gate reports an event the video does not show
    • reusing a detector on a different gait or floor friction
    “首版用**足底高度>2mm** 判离地, 在 omni_s1e 走路策略上测出 40% 假腾空 … 改用**接触力 <5N**(与 Isaac feet_contact_number 同源阈值)后 空检归零 (walk 策略 flight_frac 0.0)。**课文: 离地判定必须用接触力, 高度判 会把脚的俯仰当腾空**”
    train/README.md § run R1 立项 (2026-08-09): Mac 侧新工具 + 一次空检抓获
  • A pull-assist curriculum keyed to a global success share was satisfied by the categories that already worked and withdrew before prone learned anything - conditioning the criterion on prone took it from 2.5% to 98.7%curriculum-criterion-conditioned-on-lagging-category
    Mechanism understoodrecoverytraining-runcurriculumgate-battery

    Measure a curriculum's advancement criterion on the population the scaffold is meant to help; a global success share is met by whatever already works, and the help is withdrawn before the lagging case learns.

    Symptom

    Under the beta-anchored action, supine and side stood reliably while prone still sat (1.3%). A pull-assist curriculum added to help it was withdrawn completely within ~790 iterations and prone moved only to 2.5% (noise).

    Context

    The pull assist follows HoST: an upward force on the base, active only when the torso is within 30 deg of vertical, scaled by body weight (HoST's 200 N on G1 = 0.583 BW -> 56 N here, steps of 5.6 N, ten levels to zero); the product must pass with no assist. It had already taught sit-to-stand in V1. In V2.1 its advancement criterion was the standing-time share over all envs (threshold raised to 0.55 because the share was already ~0.53).

    Change

    V2.2: PullAssistForce with gate_category = "prone" - only envs whose first step classifies them as prone count toward the criterion - and the threshold back at 0.35. A feasibility signal was pre-registered: if prone's share stayed near zero under the full 56 N, return to the roll-over path instead of adding force.

    Outcome

    The prone-conditioned curriculum withdrew level by level only as prone itself passed: prone 98.7%, and the four-category gate passed for the first time on the line (98.6% overall, re-falls 0%, torque gate PASS).

    Mechanism

    A pooled success share is filled by the categories that already succeed (supine/side ~53%), so the scaffold is removed on their account before the lagging category has used it.

    Applies when

    • an assist, guide force or easier setting is withdrawn by a success threshold
    • one task category lags while the pooled metric looks healthy
    • a curriculum ran to completion without changing the lagging category
    “**教训:全局站立占比阈会被存量类别(supine/side ~53%)凑够,拉力在 prone 学会前就撤光了 —— metric 设计失误,不是拉力机制失效**(它在 V1 教会过 坐→站)。 … **V2.2(已启动)**:`PullAssistForce` 加 `gate_category="prone"` —— 达标判据 只统计 prone 类 env”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §34 V2.1 判决(2026-08-10)
  • Record where every failed episode ends - an end-state confusion matrix showed all failures finishing seated and overturned a "cannot roll over" diagnosis that per-category success rates hide by constructionend-state-confusion-matrix
    Mechanism understoodrecoverysim-evalmeasurementgate-batteryattribution

    For any multi-category acceptance, report where each failed episode ends, not only which category it started in; it costs a few lines and no extra simulation, and it separates "cannot reach the goal" from "reaches the wrong basin".

    Symptom

    Prone scored 0% for three generations; the working diagnosis was "prone lacks the roll-over skill", and R0.3 spent a run adding prone-to-side roll-arc start states. It bought nothing: the 45-deg roll band itself only moved from 24.2% to 26.6% after 3,000 iterations.

    Context

    Acceptance reported success per starting category. A final-state table (lying prone / on the side / supine / seated / standing for every failed episode) was added to accept_recovery.py at R0.3.

    Change

    The confusion matrix became a permanent part of the acceptance output, and the prone diagnosis was rewritten from it.

    Outcome

    The prone, side and supine columns were all zero - every failure ended seated - and prone had righted its torso in 159/159 episodes (tilt under 30 deg in 100%). The missing ability was standing up from one specific seated configuration, not rolling over, which redirected the next rungs to foot placement and to a configuration probe.

    Mechanism

    Per-category success rates collapse "reached the wrong basin" and "never reached anything" into the same zero; the end state separates them.

    Applies when

    • a category sits at 0% and the diagnosis rests on its label
    • recovery, manipulation or navigation tasks with distinct terminal states
    • an intervention aimed at the presumed cause shows no effect
    “**① 末态混淆矩阵 —— 固化(已在 `accept_recovery.py`)。** 它给出的 "趴/侧躺/仰躺三列全 0、所有失败都终于坐姿"是本线最改变决策的一个事实, 而**逐类成功率按构造看不见它**。 … 成本十来行、零额外仿真。 … **② prone 病因更正(旧诊断作废)。** 旧:"缺翻身"。新:**prone 159/159 全部 把躯干翻正**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §15 R0.3 判决 + 三件事的判断
  • Before training a one-leg stand the spec named the cheapest cheats - hopping on the support foot, a raised foot resting unloaded, a leg tripod - and gave each a countermeasure and a gate; one still appeared and was caught by exactly those gatesenumerate-cheapest-cheats-before-training
    Observed onceonelegreward-shapingreward-shapinggate-battery

    Before training, list the cheapest behaviours that would satisfy each reward term without doing the task, give each a countermeasure in the reward and a gate in acceptance, and prove the intended behaviour is reachable with a probe - then treat any gate the policy games as evidence about the reward, not the gate.

    Symptom

    The literature's single-leg benchmark reports eight state-of-the-art general policies holding a clean one-leg stand 0 times out of 90 - they survive by sneaking steps and hops - so the task's first adversary was the policy's own cheating.

    Context

    The spec's shape self-check ("what is the zero-cost option?") listed, for the one-foot bucket: the cheapest cheat, a foot resting on the ground without load, countered by a 5 N contact threshold plus positive swing income; the second cheapest, small hops on the support foot to reset balance, countered by a continuous support-air penalty plus a gate of zero support-foot flight segments. The probe that preceded training had already seen a third: early low-lift postures "survived" by pressing the swing foot at 78-95 N, a leg tripod, removed by folding the shank back. The two-foot bucket was checked too: its zero-cost behaviour is ordinary standing, with no odd base state.

    Change

    Countermeasures and gates written before training: swing-contact and support-air penalties, gate 2 (zero swing-foot contact frames above 5 N), gate 3 (zero support-foot flight segments).

    Outcome

    The first run still found the unloaded-foot cheat (a binary reward band gave it no gradient to lift) - and it was caught, by the contact gates and the cross-simulator comparison, not discovered on hardware. The retrained V0 passed all gates 40/40, including zero support-foot flight after the flight detector was corrected.

    Mechanism

    A policy optimizes the reward, not the intent; the cheapest behaviours that satisfy the reward are predictable from the reward's structure, and a gate written for each before training turns a silent cheat into a failed row.

    Applies when

    • designing rewards for balance, contact or "hold still" tasks
    • benchmark policies are known to cheat the task
    • writing acceptance gates for a new skill
    “文献里 8 个 SOTA 通用策略在单脚站基准上 0/90 干净保持, 全靠偷步偷跳活命,这是本任务的第一反作弊对象 … 单脚桶下最便宜的作弊是"脚虚放地上不受力"——接触判定 >5 N 力阈(沿用),配 swing_height_band 正收入拉开。 … 第二便宜是"支撑脚小跳重置"——support_air_penalty 连续罚 + 验收门支撑脚腾空段=0 双保险。”
    git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §0 目标口径 / §5 形状自检(零成本选项是什么)
  • A 10 s acceptance episode left 6-7 s of standing to observe - a narrow stance held for that window and split on hardware; a gate cannot see instability slower than its own horizonepisode-length-bounds-what-a-gate-sees
    Observed oncerecoverysim-evalgate-batteryreal-acceptance

    Size the standing phase of an acceptance episode, and its disturbances and floor friction, to what deployment will impose; a pass on a short static window certifies only that window.

    Symptom

    v2_6 passed every simulation gate (success 99.6%, re-falls 0-1%) and then, on the real robot, stood up and slid into the splits several times; prone starts stood and then fell backwards.

    Context

    Acceptance ran 10 s episodes; a ~1-2 s get-up left roughly 6-7 s of static standing on a nominal floor with no disturbance. The narrow stance's lateral margin and the straight-knee stance's lack of any flex buffer are both failure modes that need time, disturbance or lower friction to show.

    Change

    The gap was booked as a known blind spot of the gate ("long-duration standing stability") alongside the task-space stance criterion; later rungs added MuJoCo friction sweeps at mu 0.4 to every checkpoint scan.

    Outcome

    The spec through §50 records the blind spot but no longer standing window or disturbance row in the recovery acceptance itself.

    Mechanism

    An acceptance episode observes only the dynamics that unfold within its horizon under its conditions; slow drifts and disturbance-triggered failures are outside it by construction.

    Applies when

    • a policy passes sim gates and fails on hardware after a delay
    • acceptance episodes are short relative to deployment use
    • stability is judged without pushes or friction variation
    “**sim 门为什么没逮住**:10 s episode 起身后只站 ~6-7 s,静态窗口内窄站距 撑得住;真机站立时长/扰动谱在门口径之外 —— 长时站立稳定性记为口径缺口。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §43 判读:sim 门为什么没逮住
  • Run acceptance under measured contact parameters - honest condim/torsional-friction flipped a false PASS into a real-matching FAILeval-plant-honesty-contact-params
    Mechanism understoodwalksim2sim-gatesim2simplant-calibrationgate-battery

    Pin the evaluation plant's contact model to measured values (contact dimension, torsional/rolling friction, mu) before trusting any gate that involves slip, impact, or drift - a gate can only fail a policy for physics its simulator contains.

    Symptom

    walk_v5 passed the old acceptance battery yet failed on the real robot (footfall force, drift, kicking) - the evaluation plant was flattering the policy.

    Context

    The battery was re-run under "honest contact parameters" - condim 4 (adding torsional contact) with measured torsional friction 0.035 - and v5 then FAILED exactly the rows corresponding to its real problems: heading 185 deg (limit 30), support-foot yaw slip 284 deg (limit 80), landing force 1.72x (limit 1.5x), suspended tilt 45.9 deg (limit 10). The slip physics depends on torsional friction, which the default contact model (condim 3) does not even simulate - a slip problem is invisible to an evaluator that cannot represent yaw friction at the foot. Term-sizing measurements for the new rewards were likewise taken under the same honest parameters (cmd 0.45, skipping the 5 s start transient).

    Change

    Acceptance harness pinned to condim 4 / torsion 0.035 (measured); verdicts issued under defaults declared non-citable for these rows.

    Outcome

    Sim acceptance verdicts began agreeing with hardware ("现在失败, 与真机一致"); the v6 fixes could be developed and validated against an evaluator that could actually see the disease.

    Mechanism

    An evaluator is a plant model too: contact dimensionality and friction values decide which failure modes exist in the simulation at all. Evaluating under default contact parameters tests the policy in a world where its real failure is physically impossible, producing structurally false PASSes.

    Applies when

    • sim acceptance passes policies that fail on hardware
    • slip/drift/impact gates run under default simulator contact settings
    • setting up a cross-simulator evaluation harness
    “accept_v2.py 已加三条判据, walk_v5 在诚实的接触参数下(--condim 4 --torsion 0.035)现在失败, 与真机一致:直行 15s 航向累计 <30° | 185° ✗ … 落脚力峰值 <1.5× 体重 | 1.72× ✗”
    train/WALK_V6_MINIMAL.md § 5. 验收
  • One fixed acceptance matrix for every rung - new skill must PASS while every old skill stays within a regression budgetfixed-acceptance-matrix-per-rung
    Replicatedomnigate-batterygate-batteryprocesscurriculum

    Freeze one acceptance matrix for the whole ladder; every promotion requires the new skill's PASS plus bounded regression on every prior skill, measured against the parent's baseline under the current (bug- fixed) metric code - and include command transitions, not just steady states.

    Symptom

    Sequential skill training silently trades old skills for new ones (C1 trained away the root's backward ability); without a constant measurement frame, each rung's numbers are incomparable and regressions hide.

    Context

    The C ladder ran the same 13-cell command matrix at 20 seeds per cell at every rung (stand; vx +0.15/+0.30; vx -0.10/-0.20; vy +/-0.10; wz +/-0.20; two vx&wz combos; two vx&vy combos), with promotion requiring "新技能 PASS 且旧技能不明显退化" - old-skill regression budget <=2/20 against the parent's recorded 20-seed baseline. For the transition-rich final rung, a command-switch block was added (forward->stop, stop->backward, forward->turn, left->right, turn->forward; survival + re-track within 2 s) because "每个 steady command 都会做 ≠ 命令切换不会摔" - steady-state success does not imply switch safety, and the joystick does switches. The C4 product's gate ran 260 cells (13 x 20) all 20/20.

    Change

    Battery frozen once, reused verbatim per rung; baselines re-measured per parent (and re-measured again after the metric-frame fix, since old baselines were taken with the buggy coordinate reading - "旧基线是坏坐标系的, 不可引用").

    Outcome

    Regressions were caught at the rung that caused them (C1's backward loss, C2's vx+0.30 decay), and cross-rung comparisons stayed valid for the ladder's whole life.

    Mechanism

    A constant matrix makes every rung's output a point in the same metric space, so "did we lose anything" is a table diff, not a judgment call; the per-skill regression budget converts previously earned PASSes into standing constraints on all future training.

    Applies when

    • designing gates for sequential skill addition
    • promoting a checkpoint to be the next rung's root
    • after any evaluation-code fix (old baselines must be re-measured)
    “新技能 PASS 且旧技能不明显退化才晋级。… C5 追加:命令切换验收(steady ≠ transition) forward→stop、stop→backward、forward→turn、left→right、turn→forward,各 20 seed,判存活 + 切换后 2 s 内是否重新跟上。”
    train/C_LADDER_RUN.md § 5. 固定验收矩阵(每级跑同一张,每项 20 seed)
  • When the training reset distribution changes, freeze the acceptance distribution separately and pin its seed - the same checkpoint measured twice differed by 3.8 and 10.2 pointsfrozen-acceptance-distribution-and-pinned-seed
    Mechanism understoodrecoverysim-evalgate-batterymeasurement

    Acceptance distributions are frozen artifacts, decoupled from whatever the training distribution becomes and evaluated with a pinned seed, and every gate row carries its sample size so that a small-n row is never read as a regression.

    Symptom

    R0.3 added easier roll-arc start states to the training resets, and the acceptance script shared the training category table; separately, the same checkpoint scored side 3.8 points and mid 10.2 points differently on two runs of the same script.

    Context

    Acceptance ran 512 parallel envs drawn from the fall categories. Had it kept following the training table, 15% of acceptance samples would have landed on states easier than prone - inflated scores and generations that could not be compared. The two same-checkpoint runs had identical per-item height medians, so the ruler had not changed; the spread was reset resampling noise (side n ~ 169, sigma 2.3%; mid n ~ 35, sigma 8.3%). The spec's "20 seeds" had always meant controlled seeds.

    Change

    accept_recovery.ACCEPT_CATEGORIES pinned to the four R0-R0.2 categories and decoupled from the training FALL_CATEGORIES; --seed 20260809 pinned, after which two consecutive runs were bit-identical. The mid row (n ~ 35) was labelled the bluntest gate.

    Outcome

    Every later generation (R0.3 through V3.1) was scored on the frozen distribution and seed, which is what let R0.3's intermediate state be read as "no measurable gain" (62.7 -> 62.1%) and R0.2's mid drop be booked as noise rather than a regression.

    Mechanism

    An acceptance set that follows the training distribution measures a moving target, and an unpinned reset draw adds sampling noise that small-n rows cannot absorb.

    Applies when

    • the training reset or command distribution changes between generations
    • repeated evaluations of one checkpoint disagree
    • a small category drives a pass/fail decision
    “**分布冻结**:`accept_recovery.ACCEPT_CATEGORIES` 钉死 §5 四类 … 与训练侧 `FALL_CATEGORIES` **解耦**。 … **side 差 3.8 点、mid 差 10.2 点**(h 中位逐项一致,证明不是尺子变了)—— 纯 reset 重采样噪声 … 钉死后两次连跑逐位相同。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §14 验收尺子的两处加固(R0.3 起生效,向后兼容)
  • The 1.5 Hz step-frequency gate was retracted - the author had misread his own actuator data, and the limit fought pendulum dynamicsgate-threshold-retracted-frequency
    Mechanism understoodwalkgate-batterygate-batteryactuator-modelingprocess

    Every gate threshold must cite its measurement and survive a first-principles sanity check; when a gate keeps failing otherwise healthy behavior, re-derive the threshold from the raw data before enforcing it again - and retract wrong gates in writing.

    Symptom

    An acceptance criterion "step frequency <= 1.5 Hz" kept failing healthy policies (walk_v4 at 2.33 Hz), and an earlier attempt to force slower stepping (v3) had killed stepping altogether.

    Context

    The threshold had been derived from the author's own actuator frequency-response measurements - but re-reading the raw table showed the misread: amplitude ratio at 2.0 Hz is 0.88 (knee) / 0.83 (ankle), acceptable; the genuinely bad point was walk_v2's 3.8 Hz at 0.58. Mechanism check agreed: the leg as a compound pendulum (L ~ 0.30 m) has a natural frequency ~1.1 Hz, swing half-period 0.45 s - the observed 2.1-2.4 Hz sits near where the leg wants to swing, and forcing 1.5 Hz "是跟摆动动力学对着干" (fights the swing dynamics). The frequency definition itself was pinned by two independent methods (contact-event counting 2.39 Hz vs FFT 2.33 Hz, agreeing): reported numbers are cycle frequency = steps per leg per second.

    Change

    Gate retracted in writing: "步频 ≤1.5 Hz 应删除或放宽到 ≤2.5 Hz"; frequency definition standardized before entering any config.

    Outcome

    walk_v4/v6's 2.1-2.4 Hz reclassified from disease to normal; the v3 failure got its probable explanation (suppressing stepping to meet a wrong gate).

    Mechanism

    A gate is only as good as the measurement and the reading behind it; thresholds inherited from a misread plot become invisible design constraints that later training obeys at real cost. Cross-checking a threshold against first-principles dynamics (pendulum frequency) is a cheap way to catch such misreads.

    Applies when

    • an acceptance threshold repeatedly fails policies that look healthy
    • thresholds were set from a single person's reading of raw data
    • a forced compliance with a gate degrades the behavior it guards
    “我当初依据自己测的执行器频响定的,但看错了区间。… 2.0~2.4 Hz 的幅值比 0.83~0.88 是可接受的;真正不行的是 walk_v2 的 3.8 Hz。… 腿按复摆算(L≈0.30 m)自然频率约 1.1 Hz … 把它压到 1.5 Hz 是跟摆动动力学对着干(walk_v3 把迈步压没了,可能正是这个原因)。”
    train/WALK_DIAGNOSIS.md § ② 撤回"步频 ≤1.5 Hz"这条验收标准 —— 是我定错了
  • The hip_roll (l+r) asymmetry scalar predicted real-robot lateral drift - promote validated sim scalars into the gatehip-roll-sum-predicts-lateral-drift
    Replicatedomnireal-acceptancereal-acceptancegate-batterysim2sim

    Hunt for cheap sim scalars that predict real-robot behaviors, validate them on direction AND ordering across multiple policies, then promote them into the acceptance battery; treat later violations as debt to justify in writing, not noise to ignore.

    Symptom

    A persistent hip_roll left/right asymmetry row in the sim2sim symmetry table had been dismissed as "calibration or mechanical asymmetry" noise; meanwhile real deployments drifted sideways by policy-dependent amounts.

    Context

    Forward-kinematics analysis reframed the scalar: both hip_rolls move the feet in +y for positive angle, so a same-signed (l+r) sum IS a lateral translation mode - the scalar is a direct lateral-drift bias estimate. Checked against real deployments: s1e (l+r = -0.0178, smallest magnitude) was the steadiest with least drift; 700 (+0.0253) drifted mildly left; A800 (+0.0267) drifted clearly left with the largest tilt 12.9 deg. Direction correct 3/3, ordering correct 3/3 (the log's heading calls it "四枚四中", four-for-four).

    Change

    The scalar was promoted into the acceptance battery as a posture-class criterion alongside tilt-max median: "hip_roll 左右不对称 |l+r| 不得比父代大" - doubling as a heat proxy (error ~ torque ~ heating).

    Outcome

    Used at every later gate; when the C4 product exceeded it by +0.005 rad (~+0.3 deg vs parent), the criterion was not silently waived - it was booked as explicit debt with a mechanism argument (the increment is task-required, far smaller than the sidewalk amplitude +/-2.2 deg) plus a related account (stand saturation 32.4% -> 37.2%).

    Mechanism

    A policy's static joint-angle bias in a translation-producing mode integrates into real-world drift; sim can measure that bias precisely and cheaply. A sim scalar earns gate status exactly when its predictions are validated against hardware in both direction and ordering - and a validated gate may only be exceeded with a written mechanism-level justification, never silently.

    Conflicts

    The log's heading says "四枚四中" (4/4) but the evidence table lists three policies and the text says "方向 3/3、排序 3/3"; the fourth instance is not shown in this file.

    Applies when

    • a real robot drifts or leans in a policy-dependent way
    • deciding which sim measurements deserve gate status
    • a validated gate criterion is marginally exceeded by a new product
    “s1e | −0.0178(绝对值最小)| 微右、最不飘 | 三者中最稳、飘最小 ✓ … A800 | +0.0267 | 左、最飘 | 明显左飘、倾角最大 12.9° ✓ 方向 3/3、排序 3/3。 → 正式纳入验收表(与「倾角 max 中位」并列为姿态类判据)。”
    train/C_LADDER_RUN.md § 3e. 顺带:hip_roll 左右不对称 (l+r) 就是横移偏置 —— 四枚四中
  • Narrowing the speed range to stop high-speed falls entrenched crouch-shuffling - judge gait quality at the speed that demands a gaitlow-speed-commands-reward-dragging
    Mechanism understoodwalkcurriculumcurriculumreward-shapinggate-battery

    Set command ranges to include speeds that physically demand the target behavior, and evaluate behavior-quality gates at those speeds; when a restriction is added to suppress a failure, check what new optimum it creates at the remaining commands.

    Symptom

    After the command range was narrowed to (0.15, 0.35) m/s (to treat walk_v1's 134% overspeed and 8.3 s fall at 0.5), the policy settled into crouched foot-dragging; tracking rose monotonically with speed (63% at cmd 0.2, 76% at 0.3, 87% at 0.45), showing low speeds were where the degenerate gait was optimal.

    Context

    The narrowing advice was the author's own and is retracted in the file: it treated the symptom (falls at speed) while reinforcing the root cause (at 0.15-0.35 m/s, shuffling in a crouch is globally optimal - the Froude number is so low that even humans would not lift their feet). A zero-cost experiment confirmed the flip side: at cmd 0.5 the same policy met BOTH tracking (81%) and clearance (23.0/23.2 mm) standards.

    Change

    Speed range widened back toward (0.15, 0.5) - upper bound deliberately slightly above the mechanically feasible ~0.44 m/s so the policy finds the boundary itself; acceptance re-pointed: gait-quality criteria (tracking, clearance) judged at 0.45-0.5 m/s, low speed kept only as a survival check.

    Outcome

    v4 -> v6 progression under the widened range delivered 87% tracking with 34 mm clearance; the "low command = drag" account was confirmed by the monotone tracking-vs-speed curve.

    Mechanism

    Command distribution is part of the reward: physics prices gaits per speed, and at very low speed the energetic optimum is no swing phase at all. Restricting training to that regime makes the degenerate gait the correct answer to the posed problem - and grading a gait at a speed that does not require stepping measures nothing.

    Applies when

    • a gait degenerates after a command-range restriction
    • quality metrics improve monotonically toward the range boundary
    • writing acceptance criteria for gait quality vs survival
    “现在看那个建议可能起了反作用:0.15~0.35 m/s 下蹲着蹭就是全局最优,抬腿反而亏。收窄治的是"高速摔倒"的症状,却强化了拖地的病根。… 验收标准里的 cmd 0.2 本身就是拖地速度(Froude 数极低,人在那个速度下也不抬脚)。accept_v2 应把速度跟踪与 clearance 的判定点改到 0.45~0.5 m/s”
    train/WALK_DIAGNOSIS.md § ② 放宽速度区间 / ① 零成本实验
  • The median of a bimodal metric lands in the empty gap - check the distribution, and never judge swing on 3 seedsmedian-hides-bimodal-distribution
    Replicatedomnisim-evalmeasurementgate-batteryprocess

    Before quoting a median or mean, look at the distribution; report suspected-bimodal metrics as mode share plus per-mode ranges, use small-seed smoke runs only to screen trends, and size the seed count for decisions by the share resolution you need (here: 20).

    Symptom

    Years of "high swing variance" and undecidable 3-seed swing readings turned out to be one fact: the metric was bimodal all along - "历代 swing 高方差与 3-seed 不可判由此定性:一直在测双稳态系统" - and every median reported from it (e.g. 12.1 mm) described a value no seed ever produced.

    Context

    Concrete instances: s2e_pd-1400's 20 seeds split 2.6-4.9 mm vs 19.3-24.0 mm with zero seeds between; s1e-500 read 3.9 mm on seeds 0-2 but 23.1 mm median over 20 seeds; a 3-seed reading of 19.3 was logged as "double-peak optimism, lesson recurrence #4". The selection re-audit codified the sampling rule: "3-seed 的 swing 读数不可判点, 只能筛带,选点必须 20-seed" - 3 seeds may screen a band, only 20 seeds may pick a point.

    Change

    Swing (and any suspected multi-modal metric) reported as mode shares plus per-mode ranges instead of a bare median; 3-seed smoke numbers demoted to band-screening; all shipping/selection decisions moved to 20-seed batteries.

    Outcome

    The "swing debt" bookkeeping was reinterpreted as basin probability (see swing-bistability-damping-switch), and checkpoint selection stopped being whipsawed by which basin the first three seeds happened to fall into.

    Mechanism

    Central-tendency statistics presuppose unimodality; on a bimodal distribution the median tracks the mode SHARE, not any achievable behavior, and small samples alias the share entirely. Mode-aware reporting (share + per-mode stats) is the only faithful summary, and the needed sample size is set by the share resolution required.

    Applies when

    • a quality metric shows chronic high variance across seeds
    • 3-seed smoke readings contradict 20-seed batteries
    • reporting swing height, clearance, or any basin-prone metric
    “中位数落在空档里,「swing 债 −11mm」实为「50% 概率掉进拖地吸引子」。历代 swing 高方差与 3-seed 不可判由此定性:一直在测双稳态系统。… swing 跨 seed 双峰 (500 在 seed0~2 只读 3.9mm, 20-seed 中位 23.1) —— 3-seed 的 swing 读数不可判点, 只能筛带, 选点必须 20-seed。”
    train/README.md § swing 双稳态定性 / s1e 选点重审
  • A single run's drift direction may be a limit cycle, not a policy bias - check the sign distribution across seedsmultiseed-sign-test-for-drift
    Mechanism understoodwalksim-evalmeasurementattributiongate-battery

    Distinguish "bias" from "broken symmetry limit cycle" by the sign distribution over many seeds; report drift as (mean, sign split), and never compare single-run drift magnitudes across versions.

    Symptom

    Net yaw over 15 s appeared to worsen from -41 deg (v2) to -84 deg (v4), inviting the conclusion that the new version drifted more.

    Context

    The Isaac-side view across 32 environments told a different story: per-env yaw was mixed-sign (20 negative / 12 positive) with mean ~0 - the drift is a limit cycle whose direction depends on initial conditions, not a systematic policy bias. The single MuJoCo run had sampled one draw from that distribution, so its magnitude could not be compared across versions as if it were a property.

    Change

    Evaluation rule: before classifying drift as systematic, run multiple seeds and examine the sign distribution; single-trajectory drift magnitudes are samples, not properties.

    Outcome

    The v2-vs-v4 drift "regression" was reclassified as not-established; later drift work (hip_roll l+r bias) used cross-policy, cross-seed evidence instead.

    Mechanism

    Symmetric dynamical systems can settle into either of two mirrored limit cycles; the selected cycle is decided by noise and initial state. A statistic whose sign is initial-condition-dependent has no meaning as a single sample - only its distribution does.

    Applies when

    • comparing heading drift or lateral drift across policy versions
    • a symmetric-looking behavior shows a consistent direction in one run
    • deciding whether to fix "drift" in reward or calibration
    “偏航反而变差(−41° → −84°):注意 Isaac 侧 32 env 的逐 env 偏航是正负混合(20/12)、均值 ≈0,说明这是极限环性质(方向随初值)而非策略偏置 —— MuJoCo 单次跑测到的是分布里的一个样本,不能当作系统性偏差。要判断需多种子统计。”
    train/WALK_DIAGNOSIS.md § walk_v4 独立验收 读法 (偏航)
  • Every gate on a penalty is an exit - to stop paying a stance tax the policy parked 2 deg outside a 30 deg uprightness gate (lunging), and, from scratch, just under a height gate (crouching); a positive band and an always-on guard fixed bothpenalty-gate-is-an-escape-hatch
    Replicatedrecoveryreward-shapingreward-shapinggate-battery

    Never gate a penalty on a state the policy can leave by getting worse; use always-on guards for what must never happen and positive, gated bands for what you want, and count every gate on a penalty as one more escape route to check in the logs.

    Symptom

    V2.9 added stance_width_task = relu(0.34 m - foot spacing) x standing gates, weight -10. Three checkpoints scored 0% on acceptance: standing height reached, feet on the ground, angular rate low, but the torso leaned 32.5/32.3/31.9 deg in a fore-aft lunge. The training dashboard read "tax paid off, base_height at full value".

    Context

    The penalty was gated by uprightness (tilt < 30 deg) and standing height; its tax had no time gate (500 steps x -1.65) while the standing income sat behind a 3 s zero gate (~325 steps), so leaning just past 30 deg lost a little gated income and saved the whole tax. base_height has no upright gate, so the lunge still collected it. The first metric also measured full horizontal spacing, so a staggered lunge counted as "wide".

    Change

    Two laws written down: a penalty may only carry gates the policy cannot escape by getting worse (make it an always-on guard) or it becomes a positive band ("not earned" is not "escaped"); and width is measured laterally in the base yaw frame. From scratch (V3.1 P1) with the lateral metric but the same gates, the policy parked just under the height gate instead (base_height 1.176/1.5, h ~ 0.30 m against a 0.3264 gate), the beta curriculum never advanced in 1,700 iterations, and the run was stopped early. P1b flipped the penalty into a positive band +2.0 x clamp(lateral / 0.34) x standing gates; P1c added yaw_guard = -5 x relu(|hip_yaw| - 30 deg), always on, no gate, no exemption.

    Outcome

    P1b: lateral stance 0.364 m, all four categories 100%, MuJoCo mu 1.0/0.4 both 100%. P1c: all six acceptance criteria passed for the first time on the line (hip-yaw saturation 1.2%), the guard's tax converging to -0.016 (almost never touched).

    Mechanism

    A gated tax that is not paid is saved, so the policy moves to the cheapest state just outside the gate; a gated income that is not earned is simply lost, so a positive band has no exit. HoST's style penalties are ungated or binary - the same law seen from the other side.

    Applies when

    • adding a penalty multiplied by an uprightness, height, phase or contact gate
    • a policy settles just beyond a gate threshold
    • training metrics look paid-up while acceptance collapses
    “**罚项的门 = 策略的逃生门**。带直立门的负项可以靠"变得更差"(倾出门外) 全时免税;HoST 的 style 罚全部无门控/二值恰是同一律的反面实证。修律: 负项只许挂"变差逃不掉"的门(上限护栏 always-on),或改正向 band (收不到 ≠ 逃掉)。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §46 定案(血统内第四败 + 两条新律)
  • The walk phase machine structurally cannot express flight - rebuild the representation for duty < 0.5, teach flight with a mask tax, never a cliff bountyphase-machine-structural-limits
    Mechanism understoodrunreward-shapingreward-shapingcurriculumgate-battery

    When a new gait changes the contact pattern's structure, audit whether the phase/mask representation can express it and rebuild the representation if not; teach the new contact pattern with graded mask-mismatch pressure and count it in acceptance with artifact-proof definitions (minimum segment length), never with cliff bounties.

    Symptom

    Running requires both feet airborne, but the walk-era phase machine switches legs by the sign of sin(phase) - with duty < 0.5 the two swing windows must OVERLAP during flight, which a sign-switching representation cannot express at all.

    Context

    The run phase machine was re-architected rather than patched: per-leg phases (left = phi, right = phi+0.5 mod 1) with leg_phase < duty defining stance, aligned to the walk sin convention at duty=0.5 so the machines agree where their domains overlap. Flight is taught by the SAME mechanism that once cured foot-dragging, direction reversed: in the two planned flight windows the contact mask is (0,0) and feet_contact_number_duty charges -0.3 per foot still on the ground - a mild ~0.16/step tax, deliberately NOT a cliff: "悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文". The reference shape (half-sine bump over swing progress) is zero at window boundaries by construction, eliminating the clearing-window that the C4 probe measured to cost 13-19% on non-sinusoidal references. The shape self-check ("零代价选项是什么") was run on three behaviors: standing pays ref everywhere (known cmd=0 stepping risk, booked), walking pays only the flight-window tax, proper running collects full marks.

    Change

    New rewards.py run section (leg_phase_duty / stance_mask_duty / ref_run / clearance_run / contact_duty) with the walk versions untouched byte-for-byte; flight acceptance metric defined with a segment-length floor (>=40 ms to count) so numeric contact flicker cannot fake flight.

    Outcome

    Flight became expressible and taught by a calibrated mild pressure; the walk lineage's phase code stayed frozen as its own contract.

    Mechanism

    A phase representation defines which contact patterns exist in the reward's vocabulary; duty cycling below 0.5 introduces states (double-flight) outside a half-period sign convention's language, so no weight tuning can teach them. And rare desirable events taught by cliff-shaped bounties invite hacks (jumping in place); the graded mask tax prices the planned pattern without creating a jackpot.

    Applies when

    • extending a walking stack to running/jumping (duty < 0.5)
    • a desired contact pattern never appears despite reward increases
    • defining flight/contact acceptance metrics
    “duty<0.5 时摆动窗 (1−duty)T > T/2,两腿摆动窗在腾空段重叠 —— sin 符号切腿的机制结构上表达不了"双脚同时在空中"。… feet_contact_number_duty 对"还踩着地"持续 −0.3/脚 —— 与 walk 治拖地同一机制,方向相反。不腾空的税 ~0.16/步 … 梯度温和不构成悬崖(悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文)。”
    train/RUN_V0_SPEC.md § 5. 相位机设计
  • After installing the measured plant, re-run all generations paired on old and new plant - identity metrics carry verdicts, physics metrics re-baselineplant-swap-invariants-vs-shifts
    Mechanism understoodwalksim2sim-gatesim2simplant-calibrationgate-batteryattribution

    Treat every plant upgrade as an era boundary: re-run the retained policy set paired (same seeds/flags) on both plants, carry forward only verdicts whose metrics proved plant-invariant, re-baseline the rest - and mine the systematic shifts as measurements of the old plant's biases.

    Symptom

    With the plant finally fully measured (weighed masses 9.792 kg, bench-identified armature, in-situ friction), no historical sim number was comparable to new runs - "历史 sim 数字跨纪元不可比" - and it was unknown which historical verdicts still held.

    Context

    The era-2c cross-test ran all seven walk generations on the complete plant under one harness (14/14 survived), then re-ran the same 14 configurations on the OLD plant retrieved from git, same flags and seeds, with a self-check (one historical record reproduced digit-for-digit). The split was clean. Policy-identity metrics moved essentially zero across the plant swap - dominant frequency (v7's period-doubling 1.20 -> 1.20), knee amplitude (9.9 -> 10.0), foot distance (+/-2 mm), saturation (100 -> 100, 0 -> 0) - so all seven cross-generation verdicts (freeze signature, saturation-line closure, knee-collapse location, slip-penalty accounting, N2 lineage, v6's balance, v11's triad) were re-confirmed on the honest plant. Plant-physics metrics shifted systematically with ordering preserved: slip down 10-25% (measured friction makes ground-twisting costlier), landing vertical velocity down 15-50% ("旧 plant 高估落地 凶度"第二次独立证实), landing force mixed (mass up 2.2% vs friction braking the swing - two effects fighting). Exactly ONE behavior-level change: v8's low-speed period-doubling vanished (1.30 -> 2.50, lift normalizing) - confirming it had been machine-dependent bifurcation-edge behavior that armature+friction push off the knife edge, while v7's period-doubling stood untouched: saturation-freeze-driven, a policy property, not a numerical accident.

    Change

    Era re-baselining protocol: after any plant upgrade, one paired same-seed sweep of all retained generations on old and new plant; verdicts keyed to identity metrics carry over, thresholds re-read against the new-plant table, and the differences themselves become plant-physics findings.

    Outcome

    Seven verdicts survived with evidence rather than assumption; two causes of the period-doubling family were separated with plant-side proof; and the sim's landing-violence overestimate was independently confirmed a second time.

    Mechanism

    A policy's structural properties (frequencies, amplitudes, frozen joints) are functions of its weights and survive plant changes; contact-mediated quantities are joint properties of policy and plant and shift when the plant becomes honest. Pairing seeds across plants isolates the plant's contribution exactly, so the sweep both validates history and measures what the old plant had been lying about.

    Applies when

    • installing measured masses/armature/friction into the sim
    • historical thresholds are cited across a plant change
    • a hardware-only behavior might be bifurcation-edge sensitivity
    “策略身份指标逐位不动:主频(v7 1.20→1.20)、膝摆 … 这些是策略属性,plant 换代携带无损,历史定论因此全部成立。… 唯一行为级变化:v8@0.15 的倍周期消失(主频 1.30→2.50…)——印证当时"分岔边缘、机器相关"的判定:armature+摩擦把 v8 推离刀锋;v7 的倍周期纹丝不动(1.20→1.20),它是饱和冻结驱动的深层属性,不是数值巧合。”
    train/README.md § 纪元 2c 全代同机横测 (2026-08-04): 完全体 plant 上历史结论全部存活
  • Every rung gets written stop criteria - hit any one, stop; tightened when priors say results should come fastpreregistered-stop-criteria-per-rung
    Replicatedomnitraining-runprocessgate-battery

    Freeze per-rung stop criteria (old-skill floors, new-skill deadline, oscillation signature, known pathology signatures) before training, stop on first hit - and shorten the deadline in proportion to how fast your mechanism says results should appear.

    Symptom

    Continued training past the point of degradation had previously destroyed capabilities (C1 trained away the root's backward skill); without hard stop rules, sunk-cost reasoning keeps runs alive too long.

    Context

    Standard rung protocol - eval_c_matrix every 100 iters at 5 seeds with a fixed criteria list, frozen before training ("开训前写死,事后不许改"): (1) stand or vx+0.15 at <=3/5 for two consecutive points; (2) vx+0.15 tracking <70% for two consecutive points (vs recorded root baseline); (3) oscillation signature - survival repeatedly crossing zero between adjacent checkpoints, stop on first occurrence, do not wait for confirmation; (4) new-skill-no-progress deadline (e.g. wz tracking <30% or still same-signed after iter 400); (5) a known lineage pathology signature (s1c arm-B: scatter -> half-recover -> collapse).

    Change

    Stop budgets scale with prior knowledge: when C4-redo4's feed-forward was already proven open-loop, the no-progress deadline was tightened from +500 to +100 iters ("前馈是开环就给 71~96% 的 … 若 +100 还没有 … 不值得再烧").

    Outcome

    Multiple rungs were stopped exactly on criterion (C4-redo criterion 4 at iter 1200; redo2/redo3 at absolute 900), converting each into a clean hypothesis test instead of a drifting run; no rung burned its full cap on a dead hypothesis.

    Mechanism

    Degradation of retained skills is a lineage-level injury that later training often cannot undo, so detection latency is capital loss; and a stop rule written before the run cannot be bent by hope. Tightening deadlines when priors predict immediate results converts "no result yet" into evidence against the mechanism.

    Applies when

    • starting any resumed/curriculum training rung
    • deciding whether to keep training a run that shows early regression
    • a mechanism-backed change should produce results immediately
    “每 100 iter eval_c_matrix.py --seeds 5,命中任一即停 … 存活在相邻 checkpoint 间反复跨越 0(出现即停)… 收紧版:绝对 iter 900(= +200)时 vy±0.10 跟踪仍 <30% 即停”
    train/C_LADDER_RUN.md § 3h. 预注册停梯判据(开训前写死,事后不许改)
  • Push-test protocol - positive side first, fragile side spotted, axes aligned in the log, and cross-machine push counts stay qualitativepush-test-chirality-protocol
    Mechanism understoodomnireal-acceptancereal-acceptancegate-batteryprocess

    Order disturbance tests from the robust side to the fragile side with protection scaled to sim-measured asymmetry, align and log frame conventions before testing, and treat cross-domain disturbance counts as qualitative evidence only.

    Symptom

    Hand-push testing on hardware risked falls on a side sim had already flagged as fragile, and push counts invited apples-to-oranges comparison with sim numbers.

    Context

    Sim chirality was explicit: descendants were far more fragile in -y (fric-3000@kd1.2: +6 N*s survived 15/20 vs -6 N*s only 3-9/20) while the s1e control was perfectly symmetric (40/40). The protocol therefore: push the positive direction first, keep a spotter for the negative side; before any push, record which real-robot side corresponds to sim's +y in the log ("上机前对一次坐标"); and - citing the chaos lesson ("混沌课文") - real push results are used only as qualitative corroboration, never compared numerically with sim survival counts across machines.

    Change

    Push testing became a scripted, chirality-aware protocol with frame alignment as a logged precondition and an explicit epistemic limit on cross-domain count comparison.

    Outcome

    The fragile side was tested with protection informed by sim's quantified asymmetry; logs stayed interpretable because the frame correspondence was recorded before the first push.

    Mechanism

    Disturbance-response chirality is a real, quantifiable lineage property, so test order should follow measured fragility; and perturbation outcomes are chaotic in the details (divergent trajectories from tiny differences), so counts do not transfer across domains even when qualitative rankings do.

    Applies when

    • planning push/disturbance tests on hardware
    • sim shows directional asymmetry in disturbance survival
    • someone proposes comparing real push counts to sim counts
    “先正向后负向, 负向留人扶 —— sim 手性明确: 后代在负 y 向显著更脆 (fric-3000 @kd1.2: +6 N·s 15/20 vs −6 N·s 3~9/20), 而 s1e@0.8 两向 40/40 完全对称。上机前对一次坐标 … 跨机不做二值结论 (混沌课文): 真机推力只作定性对照, 不与 sim 计数对比。”
    train/REAL_RUN_S2.md § 3. 抗推 (可选, 人手推; 做则按此协议)
  • When the eval proxy has a known systematic bias, gate on within-proxy differences, not absolutesrelative-metrics-survive-proxy-bias
    Mechanism understoodwalkgate-batterygate-batterysim2simattribution

    Where the evaluator is known-biased, design gates as within-evaluator contrasts (left vs right, A vs B, pre vs post) that cancel common-mode error; reserve absolute thresholds for quantities whose proxy calibration has been checked.

    Symptom

    MuJoCo systematically overestimated yaw turn gain (1.58/2.45 where Isaac measured 1.04/1.02), and friction alignment recovered only part of the gap (mu 1.0 -> 0.6 pulled it to 1.43/2.27) - so any absolute turn-gain acceptance threshold in MuJoCo would be judging the proxy's bias, not the policy.

    Context

    The acceptance criterion was rewritten to use only the left/right difference of the turn gain (standard: <20%): both directions pass through the same biased proxy, so the bias largely cancels in the difference while the policy's chirality - the thing being gated - survives. The absolute-gain row was dropped: "转向只判左右差,不判绝对值". A parallel task was still opened to align MuJoCo contact parameters to Isaac's (mu, restitution), since drift proved highly friction-sensitive (straight-line drift -65 deg at mu 1.0 vs +1 deg at mu 0.4) - bias reduction and bias-robust metrics proceeded together.

    Change

    Gate metric changed from absolute turn gain to left-right gain difference; proxy-alignment work scheduled separately rather than blocking acceptance.

    Outcome

    Turn acceptance became meaningful across proxy versions (v5 70% -> v6 19% difference measured the real improvement) while the absolute bias question was pursued without holding the ladder hostage.

    Mechanism

    A systematic multiplicative or additive proxy bias applies to both arms of a mirrored measurement; differencing (or ratioing) mirrored conditions cancels the common-mode bias to first order, leaving the asymmetry signal. Metrics built this way remain valid while the proxy is imperfect - which it always is somewhere.

    Applies when

    • a sim proxy disagrees with the trainer or hardware in absolute terms
    • writing acceptance thresholds for direction-paired skills
    • proxy calibration work would otherwise block a ladder
    “转向只判左右差,不判绝对值 —— MuJoCo 的偏航增益系统性高估(实测 1.58/2.45 vs Isaac 1.04/1.02),摩擦只能解释一部分(μ 1.0→0.6 仅拉到 1.43/2.27)。差值是相对量。”
    train/WALK_V5_SPEC.md § 6. 验收
  • Removing a foot-spacing wall passed every simulated gate and made the feet collide on the real robot - nothing priced stance width in the two-foot phase, the policy narrowed to the simulator's self-collision floor, and real calibration offsets closed the last millimetres; the wall came back with a gateremoved-wall-returns-on-hardware
    Observed onceonelegreal-deployreward-shapinggate-batteryreal-acceptance

    When a constraint is removed, name what will govern that quantity instead and add a gate for it; never let a simulator's collision floor be the margin, and when a gate is exceeded by a hair, record the exact numbers and hand the release decision to a person instead of quietly passing it.

    Symptom

    On the first real-robot try of oneleg_v0 (2026-09-16) the two feet collided; the user also judged the folded foot not high enough.

    Context

    The V0 reward table had dropped the feet_lateral_distance wall because it seemed to conflict with the hip adduction single support needs. In the two-foot command bucket no remaining term governed stance width, so the policy drifted narrower until the simulator's self-collision stopped it; the sim acceptance had no foot-spacing gate, so 40/40 said nothing about it. On the robot, calibration offsets consumed the margin.

    Change

    V0.1: the wall restored (-10, minimum 0.16 m), re-checked against measured numbers (a swing-phase lateral spacing of ~148 mm costs 0.12 per step, acceptable); fold weight 0.8 -> 2.0; a ninth gate: minimum foot spacing >= 100 mm and zero leg-contact frames. The removal was kept on record.

    Outcome

    oneleg_v0_1 (V0r2 model_2200) passed 39/40 with the spacing gate 40/40. The single miss (a 15.4 deg tilt transient against a < 15 deg limit during a side switch, steady 6.9 deg, everything else green) was recorded with its numbers and released for the user to overrule.

    Mechanism

    An unpriced degree of freedom drifts to wherever the simulator stops it; if that stop is the simulator's own collision model, the policy's margin on hardware is whatever the calibration error leaves.

    Applies when

    • dropping a reward term that looked redundant or conflicting
    • hardware shows a failure no simulated gate measures
    • a release candidate misses one gate row by a small amount
    “V0 撤墙被真机证伪(2026-09-16):双脚桶没有任何项管站宽,策略贴 sim 自碰撞底线收窄,真机标定偏差一吃**双脚相碰**。 … min ≥ 100 mm 且腿碰 0 帧(eval_straight 同判据)—— … V0 真机双脚相碰暴露 sim 门未看脚距的缺口 … L s2 标称 tilt 瞬态 15.4°(门限 <15, 超 0.4°, 稳态 6.9°, 该跑其余全绿)——换侧瞬态蹭线, 判定放行留档, 用户可否决。”
    git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §4 feet_lateral_distance 行 / §6 验收门 ⑨ / §8 核查单 7
  • The best checkpoint to SHIP is not the best checkpoint to CONTINUE FROM - maturity is capital against adaptation shockroot-maturity-vs-product-quality
    Mechanism understoodomnifork-selectionfork-selectioncurriculumgate-battery

    Decide shipping points and fork roots separately: gates rank products, but a root candidate must prove itself by surviving a continuation under the next rung's shift (dual-arm if in doubt) - and prefer the more-trained point as root when product metrics conflict with maturity.

    Symptom

    A band re-audit found s1e-300 beat the incumbent root s1e-500 on nearly every quality gate (stepping 19/20 vs 13/20 with historically-best 26.9 mm swing, speed gate 14/20 vs 2/20, heading 26 vs 54 deg/20 s) - suggesting the root had been mis-picked and the younger point should take over.

    Context

    The dual-arm control settled it the other way: continuing the S2 PD rung from s1e-500 adapted smoothly (3/3 smoke throughout), while the b300 control arm (same config, from s1e-300) fell into a survival valley under the PD shock (+100 iters: 1/3 -> 0/3), never climbed out within budget, and its 800-iter product scored 13/20 survival - eliminated. Verdict: "幼年点自身指标再好也扛不住新 DR 适应冲击, 成熟度是本钱,s1e-500 根被数据背书" - a young point's own metrics, however good, do not survive new-DR adaptation shock; maturity is capital. The audit still yielded value: the band scan (200-1000, per-100) mapped the lineage's arc (200 dragging -> 300 peak -> 400+ decay -> 900+ drift blowout), and 300 remains the better PRODUCT answer for shipping-as-is questions.

    Change

    Selection doctrine split into two questions with different answers: best-product point (quality gates at the point itself) vs best-root point (survives adaptation shocks; more training age = more capital), each decided by its own evidence - and root claims settled by a dual-arm continuation test, not by point metrics.

    Outcome

    s1e-500 kept the root role with data behind it; the S2e ladder built on it passed rung after rung, while the b300 line was closed at the cost of one control arm.

    Mechanism

    Early checkpoints sit near sharp optima with less accumulated robustness structure; their headline metrics reflect the narrow training distribution, not resilience to distribution shifts. A continuation rung is itself a distribution shift, so the root property being selected for is shock tolerance - observable only by actually continuing, never by static gates.

    Applies when

    • a younger checkpoint outscores the current root on quality gates
    • choosing the base for a robustification or command ladder
    • a continuation run stalls in an early survival valley
    “b300 对照臂 … PD 冲击下存活谷(+100 起 1/3→0/3),预算尽未爬出,800 档 20-seed 存活 13/20 出局——幼年点自身指标再好也扛不住新 DR 适应冲击,成熟度是本钱,s1e-500 根被数据背书”
    train/README.md § omni_s2e_pd (b300 对照臂) / s1e 选点重审
  • A nonzero response with the same sign for + and - commands is bias, not abilitysame-sign-response-is-yaw-bias
    Replicatedomnisim-evalmeasurementgate-batteryattribution

    Before crediting any directional skill, test both command signs: response must flip sign with the command; a same-signed pair is a bias to subtract, not an ability to report.

    Symptom

    Root-selection probe showed nonzero wz "tracking percentages" on turn commands, tempting the read that candidates could partially turn.

    Context

    During C-ladder root selection, s1e-500's measured yaw rate was +0.084 rad/s for cmd +0.3 and +0.093 rad/s for cmd -0.3 - same sign both ways. The same check on the C2 baseline gave wz+0.20 -> -0.13 and wz-0.20 -> +0.12 (again same sign), while the alternative root s2e_pd-1400 gave +0.16 / -0.16 - opposite signs, i.e. a genuine 16% command response.

    Change

    Reading corrected and written into the execution sheet: percentages on directional commands are meaningless unless the +cmd and -cmd responses have opposite signs; all three candidates were re-classified as "cannot turn, cannot sidewalk - C2/C3/C4 learn from zero". Acceptance criteria thereafter required "tracking >=50% AND left/right opposite-signed".

    Outcome

    Prevented crediting turn/sidewalk ability that did not exist; the antisymmetry clause became a standing part of every turn and sidewalk PASS condition (C2, C4, C4-redo levels all carry "且左右反号").

    Mechanism

    A constant yaw (or lateral) bias projects onto any command's sign convention and shows up as fake fractional tracking; only sign-antisymmetry under command reversal distinguishes a feedback response to the command from an open-loop offset.

    Applies when

    • evaluating turn/sidewalk/any signed-command tracking percentages
    • a candidate shows partial tracking on an axis it was never trained on
    • writing PASS criteria for a new directional skill
    “C2/C3 那些非零的 wz 百分比不是转向能力 —— 转向+ 与 转向− 的实测同号(s1e:cmd +0.3 → +0.084,cmd −0.3 → +0.093 rad/s),那是恒定偏航偏置。… 三个候选都不会转、都不会侧走。”
    train/C_LADDER_RUN.md § 0. 读数纠正(重要,别引错)
  • A joint frozen at the action clamp pays zero action_rate forever - penalize pre-clip saturation to make the cheat cost moneysaturation-cheating-zero-rate-cost
    Mechanism understoodwalkreward-shapingreward-shapingaction-rategate-battery

    Whenever actions are clipped and any smoothness/rate penalty exists, add a pre-clip saturation penalty so living at the clamp costs more than oscillating - and audit for frozen-at-clamp joints (action std ~0, |a| at exactly the clip value) as a standing acceptance row.

    Symptom

    With action_rate_l2 raised to -0.2, walk_v7's hip_pitch actions froze at exactly +/-1.000 (the clamp), reproduced bit-for-bit on hardware (splits frozen at +/-0.35 rad); the gait-shaping term joint_pos_ref collapsed to 0.026-0.035. The repo had died in the same trap once before (walk_v0: four joints pinned at +/-1.0).

    Context

    Mechanism: a joint pinned at the clamp has action-rate cost exactly zero and forever zero - under a strong smoothness tax, "push to the clamp and freeze" becomes the dominant optimum. Lowering the weight (-0.2 -> -0.1) only reduces temptation; the frozen state still costs nothing, so the structural fix adds action_saturation = sum(relu( |a_raw| - 0.9)) at weight -1.0, computed on the PRE-clip network output - post-clip, |a|=1.01 and |a|=3 punish identically and the out-of-range gradient dies (v0's old disease: mean |a| 1.71 soaked in saturation). Economics: freezing at |a|=1.0 now pays 0.1/joint/step (two hips = 40% of alive) vs ~0.0004/step for the healthy reference oscillation - the cheat flips from free to ~250x negative. Honest limits were recorded: A1 does not forbid freezing at 0.89 (the anti-freeze pressure must come from the oscillation demand of joint_pos_ref), and the alternative "rate on post-clip target" was rejected as 换汤不换药 - a pinned target also has zero rate.

    Change

    v8-A: add action_saturation (-1.0, thresh 0.9, pre-clip) AND halve action_rate_l2 (-0.2 -> -0.1, still 3.3x the v5 value); success criterion pre-declared (joint_pos_ref telemetry returns to v6 scale).

    Outcome

    Booked as the structural repair of the v7 freeze; also fixed a config hygiene trap discovered on the way - action_rate was assigned twice in __post_init__ (v5 comment line then v7 line), merged to one assignment "别再留两处赋值给下次审计埋雷".

    Mechanism

    Clipping creates a zero-gradient, zero-cost absorbing region in action space; any penalty on action derivatives makes that region strictly optimal once entered. Only a penalty on clamp proximity itself (measured pre-clip so depth of violation is visible) restores a slope out of the absorbing region.

    Applies when

    • joints sit at exactly the action clip with near-zero variance
    • raising a smoothness penalty degrades gait amplitude
    • shaped-oscillation terms collapse after a rate-weight increase
    “钉死在钳位的关节 action_rate 代价精确为零且永远为零;−0.2 之下"推到钳位冻起来"成了压倒性最优 … 本仓第二次栽在同一坑(walk_v0 死于四关节钉死 ±1.0)。回调权重(−0.2→−0.1)只降低诱惑不消除作弊 … 算在 clip 前的原始网络输出上 … 作弊收支从"白赚"变成"倒贴 ~250 倍"。”
    train/WALK_V8_SPEC.md § 1. 改动 A — 治饱和作弊
  • Cross-simulator gate (Isaac Lab to MuJoCo) comes before any hardware attemptsim2sim-gate-before-sim2real
    Replicatedwalksim2sim-gatesim2simprocessgate-battery

    Gate every policy through a second simulator with an independently built plant before hardware; treat sim2sim failure as a contract or overfitting bug, and sim2real failure after a sim2sim pass as a plant/actuator gap.

    Symptom

    A policy that only ever ran in its training simulator carries untested dependencies on that simulator's solver, contact model, and defaults; the first place those dependencies surface should not be the real robot.

    Context

    Standing order of operations for the whole Lucen program, recorded as the opening line of the experience log: train in Isaac Lab, gate in MuJoCo, only then go to hardware. The MuJoCo side is the same plant used for evaluation batteries, so a sim2sim pass also validates the exported policy + contract (obs ordering, scales, defaults) outside the training stack.

    Change

    Pipeline rule adopted - every checkpoint must pass the MuJoCo evaluation battery (sim2sim) before it is considered for real deployment (sim2real).

    Outcome

    Used generation after generation as the cheap filter; hardware sessions only ever received policies that had already survived a second simulator.

    Mechanism

    Two simulators disagree exactly where a policy is overfit to simulator-specific artifacts (contact softness, integrator, default parameters, obs conventions); a cross-sim transfer catches contract bugs and solver overfitting at zero hardware risk, so real-robot failures that remain are attributable to genuine plant/actuator gaps.

    Applies when

    • planning the path from training to first hardware trial
    • exported policy behaves differently outside the training framework
    • triaging whether a real-robot failure is contract vs plant
    “先sim2sim - 从isaaclab 到mujoco / 再sim2real”
    Experience.md § opening lines (1-2)
  • Single-impulse push recovery is a binary chaotic quantity - cross-machine floating-point divergence can flip the outcomesingle-impulse-recovery-is-chaotic
    Mechanism understoodomnisim-evalmeasurementgate-batteryattribution

    Never gate or compare single-event recovery outcomes across machines or domains: evaluate disturbances as survival distributions over phases and seeds, compare longitudinally on one machine, and treat any single-point cliff as unconfirmed until it survives the statistical protocol.

    Symptom

    Mac evaluation found a hard "0.8 N*s cliff" (0/3 survival) that the training machine flatly contradicted: the identical protocol (0.8 impulse at 8 s, cmd 0.2) survived 3/3 there, and a 0.6/0.8/2/4 cross sweep survived everything.

    Context

    The verdict became a named lesson ("跨机混沌课文"): whether one specific push at one specific phase is survived depends on a trajectory that diverges across machines from floating-point differences alone - "单次冲量恢复是二值混沌量, 跨机浮点发散可翻结局". The boundary was drawn precisely: the 20-seed statistical gates DO agree across machines (established precedent), but that agreement cannot be extrapolated to single-point recovery tests. Protocol amended: disturbance evaluation uses multiple push phases (8/10/12 s), >=10 seeds, and only same-machine longitudinal comparisons; the Mac-side recommendation built on the unreproducible cliff was not adopted, while its directionally-consistent small-impulse data was kept.

    Change

    Push evaluation redefined from single-event pass/fail to multi-phase multi-seed statistics, with cross-machine comparison banned for event-level results and allowed for distribution-level ones.

    Outcome

    A false hardware-relevant "cliff" was prevented from steering the ladder (the s2e push rung decisions were made on same-machine statistics); the chaos lesson was cited again when real push tests were restricted to qualitative cross-domain use.

    Mechanism

    Perturbation recovery near the viability boundary has sensitive dependence on initial conditions; different BLAS/GPU reduction orders yield different trajectories from identical configs, so a binary outcome at one phase is machine-specific noise. Averaging over phases and seeds restores a quantity whose expectation is machine-stable.

    Applies when

    • a push/disturbance result differs between machines or sim and real
    • designing push-recovery acceptance tests
    • a sharp pass/fail cliff appears in a chaotic-regime evaluation
    “训练机上 Mac 原协议 (0.8 @8s cmd0.2) 3/3 全活 … 与 Mac 的 +0.8 0/3 直接矛盾。定性: 单次冲量恢复是二值混沌量, 跨机浮点发散可翻结局;统计门 (20-seed 八门) 跨机吻合的先例不能外推到单点恢复测试。协议改判: 抗推评测多相位 (push 时刻 8/10/12s) + ≥10 seed + 只做同机纵向比”
    train/README.md § s2e 支线终章 (跨机混沌课文)
  • The smoke watcher panicked at the wrong operating point and its score goes blind when gates saturate - treat it as a survival sentinel, not a judgesmoke-watcher-operating-point
    Replicatedomnisim-evalmeasurementprocessgate-battery

    Configure every automated evaluator at the lineage's declared deployment operating point, give it graded metrics that cannot saturate, and until then scope its authority to catastrophe-detection - never let a mis-configured watcher stop or rank a rung on its own.

    Symptom

    Two watcher misfires in one ladder: (1) during the friction rung the watcher (evaluating at kd 1.0) reported panic-level 1/3 survival from iter 3300 - falsified by the official kd 1.2 scan, because the lineage's design operating point was kd 1.2 and the watcher lacked the --kd-scale passthrough; (2) during the PD rung the watcher's early-stop score froze at iter 1050 despite ongoing drift improvements, because with all eight gates passing (constant 0/3 failures) the score has no gradient left - "八门全过恒 0/3 时 score 对漂移改善盲, s1f 课文三现".

    Context

    Both are the same category: the in-training smoke loop is an instrument with its own configuration (operating point, score design), and its verdicts are only as aligned as that configuration. The booked doctrine: "冒烟只当存活哨兵" - until the watcher evaluates at the deployment operating point with graded metrics, its role is detecting catastrophes, not ranking checkpoints; ranking belongs to the full battery at the design operating point (and the watcher's scoring was separately patched to weight survival 3x so recoveries during hard phases are not early-stopped away).

    Change

    Watcher debt booked (--kd-scale passthrough); score saturation acknowledged with graded columns planned; selection authority kept with 20-seed batteries at the declared operating point.

    Outcome

    A false panic did not abort a rung that was actually passing at its design point; a frozen score did not hide real drift gains; the instrument's authority was scoped to what its configuration can actually see.

    Mechanism

    An evaluator is itself configured (gain profile, delay, metrics); evaluating a policy away from its design operating point measures a counterfactual robot, and bounded scores saturate once binary gates pass, losing all sensitivity. Instruments need the same operating-point discipline as deployments and graded outputs to retain gradient.

    Applies when

    • an automated smoke loop contradicts the official battery
    • early-stop scores freeze while graded metrics still improve
    • lineages with non-default deployment gain/delay profiles
    “watcher (kd1.0 口径) 3300 起 1/3 恐慌被 kd1.2 正式扫描证伪为考纲外假象 —— 工作点评测口径教训: watch_ckpt 缺 --kd-scale 透传 (待补), 冒烟只当存活哨兵。… watcher score 饱和误停 @1050(八门全过恒 0/3 时 score 对漂移改善盲, s1f 课文三现)”
    train/README.md § omni_s2e_fric (watcher 恐慌被证伪) / omni_s2e_pd (500 臂)
  • A stand gate judged by survival passes a robot that wanders a meter - judge posture insteadstand-gate-posture-not-survival
    Mechanism understoodomnireal-acceptancegate-batteryreal-acceptance

    For every gate, ask what behavior the metric is a proxy for and validate its ordering against real observations; replace metrics whose ordering disagrees with reality, and demote them explicitly rather than silently.

    Symptom

    Real robot "standing" drifted 0.5-1.4 m across the floor while the sim stand gate scored a clean 20/20 - because the gate's metric was episode survival, which wandering does not violate.

    Context

    C2's stand condition was originally written as "survival regression <=2/20". A real-robot counter-example on 2026-08-08 forced the re-judgment: wandering robots survive. Cross-checking candidate sim metrics against real-robot feel showed max-tilt median ordering agreed with hands-on ranking, while displacement ordering was actually OPPOSITE to real impressions - so displacement was demoted to a reference quantity, not a gate.

    Change

    Stand PASS criterion rewritten from survival to posture: "stand tilt max median <= root baseline +1.5 deg"; displacement kept only as reference. Applied to all subsequent rungs (C4 and redo levels inherit it).

    Outcome

    Later rungs gated stand on tilt (e.g. C4 product: 7.7 deg vs parent 7.1 deg, +0.6 deg PASS); the wandering failure mode became visible to the battery instead of hidden by survival.

    Mechanism

    A gate metric is a proxy for an intended behavior; survival is a proxy for "did not fall", not "stood still". Metric choice must be validated against ground truth (real-robot feel/measurement), and a proxy whose ordering disagrees with reality on real data is worse than no metric - it steers selection backwards.

    Applies when

    • writing PASS conditions for stand/idle/hold behaviors
    • a gate passes policies that visibly misbehave on hardware
    • choosing between candidate metrics for an acceptance battery
    “stand 必须用位姿判,不能用存活判(2026-08-08 真机反证改判):站着乱走 0.5~1.4 m 时存活照样 20/20 —— C2 的 stand 条件原写「存活退化 ≤2/20」,选错了指标。改为: stand 倾角 max 中位 ≤ 根基线 +1.5°(倾角与真机手感排序一致;位移排序与真机相反,降为参考量)”
    train/C_LADDER_RUN.md § 3b. PASS 条件 ⚠️ stand 必须用位姿判
  • Add a single-point-suspension test to acceptance - the ground is a free stabilizer that hides divergencesuspension-probe-removes-free-stabilizer
    Mechanism understoodwalkgate-batterygate-batteryreal-acceptancesim2sim

    Include at least one acceptance condition that strips the environment's free stabilization (suspension, or equivalent) - the sim-passing policy that fails on hardware is often failing a condition the battery never posed.

    Symptom

    walk_v5 looked healthy in every on-ground sim test yet diverged on the real robot - the acceptance battery had never measured a condition that would have revealed it.

    Context

    The battery gained a single-point-suspension probe (robot hung, feet free): measure torso tilt while the policy runs without ground contact. v5 scored 45.9 deg mean tilt suspended - wildly unstable - which the file calls "最灵敏的失稳探针(拿掉地面这个免费稳定器)": ground reaction forces passively stabilize a marginal policy, so on-ground metrics saturate long before the policy's internal balance is actually sound. v6 halved it (23.0 deg, target <10 deg) - progress visible on a scale where on-ground numbers showed nothing.

    Change

    Suspended-tilt added as a standing acceptance row; run under the honest contact parameters battery (accept_v2 with measured condim 4 / torsional friction 0.035), under which v5 correctly FAILS in agreement with the real robot.

    Outcome

    The sim battery's verdict on v5 flipped from pass to fail, matching hardware; suspended tilt became the discriminating metric between v5 and v6 (45.9 vs 23.0 deg) when ground metrics differed little.

    Mechanism

    Contact with the ground closes a stabilizing feedback loop the policy gets for free; removing it exposes the policy's own attitude control authority. A metric measured only in the assisted condition cannot rank policies by the unassisted quantity that hardware will actually demand during perturbations and flight phases.

    Applies when

    • sim acceptance passes but hardware diverges
    • designing an acceptance battery for a legged robot
    • two candidates tie on ground metrics
    “单点吊那条是最灵敏的失稳探针(拿掉地面这个"免费稳定器"), v5 在地上一切正常却在真机发散, 就是因为验收从没测过这个工况。”
    train/WALK_V6_MINIMAL.md § 5. 验收
  • Three same-shaped judging errors - task metrics (survival, tracking, displacement) cannot stand in for posture metricstask-metrics-vs-posture-metrics
    Replicatedomnireal-acceptancereal-acceptancegate-batteryattribution

    Keep validated posture-class rows (tilt, per-joint L/R asymmetry, temperature) in every acceptance battery alongside task rows; when operator feel contradicts the gates, suspect the metric class before the operator - and never build a new skill on what is actually an asymmetry defect.

    Symptom

    The C2 product judged "full pass" on task metrics (A800: turn-gap 7 pp, vx+0.30 19/20) felt WORSE in the operator's hands than the half-pass 700: A800 tilted up to 12.90 deg (700: 6.64), drifted left while standing, showed larger per-joint asymmetries, and ran its hip_rolls 5 degC hotter.

    Context

    The re-judgment catalogued three same-type metric errors in one campaign: (1) stand judged by SURVIVAL - missed 0.5-1.4 m wandering; (2) stand ranked by DISPLACEMENT - ordering was opposite to real feel (tilt ordering matched); (3) chirality judged by wz-tracking GAP - measured turning symmetry while the robot's actual disease was postural left/right asymmetry, "两个不同的东西,且结论相反". Common pattern named: "我一直用「任务指标」当判据,而真机手感对应的是「姿态 指标」… 任务类指标不能替代它". The fix was already in the data: the per-joint left/right asymmetry table (printed identically by sim2sim and deploy) agreed with hardware in direction on every row - "判据可用、有预测力,我只是没把它写进 PASS 条件". Shipping decision followed the posture read: product reverted to 700 ("又一次「买到 精度、卖掉别的」"), and the C4 root moved to 700 as well, with the sharpest line of the episode: A800's left-drift "like sidewalking" is probably its frontal-plane asymmetry defect, not a capability - "在缺陷上建能力是危险的".

    Change

    Two posture quantities with demonstrated real-robot predictive power promoted into every PASS battery: tilt-max median and per-joint left/right asymmetry (both sim-computable, deploy-homologous); motor-temperature readout added to session close-out.

    Outcome

    Deployment flipped to the posture-better checkpoint; the hip_roll temperature table (43-48 degC vs 25-28) confirmed the earlier 90%-of-heat account; the run-line acceptance battery inherited the posture rows from birth ("任务类替代不了姿态类").

    Mechanism

    Task metrics measure goal attainment under the evaluator's episode definition; posture metrics measure the body state trajectory that operators, motors, and downstream skills actually experience. The two can rank candidates oppositely because task success tolerates postural pathology - so a battery without posture rows is blind to exactly what hardware feel reports first.

    Applies when

    • hardware feel disagrees with a green acceptance table
    • choosing between checkpoints that split task vs posture metrics
    • selecting the root for a skill that resembles an existing defect
    “共同模式:我一直用「任务指标」(存活 / 跟踪率 / 位移)当判据,而真机手感对应的是「姿态指标」(倾角、逐关节左右不对称)。→ 验收判据里必须有姿态类指标,任务类指标不能替代它。… A800 的「左飘像 side walk」很可能 … 是它更大的额平面不对称的表现 —— 在缺陷上建能力是危险的。”
    train/README.md § C2 选点改判 (2026-08-09): 手性判据第三次选错指标
  • Every acceptance run records video of the very rollout that produced the numbers - the seated basin, edge-standing feet, tangled legs and the narrow stance were all seen on video before, or instead of, a metric catching themvideo-as-acceptance-record
    Replicatedrecoverysim-evalgate-batteryreal-acceptancemeasurement

    Make video a default output of every acceptance run, rendered from the same rollout the metrics come from (fixed views including the feet), and watch it - posture failures are visible before any gate row exists for them; never let video replace or override the numeric gate.

    Symptom

    Posture failures in the recovery line kept arriving as things the numbers had no row for: a kneeling W-sit, feet standing on their outer edges, crossed legs, a stance too narrow to hold on hardware.

    Context

    From 2026-08-09 (user decision) accept_recovery renders offscreen by default, following one env for the whole episode and archiving the clip. The MuJoCo gate's --video renders three views (side, front, feet) of the same rollout the metrics come from, with all plant modelling (delay, push); the older replay-based renderer produced an independent trajectory without delay and was not used for acceptance. Video never gates: if rendering breaks, --no-video keeps the numeric gate running.

    Change

    Video as a default acceptance artifact, named per policy, category and view, reviewed by the user.

    Outcome

    The R0.1 prone clip showed the same kneel-sit as R0; the R3.1 failure clip showed the crossed legs; the v2_5 feet view showed edge standing and led to V2.6; the v2_6 videos led the user to order a real-robot A/B between v2_5b and v2_6. Once, the recorder did not start (P1c final acceptance) and the visual material had to be produced separately.

    Mechanism

    Metrics exist only for failure modes someone anticipated; video shows the unanticipated ones, and rendering the metric rollout itself guarantees the picture and the numbers describe the same episode.

    Applies when

    • setting up an acceptance pipeline for posture-sensitive skills
    • numbers pass but a human reviewer is uneasy
    • sim videos are rendered by a separate replay tool
    “**验收存视频(用户定 2026-08-09)**:`accept_recovery.py` 默认开 Isaac 离屏 渲染,跟拍一个 env 的整局并归档 … 跪坐这类盆地在数字表出现前肉眼先看见,视频是验收的 定性存档,数字门不受它影响”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §5 R0 验收门(预注册)验收存视频(用户定 2026-08-09)
  • Under continuous 3-axis uniform sampling, pure straight-line walking is a zero-measure event the policy never trainedzero-measure-commands-need-mode-sampling
    Mechanism understoodomnicurriculumcurriculumobservation-honestygate-battery

    Enumerate the exact command points users will actually issue (straight, stop, in-place turn) and give each explicit probability mass via mode sampling with off-axes pinned to zero - never assume a continuous sampler covers its measure-zero subsets.

    Symptom

    "The robot drifts even in sim when told to walk straight" persisted across reward tunings - because with commands drawn as vx in [0.15,0.5] x vy ~ U(+/-0.2) x wz ~ U(+/-0.6), the event vy=0 AND wz=0 has probability zero: pure straight-line walking was never sampled even once.

    Context

    Restart evidence item #3: "纯直行是零测度点 … 'sim 里直行就漂'是分布的 必然,不是 reward 没调好" - the drift metric was legitimately drowned by commanded turning (v11's own comment self-documented this). The structural fix is discrete mode sampling: a custom ModeVelocityCommand that first draws a mode by share (stand/forward/back/turn/side/mixed), then draws values only on that mode's axes with all others pinned to exact zero - which is also what preserves single-variable discipline in the C ladder (native 3-axis uniform "采不出'离散模式桶' … 把 C1~C4 的单变量纪律直接毁掉"). The mixed mode later got an ellipsoid constraint rather than a cube for the same reason in reverse - corner combinations of a cube are unrepresentative extremes.

    Change

    Command generation moved from independent per-axis uniforms to mode-bucket sampling with pinned-zero off-axes (plus 20% rel_standing); acceptance likewise evaluates per mode.

    Outcome

    Straight-line behavior became a trained, testable mode instead of a measure-zero hope; the C ladder could add one mode per rung with provable isolation.

    Mechanism

    A policy optimizes expected reward under the command distribution; events of probability zero contribute nothing to the objective, so exact-zero-command behaviors (straight walk, stand, in-place turn) are only learned if the sampler gives them mass. Product-of-uniforms distributions concentrate mass on mixtures and give none to the pure behaviors users actually command.

    Applies when

    • a "simple" command (straight, stop) underperforms mixtures in sim
    • designing command distributions for velocity-tracking tasks
    • a ladder needs per-mode isolation for attribution
    “纯直行是零测度点:最终 command 为 vx∈[0.15,0.5] × vy∈U(±0.2) × wz∈U(±0.6) 连续均匀,vy=0∧wz=0 从未被专门采样 —— "sim 里直行就漂"是分布的必然,不是 reward 没调好 … Isaac 原生 UniformVelocityCommand 是三轴各自 uniform,采不出"离散模式桶"”
    train/OMNI_V0_SPEC.md § 0. 为什么从零 (3) / 三件前置 (1)