Skip to content

Training Coach

The coach reads a training run and returns a diagnosis, in-whitelist proposals and an experiment plan — and every claim cites one of the cards below. This is that corpus: 22 methodology rules distilled from a real sim2real programme, and 171 episode cards behind them, each with the sentence in the war history it came from.

Doctrine

A report may cite any of these as doctrine-N.

  1. doctrine-1Contract freeze and fingerprint discipline

    The policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.

    Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).

    Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.

    contract-freeze-and-checkerlegacy-profile-pinningderived-asset-staleness-checkgain-profile-belongs-in-the-stamprecovery-two-policies-and-a-state-machinewalk-recovery-fsm-handoff

  2. doctrine-2Attribution by resolved training params - never eval-override knobs

    Capability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.

    Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).

    Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.

    kd-bandwidth-mu-law-attributiondeploy-knob-attribution-before-retrainingcycle-time-override-is-oodresolved-config-is-source-of-truth

  3. doctrine-3PASS gates become constraints; FAIL gates become objectives

    Once a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.

    Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.

    Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".

    fixed-acceptance-matrix-per-rungpreregistered-stop-criteria-per-rung

  4. doctrine-4One variable per ladder rung - counted against what the checkpoint saw

    A rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.

    Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.

    Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.

    resume-state-dr-auditorthogonal-batch-with-ablation-orderfill-the-missing-factorial-cell

  5. doctrine-5Pre-register risks, readings, and stop criteria before the ladder

    Before a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.

    Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).

    Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.

    preregister-risks-and-fork-readingspreregistered-stop-criteria-per-rungpreregistered-real-expectationsfeasibility-accounts-lock-design-point

  6. doctrine-6Plant parameters are measured, never invented

    Every plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.

    Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).

    Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.

    friction-measured-not-guessedarmature-n2-rotor-inertiatorque-limit-shape-by-measured-peakslatency-lerp-reverse-extrapolationfeasibility-accounts-lock-design-pointsim-api-friction-columnsget-up-feasibility-accounts-before-trainingsingle-support-gain-authority-probe

  7. doctrine-7Sim2sim gate before sim2real - under deployment conditions

    Every checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).

    Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).

    Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.

    sim2sim-gate-before-sim2realeval-plant-honesty-contact-paramspipeline-latency-is-plant-not-drbody-frame-velocity-api-audittorque-penalty-bought-by-leg-bracingtorque-disagreement-between-simulators-unresolved

  8. doctrine-8Observation honesty - the actor's inputs are a hardware contract

    The actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.

    Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).

    Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.

    observation-honesty-critic-onlyhistory-obs-needs-plant-variationreward-observability-limitdeploy-heading-loop-and-align-training

  9. doctrine-9Reward economics are audited in realized currency

    Reward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.

    Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).

    Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.

    realized-contribution-auditreward-cost-of-ignoring-auditgate-new-reward-terms-by-commandignore-floor-diagnosiscalibrate-threshold-between-healthy-and-sickcalibration-threshold-with-withdrawal-clauseinert-reward-term-auditseated-basin-dead-exp-kerneltail-torque-needs-hinge-on-computed-demand

  10. doctrine-10The zero-cost option must be the desired behavior

    For every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.

    Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).

    Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.

    penalize-the-slip-not-the-jointsaturation-cheating-zero-rate-costmoving-gate-42x-stand-taxtermination-closes-degenerate-basinpenalty-gate-is-an-escape-hatchsoft-limit-penalty-charges-nominal-poseunpriced-foot-attitude-is-a-free-variableenumerate-cheapest-cheats-before-trainingbinary-band-reward-fake-touchdown

  11. doctrine-11Measurement discipline: independent referees, signs, distributions

    A disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.

    Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).

    Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.

    independent-referee-for-metric-disputesbody-frame-velocity-api-auditsame-sign-response-is-yaw-biasmedian-hides-bimodal-distributionratio-metrics-need-absolute-checkheading-integral-not-body-ratesame-distribution-reward-comparisonmultiseed-sign-test-for-driftsingle-impulse-recovery-is-chaotic

  12. doctrine-12The deployment pipeline is plant

    Irreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.

    Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).

    Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.

    pipeline-latency-is-plant-not-drpower-scale-hurts-nonforward-axesdeploy-scaling-not-training-equivalentteleop-command-band-per-axislatency-dr-covers-measured-pipelinedeploy-rate-limiter-windupslew-anchor-is-an-integratorbeta-anchored-action-targetpower-derating-cuts-full-range-contract

  13. doctrine-13DR budget is finite; its distribution is the measured support

    Robustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.

    Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).

    Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.

    push-dr-conditional-budget-conservationdr-tail-plant-continuationconstant-value-dr-overfits-margintask-shaping-before-plant-hardeningcom-randomization-forces-leg-spreadcom-dr-rollback-on-symptomthin-dr-judged-by-channel-coveragefriction-priority-re-measured-after-plant-change

  14. doctrine-14Gates measure what hardware feels: posture, margins, stripped assists

    Acceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.

    Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).

    Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.

    task-metrics-vs-posture-metricsstand-gate-posture-not-survivalconstant-value-dr-overfits-marginsuspension-probe-removes-free-stabilizerhip-roll-sum-predicts-lateral-driftchirality-scored-separatelylow-speed-commands-reward-draggingend-state-confusion-matrixfrozen-acceptance-distribution-and-pinned-seedvideo-as-acceptance-recordepisode-length-bounds-what-a-gate-seesremoved-wall-returns-on-hardware

  15. doctrine-15Fork and root selection: recoverability, maturity, frozen rewards

    Choose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.

    Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).

    Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.

    fork-root-recoverable-shortfallroot-maturity-vs-product-qualityfine-tune-reward-change-falsified

  16. doctrine-16Curricula: verified engagement, lineage counters, disease-phase gating

    Automatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.

    Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).

    Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.

    auto-curriculum-engagement-checkcurriculum-counter-lineage-stepsgate-penalties-to-the-disease-phaseaggregate-metrics-mask-subgroup-failurebucket-share-is-not-a-gradient-levercurriculum-criterion-conditioned-on-lagging-categoryper-step-income-drives-speed-time-gatetime-gate-vs-wide-stance-retire-the-fix

  17. doctrine-17Probe before training: feasibility first, hypotheses in tables

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.

    Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).

    Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).

    open-loop-probe-before-reward-tuninghypothesis-table-code-auditperiod-doubling-evidence-racesuspended-test-isolates-actuator-authorityconfiguration-probe-wall-not-slopeprone-dead-end-is-foot-placementdof-vel-penalty-is-not-a-pacing-knobamplitude-cut-falsified-yaw-fix

  18. doctrine-18External advice is recomputed locally; values transfer as ratios

    Every external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.

    Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).

    Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.

    external-advice-audit-against-own-arithmetictransfer-ratios-not-absoluteslatency-dr-covers-measured-pipelinereference-structure-fk-amplitude-divisioncycle-average-tracking-for-gait-quantitiesadvisor-paraphrase-vs-paper

  19. doctrine-19Hardware sessions are scripted experiments, not tuning sessions

    Real-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.

    Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).

    Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.

    risk-ordered-real-deploymentbattery-bracketed-real-abknow-zero-command-behaviorpush-test-chirality-protocolno-field-tuning-protocolreversible-single-variable-field-experimentssim-veto-needs-real-confirmationfirst-real-get-up-violent-stage-one-policystaged-hang-mat-floor-for-get-uppower-cycle-preflighttwo-machine-config-disciplinefall-guard-becomes-a-statehardware-log-is-the-attribution-input

  20. doctrine-20Close questions in writing; restart when the debt is structural

    Audited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.

    Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).

    Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.

    frozen-verdicts-semantic-boundariesstale-verdicts-under-old-stackzero-offset-calibration-shifts-envelopeplant-swap-invariants-vs-shiftsfreeze-lineage-fix-structure-restartminimal-reward-table-with-provenancewrite-hardware-verdicts-back

  21. doctrine-21Name the quantity in the space it lives in

    A goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.

    Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).

    Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.

    joint-space-proxy-for-task-space-quantitycontact-detector-single-signal-lieszero-partial-credit-is-not-an-iteration-problemheading-integral-not-body-rate

  22. doctrine-22Continuation needs a live gradient; a release is chosen by a scan

    Continue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.

    Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).

    Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.

    converged-continuation-is-poisoncheckpoint-choice-is-a-full-gate-scanstance-decided-by-get-up-pathstop-stacking-roll-back-and-auditcurriculum-history-is-part-of-the-productcontinuation-budget-not-from-zeroroot-maturity-vs-product-quality

Experience cards

101 cards matching “task-metrics-vs-posture-metrics”.

  • The restart reward table lists a reason for every term AND a lesson for every exclusion - absent terms are removed, not zero-weightedminimal-reward-table-with-provenance
    Mechanism understoodomnireward-shapingreward-shapingprocesscontract-freeze

    Maintain the reward table as an evidence ledger: every term cites the episode that justifies it, every excluded term cites the episode that convicted it (including the development stage it is valid at), and retired terms are deleted from the config, never left at weight zero.

    Symptom

    Seven walk generations had accumulated an entangled reward table where nobody could say which term earned its place; the restart needed a table that could be audited line by line.

    Context

    The minimal table v2 was built under three written principles: "一项管 一件事、结构性反抬脚的项一个不留、塑形只留一套相位逻辑;不在表里的一律 不加" (one term per job; zero structurally-anti-lift terms; exactly one phase-shaping logic; nothing outside the table gets added). Every row carries its provenance (e.g. base_height target = standing height cites the v4 crouch lesson; split x/y tracking cites the merged-exp gradient hole; world-frame yaw cites the v1/v3 body-frame lesson). Every EXCLUSION carries its same-type precedent: feet_landing_vel out because it is poison while the gait is unformed (drag pays 0, lifting pays - the v4-clearance / v8a-B "reverse threshold" family) though it was fine in v7 when the gait already existed - term validity depends on developmental stage; feet_air_time out with its measured non-lever evidence (v6 had it at 2.0 and still lifted 4 mm); and replaced terms are REMOVED from the config ("置 None,不是权重 0 挂着") so audits see truth, not dormant weight.

    Change

    Reward table rebuilt as ~20 rows each with weight + provenance column; exclusion list maintained alongside with the falsifying episode for each; dormant terms deleted rather than zeroed.

    Outcome

    Later revisions (S1.2/S1.3) modified the table by citing and updating specific rows' evidence rather than re-arguing the whole design; the table doubled as the lineage's reward-lesson index.

    Mechanism

    A reward table is a set of standing hypotheses; attaching each row's evidence makes revisions targeted and reversible, and recording why a term is absent prevents the cycle of re-adding known poisons. Deleting vs zero-weighting matters because config audits and DR interactions see the term either way - a zero-weight term is dormant complexity waiting to be flipped on wrongly.

    Applies when

    • designing a reward table for a restart or new task
    • someone proposes re-adding a previously removed term
    • auditing which reward rows still earn their place
    “原则:一项管一件事、结构性反抬脚的项一个不留、塑形只留一套相位逻辑;不在表里的一律不加 … feet_landing_vel(评审 #4):拖地时代价恒 0、抬脚才收费——与 v4-clearance/v8a-B 同属「反抬脚门槛」家族,步态未成形时是毒;v7④ 加它时步态已存在。… 已从 cfg 移除(置 None),不是权重 0 挂着。”
    train/OMNI_V0_SPEC.md § 3. 最小奖励表 v2 / 明确不带
  • mj_objectVelocity returns inertial-principal-axis frame - one API assumption poisoned eval and observations for a whole linebody-frame-velocity-api-audit
    Mechanism understoodomnisim2sim-gatemeasurementsim2simobservation-honestyattribution

    Verify every frame-sensitive API against a hand-computed truth (rotate raw qvel yourself, or command a known world velocity and check where it lands) before trusting any evaluation or observation built on it - especially when a model's inertial frame is rotated from its body frame.

    Symptom

    Sidewalk vy read ~0 under every condition; separately, whole-policy performance was mysteriously mediocre in sim2sim while training-side numbers looked fine. Four training rungs were declared FAIL partly on these readings.

    Context

    base_link's URDF inertial frame is rotated 90 deg about x relative to the body frame (iquat = [0.7071, 0.7071, 0, 0]). mj_objectVelocity(flg_local=1) rotates into ximat - the inertial principal-axis frame - not the body frame, and returns center-of-mass point velocity, not body-origin velocity. Consequences measured: the "vy" column was actually vertical velocity vz (walking at cmd 0.25: old reading +0.0093 vs true -0.0424); the angular velocity fed to the policy in sim2sim was [wx, wz, -wy] - a different quantity than Isaac and the real IMU provide. RMS check over 8 s of walking: y/z axes swapped between v6[:3] and the qvel truth.

    Change

    Fixed sim2sim and both probes to compute ang_b = qvel[3:6] and lin_b = xmat.T @ qvel[0:3] (identical quantity to Isaac's root_ang_vel_b / root_lin_vel_b), with a standalone reproduction script (frame_bug_repro_0809.py).

    Outcome

    Re-scoring the "failed" C4 lineage under correct coordinates reversed the verdicts: c4r4 checkpoints showed vy 80-126% tracking (old reading: +/-2%) and vx+0.30 at 91-95% where the old metric said 0/5 - the bad frame both mis-measured vy and, via corrupted policy observations, systematically depressed all measured performance. Final product passed 260/260 cells.

    Mechanism

    A simulator API's frame convention is part of the observation contract; when the model's inertial frame is rotated relative to the body frame, frame-agnostic use of a "local" velocity silently permutes axes. Feeding a policy an axis-permuted angular velocity is an observation corruption that degrades behavior everywhere, not just on the axis being studied.

    Applies when

    • building or auditing a cross-simulator evaluation harness
    • one measured axis reads near-zero under all conditions
    • sim2sim scores are inexplicably worse than training-side metrics
    • URDF/MJCF inertial frames are rotated relative to body frames
    “base_link 的 iquat = [0.7071, 0.7071, 0, 0] … mj_objectVelocity 用的是这个 … 喂给策略的 base_ang_vel 是 [wx, wz, −wy] —— MuJoCo 侧观测与 Isaac / 真机 IMU 不是同一个量;vy_mean 报的是竖直速度 vz —— 前进 cmd 0.25 时旧读数 +0.0093,真值 −0.0424。”
    train/C_LADDER_RUN.md § 3l. ⚠️ mj_objectVelocity 读的是惯性主轴系 / 3m. 一 bug 坐实
  • A joint-velocity penalty meant to slow the get-up cut joint speed 16% and left the get-up time unchanged - the knob never moved the variable, so the idea it was meant to test stayed untesteddof-vel-penalty-is-not-a-pacing-knob
    Mechanism understoodrecoveryreward-shapingreward-shapingattribution

    Before reading a result as a test of an idea, check that the knob actually moved the independent variable; velocity regularizers smooth a schedule they do not set, and a schedule driven by per-step task income moves only when that income's time structure does.

    Symptom

    The get-up took 0.6-0.9 s in Isaac with large torque demand; the user proposed getting up more slowly so less torque would be needed.

    Context

    The idea had support in the accounts: the acceptance bound is an upper bound of 5 s (5-8x margin), the quasi-static squat path peaks at 25% of the limits, and rolling over needs no momentum. R3.2 raised dof_vel from -1e-3 to -5e-3 as the single variable.

    Change

    dof_vel -1e-3 -> -5e-3 (child-run from R3.1).

    Outcome

    Get-up medians moved +0.02-0.04 s (noise); raw joint velocity -16%; torque demand median got worse (hip_pitch 46-48% -> 63-67%) as the new term competed with torque_headroom on the same joints; MuJoCo 98 -> 96%. Verdict FAIL on the knob, not on the idea, and the rung was not adopted. When pace was later attacked through the income's time structure (V2.5/V2.5b), the MuJoCo get-up moved into the 3.5-4.5 s design band.

    Mechanism

    The pace was set by base_height_progress paying for every step spent high (stand earlier, earn more); a velocity regularizer only smooths motion along the same schedule and does not change when the robot stands up.

    Applies when

    • trying to make a skill slower or gentler with smoothness penalties
    • an experiment's primary metric did not move and a verdict is being written
    • two penalties act on the same joints
    “**关键判读:`dof_vel` 罚只把关节速度压了 16%,而起身用时一点没变。** 也就是说**这一级根本没有把"慢下来"这个自变量推动起来** —— 所以它**不构成对 用户假说的检验** … 起身节奏由 `base_height_progress` 的逐步计酬决定(早站起来就多 拿),速度正则只在同一条时间轨迹上把动作抹匀,不改变何时站起来。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §25 R3.2(dof_vel −1e-3→−5e-3,慢一点起身)
  • Audit which joints your imitation term constrains - a task that needs deviation is fighting the referenceimitation-term-scope-audit
    Mechanism understoodomnireward-shapingreward-shapingcurriculum

    List which joints your imitation/deviation terms actually constrain and check the new skill's required motion against that list; for balance-coupled joints deliver references as feed-forward residuals, not absolute-position targets - and never assume "reference = 0" is neutral.

    Symptom

    Sidewalk would not learn despite a dedicated tracking reward; meanwhile the gait-shaping imitation term (joint_pos_ref) computed its error norm over ALL 12 joints while its reference covered only the 6 sagittal joints - roll/yaw reference was constantly 0.

    Context

    Two prior generations had shown the forward gait itself was taught by joint_pos_ref, not discovered by PPO (v6 halved the shaping and swing height collapsed 35 mm -> 4 mm). So the reference is load-bearing - but sidewalk requires hip_roll to deviate from nominal, and the all-joints norm punished exactly that deviation: "一边悬赏一边罚过程" (posting a bounty while punishing the process). A follow-up experiment (C4-redo3, free_roll=True releasing the 4 roll joints from the norm) raised the regularization headroom 6x -> 27x yet sidewalk stayed flat and released hip_roll wandered, killing other skills - net negative, withdrawn. A --roll-absolute probe showed the converse failure: pinning roll to a clock-driven absolute trajectory drove tilt 6.9 -> 13.7 deg. Conclusion recorded: absolute-position imitation cannot teach actions that must be superimposed on state feedback.

    Change

    The audit reframed the problem: neither punishing roll deviation nor freeing roll nor absolute roll tracking works; the reference for a balance-coupled joint must be delivered as feed-forward under the policy's residual control (see feedforward-for-phase-locked-skills).

    Outcome

    free_roll rung: joint_pos_ref term rose 0.887 -> 1.104 (release confirmed effective) but vy stayed flat; regularization hypothesis eliminated by experiment.

    Mechanism

    An imitation error norm defines a cage: joints inside it are pulled to the reference in absolute position, so any skill requiring systematic deviation is taxed per step; but joints carrying active balance cannot follow absolute references either, since their correct position depends on state. The scope and the delivery mechanism of the reference are therefore design decisions per joint, not defaults.

    Applies when

    • adding a skill that moves joints your reference sets to zero/nominal
    • an imitation or deviation penalty coexists with a new tracking reward
    • considering releasing joints from a shaping term mid-lineage
    “前进步态也不是 PPO 自己发现的,是 joint_pos_ref 教出来的(v6 砍半塑形 → 抬脚 35 mm 塌到 4 mm…)。而 ref_joint_offset 原本只写 6 个矢状面关节,roll/yaw 参考恒 0 —— 侧走既没被教,roll 一偏离 nominal 反被 joint_pos_ref 扣分。一边悬赏一边罚过程。”
    train/C_LADDER_RUN.md § 3e. 为什么首战 FAIL / 3i. 解锁笼子
  • Drop the frozen policy into chosen configurations - a squat 2.7 cm lower than the stuck pose stood 52% of the time, the stuck W-sit 0%, and the interpolation between them showed a wall, not a slopeconfiguration-probe-wall-not-slope
    Mechanism understoodrecoveryattributionattributionmeasurementcurriculum

    When a policy is stuck, probe the frozen policy from a grid of hand-placed start configurations, including interpolations between the stuck state and a nearby state it escapes from; one read-only experiment separates height, torque, sampling and configuration and tells you whether to prevent entry or train the exit.

    Symptom

    After R0.3 the policy stood from 62% of starts and never from the W-sit it fell into; height, torque, missing samples and reward were all plausible suspects.

    Context

    A read-only probe placed the R0.3 policy directly into specified configurations. Squats (hip, knee, ankle) = (-.65,-1.3,-.65) stood 100%, (-1.0,-2.0,-1.0) 89.8%, (-1.2,-2.4,-1.2) at 0.176 m 52.3%; the measured W-sit at 0.203 m 0.0%; the account-(3) hand-over state (146 deg tilt) 34.4%; linear interpolations from the W-sit toward the squat at 25/50/75% stood 0.0/0.0/3.1%. The squat family's quasi-static torque is 16% of the limits, and the W-sit was visited ~9 s per episode in training. FK showed the squat family (-a,-2a,-a) keeps the torso vertical, the feet flat and the COM over the feet all the way from 0.146 m to 0.384 m.

    Change

    Height, torque and sampling were eliminated in one experiment; the next rungs targeted entering the W-sit (foot placement) instead of escaping it, and seeding the dead point itself was ruled out because it was already visited every episode.

    Outcome

    Pure configuration: the W-sit (hips externally rotated +/-47 deg, knees folded 110 deg, shins flat, feet beside the body) is a different place from the sagittal squat (feet flat under the COM). The policy's standing skill was bound to a narrow sagittal family, and the wall was confirmed by the interpolation. The foot-placement rungs that followed took prone from 0/159 to 158/159.

    Mechanism

    A learned skill covers the neighbourhood of the states it succeeded from; a start state outside that neighbourhood fails regardless of height or torque, and an interpolation that stays at zero until close to a working state shows the boundary is sharp.

    Applies when

    • a policy stalls in a specific posture and several causes are plausible
    • deciding between reverse-curriculum seeding and entry-prevention shaping
    • a feasibility account says a path exists but the policy does not take it
    “**决定性对比:比死点矮 2.7 cm 的蹲姿站立 52.3%,死点 0.0%。** 所以不是高度、 不是力矩(蹲姿族准静态力矩膝 1.96/12、踝 1.24/17,只占 16%)、也不是训练采样 (死点每局被访问 ~9 s)。**是纯位形问题** … 插值实验进一步显示这**不是坡是墙** —— 走到 75% 仍只有 3.1%”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §15 死点位形实验(只读探针,同一个 R0.3 策略放进指定位形)
  • Prone get-up sat at 0/159 until two gated hinge terms moved the seated feet - first sideways (561 -> 360 mm), then fore-aft (-168 -> +56 mm) - and success went to 158/159 with nothing else changedprone-dead-end-is-foot-placement
    Mechanism understoodrecoveryreward-shapingreward-shapingattributioncurriculum

    When a stuck state and a successful state differ geometrically, penalize the discriminating quantity with a gated hinge that is exactly zero in the state the policy actually reaches (measure it - not the nominal), then re-probe: flattening one axis can move the discriminant to another.

    Symptom

    Prone falls always righted and then sat with the feet splayed wide or tucked behind the hips, from where the policy never stood (0% for four generations).

    Context

    Four lines of evidence pointed at foot position: the configuration probe (ankles 215 mm apart stood 52.3%, 561 mm apart 0.0%); FK showing the action contract's nominal (a = 0) is itself a 465 mm straddle, so the action_rate and still terms were pulling toward the splits; biomechanics (feet tucked under the body cut peak hip-extension torque 148.8 -> 32.7 N*m, -78%); and HoST's foot-displacement term, which this reward table lacked. The earlier "not a reward hole" reading was corrected to "a gradient hole, not a level hole": at the dead point the heaviest term (upright) was saturated with zero gradient, still paid for not moving, and the one live gradient (base_height) pointed at the thigh-horizontal torque barrier. A prone ROM scan had already ruled out pushing up from prone.

    Change

    R0.4: feet_spread_excess = clamp(ankle distance - 0.215, 0, inf) x upright gate, weight -2.0, plus a height-decay factor added after measuring that the policy's real standing stance was 406 mm, not the 215 mm nominal (the plain version would have taxed every successful stand 0.38/s). R0.5: the same shape on the fore-aft axis, feet_fore_seated = |fore-aft offset - 0.05| x upright gate x height decay, target +50 mm (the measured natural offset of standing postures). One variable per rung.

    Outcome

    R0.4: seated ankle distance 561 -> 360 mm, supine/side exactly unchanged, prone 0 -> 1.9%, mid 45.9 -> 62.2%; a probe then showed the discriminant had moved to the fore-aft axis (standing starts +42 to +51 mm, the prone seat -168 mm). R0.5: supine 99.4, prone 99.4, side 100, mid 100%, re-falls 0%; both geometry terms collapsed to ~0 near iteration 13,100 as base_height rose, and the prone fore-aft offset went -168 -> +56 mm - the term's own target, closing the causal chain. The cost, unmeasured at the time: action jitter rose 33% (sum |da|^2 6.82 -> 9.06).

    Mechanism

    An upright-gated hinge is inert while the robot rolls and exactly zero in the achieved stance, so it adds gradient only inside the stuck basin; a seated robot with its feet behind or outside its COM must make a kinematically unfavourable transition to stand, and moving the feet under the body removes it.

    Applies when

    • a get-up or transition skill fails from one start category only
    • successful and failed episodes differ in a measurable geometric quantity
    • a shaping term might tax the posture successful episodes already use
    “`recovery_r0_5`,唯一变量 = 追加 `feet_fore_seated`(与 R0.4 同形状,只换测量轴)。 … 对照 R0.4 的 prone(3.1%,360 mm,**−168 mm**):前后偏移从 −168 走到 +56, 正是这一项的目标量,**判别量被消掉后成功率随之到顶** —— 因果链完整。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §20 R0.5(前后向脚位置项):成功率门全过 —— prone 0/159 → 158/159
  • Real-robot trials of a new skill were staged by risk - a hanging dry run with the robot posed by hand, then one short try per category on a mat with the hardest last, then the composed behaviour (switch + walking) last - with the user present and a log every timestaged-hang-mat-floor-for-get-up
    Replicatedrecoveryreal-deployreal-acceptanceprocess

    Stage a new skill's hardware trials by risk - hanging dry run posed by hand, one short try on a mat per start category with the hardest last, the composed behaviour last - with an operator ready to cut enable and a log for every try; relax a safety ban only for short, attended runs and say so in writing.

    Symptom

    A get-up policy acts violently near the ground by design, and the first unstaged real run of the line was stopped as dangerous.

    Context

    The hanging checklist written with the first stamped recovery product (v2_5, 2026-08-11), to be ticked item by item with the user present: both machines on the same commit and firmware torque limits checked; the robot hung from a single point about 0.1 m off the ground; a dry run with the robot posed by hand into supine and prone to watch that the target stream is gentle (the beta contract keeps targets within +/-0.25 rad of the measured pose, so enabling causes no homing fling); the first floor try is supine only, on a mat, once, with torque and joint logs; categories are added one at a time, prone last; any kicking or oscillation cuts enable immediately. For the switch (08-14) the runbook orders: hang with the standing policy as the locomotion side, then on a mat push the robot over and let it recover, and only last swap in the walking policy. An exemption was also written: edge-standing policies stay banned from long or unattended runs, but a short single A/B with the user present, hung or on a mat, is allowed. The one-leg line reused the same order (hang, then floor with a spotter, 60 s segments with a temperature check).

    Change

    Real trials as a checklist of stages, each gated on the previous one, with the composed behaviour last.

    Outcome

    The line's first real get-up (v2_6, 08-11) came through this protocol and was reported "fairly stable"; no further hardware outcomes of the switch are recorded in the spec.

    Mechanism

    Each stage exposes one new risk (commanded targets without contact, a single category with contact, harder categories, then the interaction of two policies), so a failure is attributable and cheap.

    Applies when

    • first hardware trial of a recovery, jumping or other high-impact skill
    • switching between two policies on hardware for the first time
    • a policy with a known posture defect needs a comparison run
    “吊挂空跑: 手动摆到 supine/prone 姿态, 看目标流是否温和 (β 帽 7.5 N·m, 目标永远贴着当前 q ±0.25 rad —— 使能瞬间无归位甩动, 这是 β 契约附带保证) … 落地首试: supine 一类, 垫子, 单次; τ/q --log 全程记录 … 逐类别扩展 (prone 最后), 每类先单次”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §40 吊挂执行单(真机首试;需用户在场,逐项打勾)
  • Cutting swing amplitude 40% raised yaw-momentum demand 53% - the falsified fix is recorded so nobody walks that road againamplitude-cut-falsified-yaw-fix
    Mechanism understoodwalksim-evalmeasurementattributionreward-shaping

    Test gait fixes against the quantity the ground must actually supply (torque/force rates vs friction ceilings), not against kinematic proxies; record falsified fixes with their mechanism so the search space shrinks permanently.

    Symptom

    Support-foot yaw slip stayed at ~223 deg (vs 225 deg) after walk_v6 cut joint swing amplitudes by ~40% (hip_pitch 20.6->10.4 deg, knee 31.3->18.8 deg) - the change built on the theory "smaller swing = less yaw momentum to dump into the ground".

    Context

    Direct measurement inverted the theory: yaw-momentum amplitude ROSE 16% (+/-0.1000 -> +/-0.1159) and its rate of change rose 53% (1.86 -> 2.84 N*m demanded from the ground), pinned exactly at the foot's supply ceiling (2.8-3.3 N*m at mu 0.6-0.7) - so slip could not drop. The extra demand lives in higher harmonics: v6's crisper foot placement (better clearance 34 mm, lower landing force) shortens the momentum- exchange window, concentrating the same exchange into less time. The verdict was written as a closed road: "下一轮不要再走这个方向". An honest residue was also booked: WHY amplitude down but momentum up 16% remained unresolved, with the named next step (per-rigid-body decomposition of H_z, since per-joint RMS sensitivity ignores phase correlations).

    Change

    The "reduce amplitude to reduce yaw momentum" lever was removed from the planning space; future yaw-slip work redirected toward the supply side (friction) and momentum-rate mechanics.

    Outcome

    Slip unchanged (225 -> 223 deg); the falsification and its mechanism became a permanent constraint on the fix search space.

    Mechanism

    Ground yaw torque demand scales with the rate of change of angular momentum, not its amplitude; kinematic amplitude cuts that also sharpen contact timing can raise dH/dt while lowering range. When demand exceeds the friction-limited supply ceiling, slip is set by the ceiling, so demand-side changes below the ceiling do nothing visible.

    Applies when

    • attacking foot slip or yaw drift via gait shape changes
    • a fix targets an amplitude while the constraint is a rate
    • documenting a failed intervention after a version comparison
    “walk_v6 | ±0.1159 (+16%) | 2.50 Hz | 2.84 N·m (+53%) … 脚的供给上限 2.8~3.3 N·m(μ 0.6~0.7)—— v6 正好顶在天花板上, 所以滑移一点没降。… 结论: "减小摆动幅度以降低偏航动量"这条被 v6 证伪 —— 砍 39% 幅度, 需求反升 53%。下一轮不要再走这个方向。”
    train/WALK_DIAGNOSIS.md § ② 未生效的机理: 减小摆动幅度反而让偏航需求上升
  • A field experiment is allowed when it is pre-scripted, single-variable, and self-reversing (RAM-only writes)reversible-single-variable-field-experiments
    Observed oncewalkreal-deployhardwarereal-acceptanceprocessattribution

    Permit hardware-side experiments only when scripted in advance with one variable, a log, and automatic reversion (volatile writes, git restore); keep every config mirror (yaml vs firmware) changed and restored as a unit.

    Symptom

    A hypothesis needed a hardware test - v5's "wild kicking" might trace to the RS00 torque cap being deployed at 11 N*m vs its trained 14 (-21%), on exactly the ankle-roll/hip-yaw joints doing lateral-yaw work - but changing limits in the field is the classic way to lose track of robot state.

    Context

    The experiment was written to be safe by construction: change exactly one number (robot.yaml RS00 tau_limit 11.0 -> 14.0), write to firmware RAM only (--write without --save, so a power cycle automatically rolls back to the saved 12/17/11), run the single logged trial, then restore both yaml (git checkout) and RAM immediately. The consistency requirement is explicit: the deploy tool's torque self-check compares against robot.yaml, so yaml and firmware must change and restore together; and the hypothesis scoping is itself single-variable - RS06's 67% cut was measured irrelevant (gait uses only 16% of rated) and hip torque was left alone for safety.

    Change

    Field experimentation policy refined: not "never touch hardware settings" but "only pre-scripted, one-variable, logged, auto-reverting changes with config/firmware kept consistent".

    Outcome

    The sub-experiment could answer the torque-cap hypothesis without any risk of the robot persisting in an undocumented state - forgetting to restore costs nothing but a git checkout.

    Mechanism

    The danger of field changes is state divergence (robot config drifting from the repo's record), not the change itself; volatile (RAM-only) writes bound the divergence lifetime to one power cycle, and single-variable scoping preserves attributability even in a field setting.

    Applies when

    • a hypothesis requires changing firmware limits or gains on the robot
    • field debugging tempts persistent config writes
    • designing safe escape hatches for deployment tooling
    “程序(--write 不带 --save = 只写 RAM,断电自动回滚)… deploy 的限扭自检是对 robot.yaml 比对的,所以 yaml 和固件必须同改同还原;忘了还原也没事,断电重启即回 12/17/11(上次 --save 的值),但 yaml 要 git checkout。”
    train/REAL_SWEEP_V5_V8.md § 4. 限扭子实验(可选二期,只对 v5,单变量)
  • Continuing a converged policy on a change that carried no new gradient drifted its transfer from 100/98% to 80/28% over 3,000 iterations while every Isaac gate stayed perfect - scan every checkpoint on the second simulator's friction axisconverged-continuation-is-poison
    Observed oncerecoverytraining-runfork-selectionsim2simcurriculum

    Before continuing a converged policy, check that the change creates a live gradient; if it does not, cap the budget at a few hundred iterations, and in every continuation scan each checkpoint on the second simulator's transfer axis (for example low friction) - trainer-side gates can stay perfect while transfer decays.

    Symptom

    V2.7-A (swap the flat_feet term for a compensated version, continue from v2_6c) finished with the line's best Isaac score (100%) and a MuJoCo transfer collapse: mu 1.0 98 -> 80%, mu 0.4 98 -> 28%; the stance it was meant to widen had not moved.

    Context

    The new term's calibration run showed a near-zero tax from the start: the policy already satisfied it, so the reward landscape offered nothing new. A checkpoint scan on MuJoCo mu {1.0, 0.4} located the damage: +100 iterations 100/98% (better than the baseline), then 86/54, 60/38, 80/28 - monotonic decay with training length, while entropy and action noise rose (7.77 -> 8.18, 0.588 -> 0.612): drift, not sharpening.

    Change

    Rule written in: with no new gradient, a continuation budget is short (at most a few hundred iterations) and the MuJoCo transfer axis enters every checkpoint scan. The next rung (V2.7b, a live stance-width gradient) was budgeted at 1,000 iterations with mu {1.0, 0.4} scans every 100 and a stop-on-signal rule.

    Outcome

    V2.7b kept transfer at the same depth (mu 1.0 98% / mu 0.4 92% at +1,000, where A had already rotted to 86/54) and at +3,000 (100/96%): a live gradient preserved transfer. V2.8 then broke that pattern (mu 0.4 2%): the gradient must also be compatible with the policy's existing form.

    Mechanism

    On a converged reward landscape PPO keeps updating without a signal to follow, and the random walk is pulled toward whatever the training plant rewards idiosyncratically - invisible in the trainer's own gates.

    Conflicts

    The drift mechanism is the spec's reading of one decay series plus one contrasting run; V2.8 is recorded as an exception to "live gradient keeps transfer".

    Applies when

    • fine-tuning a converged policy with a small reward change
    • a continuation run's trainer-side metrics improve while real or cross-sim results worsen
    • choosing which checkpoint of a continuation to ship
    “**checkpoint 扫定死因**(μ1.0/μ0.4):**29500(+100 iter)= 100/98%** (优于基线!)→ 30400 = 86/54 → 31400 = 60/38 → 32398 = 80/28 —— **迁移随续训长度单调衰减**。 … **教训入库:收敛均衡上的长续训是毒药 —— 无新梯度时 续训预算须短(≲数百 iter),且 MuJoCo 迁移轴必须进 checkpoint 扫描。**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §43 结果:V2.7-A 判 FAIL —— 换刀本身无罪,毒在续训预算
  • The walk phase machine structurally cannot express flight - rebuild the representation for duty < 0.5, teach flight with a mask tax, never a cliff bountyphase-machine-structural-limits
    Mechanism understoodrunreward-shapingreward-shapingcurriculumgate-battery

    When a new gait changes the contact pattern's structure, audit whether the phase/mask representation can express it and rebuild the representation if not; teach the new contact pattern with graded mask-mismatch pressure and count it in acceptance with artifact-proof definitions (minimum segment length), never with cliff bounties.

    Symptom

    Running requires both feet airborne, but the walk-era phase machine switches legs by the sign of sin(phase) - with duty < 0.5 the two swing windows must OVERLAP during flight, which a sign-switching representation cannot express at all.

    Context

    The run phase machine was re-architected rather than patched: per-leg phases (left = phi, right = phi+0.5 mod 1) with leg_phase < duty defining stance, aligned to the walk sin convention at duty=0.5 so the machines agree where their domains overlap. Flight is taught by the SAME mechanism that once cured foot-dragging, direction reversed: in the two planned flight windows the contact mask is (0,0) and feet_contact_number_duty charges -0.3 per foot still on the ground - a mild ~0.16/step tax, deliberately NOT a cliff: "悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文". The reference shape (half-sine bump over swing progress) is zero at window boundaries by construction, eliminating the clearing-window that the C4 probe measured to cost 13-19% on non-sinusoidal references. The shape self-check ("零代价选项是什么") was run on three behaviors: standing pays ref everywhere (known cmd=0 stepping risk, booked), walking pays only the flight-window tax, proper running collects full marks.

    Change

    New rewards.py run section (leg_phase_duty / stance_mask_duty / ref_run / clearance_run / contact_duty) with the walk versions untouched byte-for-byte; flight acceptance metric defined with a segment-length floor (>=40 ms to count) so numeric contact flicker cannot fake flight.

    Outcome

    Flight became expressible and taught by a calibrated mild pressure; the walk lineage's phase code stayed frozen as its own contract.

    Mechanism

    A phase representation defines which contact patterns exist in the reward's vocabulary; duty cycling below 0.5 introduces states (double-flight) outside a half-period sign convention's language, so no weight tuning can teach them. And rare desirable events taught by cliff-shaped bounties invite hacks (jumping in place); the graded mask tax prices the planned pattern without creating a jackpot.

    Applies when

    • extending a walking stack to running/jumping (duty < 0.5)
    • a desired contact pattern never appears despite reward increases
    • defining flight/contact acceptance metrics
    “duty<0.5 时摆动窗 (1−duty)T > T/2,两腿摆动窗在腾空段重叠 —— sin 符号切腿的机制结构上表达不了"双脚同时在空中"。… feet_contact_number_duty 对"还踩着地"持续 −0.3/脚 —— 与 walk 治拖地同一机制,方向相反。不腾空的税 ~0.16/步 … 梯度温和不构成悬崖(悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文)。”
    train/RUN_V0_SPEC.md § 5. 相位机设计
  • A plateau in a training curve was a population mix, not a half-learned skill - 60% standing at 0.372 m and 40% sitting at 0.19 m - and a category at hard zero stayed at zero through 3,000 more iterationszero-partial-credit-is-not-an-iteration-problem
    Mechanism understoodrecoveryattributionmeasurementattributioncurriculum

    Before buying iterations for a plateau, split the metric by category and check whether it is bimodal; a category at hard zero with no partial credit is missing a capability or a reachable state, and more iterations under an unchanged config will only polish the categories that already work.

    Symptom

    After R0.1 the training-side base_height sat near 0.29 m and the curve was still climbing at the iteration cap, which read as "train it longer".

    Context

    Candidate A (user decision) was a child-run from R0.1's last checkpoint with zero config change - the logged env.yaml files differ only in log_dir - for 3,000 more iterations (R0.2). The per-category acceptance split was already available: prone had scored 0/156 with no partial credit.

    Change

    Continue training unchanged, then read the result by category rather than by the pooled curve.

    Outcome

    supine 91.5 -> 96.4%, side 82.2 -> 87.9%, get-up 0.96 -> 0.84 s, pose error 0.91 -> 0.54 - all improvements to categories that already stood. Prone stayed 0/153; mid 55.9 -> 41.2% was within noise (n=34). Height by category was binary - standing groups 0.372/0.373 m, seated groups 0.187/0.194 m, nothing between - so the pooled 0.29 was 0.61 x 0.372 + 0.39 x 0.19 = 0.30 (measured 0.307): six in ten standing, four in ten sitting. The rendered prone episode was still kneel-sitting at t = 8 s.

    Mechanism

    The pooled mean of a binary outcome only moves when the mix moves; PPO kept polishing the subpopulation that already succeeded while the failing one produced no advantage signal to follow.

    Conflicts

    R0.2 recorded the missing capability as "prone lacks rolling over"; R0.3's end-state confusion matrix retracted that - prone had righted its torso in 159/159 episodes and was failing to stand from the W-sit. The lesson that iterations could not fix it holds; the named cause was wrong.

    Applies when

    • a training curve plateaus while acceptance shows one category at zero
    • deciding between "train longer" and "change something"
    • pooled training metrics are read as the typical episode
    “**分类别 h 中位把"平台 = 人口混合"钉死了**:数值是**二值**的 —— 站立组 0.372/0.373,坐姿组 0.187/0.194,**中间没有过渡态**。 … **这也是本仓此后读该指标的通用告诫:全体混合的期望会把 双峰分布平均成一个不存在的中间值,必须分类别看。** … **结论:A 不能过门,原因确定为 prone 缺"翻身"这一技能,不是迭代不够。**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §13 R0.2(recovery_r0_2,child-run 续训):A 走完了 —— 推不动 prone
  • A pull-assist curriculum keyed to a global success share was satisfied by the categories that already worked and withdrew before prone learned anything - conditioning the criterion on prone took it from 2.5% to 98.7%curriculum-criterion-conditioned-on-lagging-category
    Mechanism understoodrecoverytraining-runcurriculumgate-battery

    Measure a curriculum's advancement criterion on the population the scaffold is meant to help; a global success share is met by whatever already works, and the help is withdrawn before the lagging case learns.

    Symptom

    Under the beta-anchored action, supine and side stood reliably while prone still sat (1.3%). A pull-assist curriculum added to help it was withdrawn completely within ~790 iterations and prone moved only to 2.5% (noise).

    Context

    The pull assist follows HoST: an upward force on the base, active only when the torso is within 30 deg of vertical, scaled by body weight (HoST's 200 N on G1 = 0.583 BW -> 56 N here, steps of 5.6 N, ten levels to zero); the product must pass with no assist. It had already taught sit-to-stand in V1. In V2.1 its advancement criterion was the standing-time share over all envs (threshold raised to 0.55 because the share was already ~0.53).

    Change

    V2.2: PullAssistForce with gate_category = "prone" - only envs whose first step classifies them as prone count toward the criterion - and the threshold back at 0.35. A feasibility signal was pre-registered: if prone's share stayed near zero under the full 56 N, return to the roll-over path instead of adding force.

    Outcome

    The prone-conditioned curriculum withdrew level by level only as prone itself passed: prone 98.7%, and the four-category gate passed for the first time on the line (98.6% overall, re-falls 0%, torque gate PASS).

    Mechanism

    A pooled success share is filled by the categories that already succeed (supine/side ~53%), so the scaffold is removed on their account before the lagging category has used it.

    Applies when

    • an assist, guide force or easier setting is withdrawn by a success threshold
    • one task category lags while the pooled metric looks healthy
    • a curriculum ran to completion without changing the lagging category
    “**教训:全局站立占比阈会被存量类别(supine/side ~53%)凑够,拉力在 prone 学会前就撤光了 —— metric 设计失误,不是拉力机制失效**(它在 V1 教会过 坐→站)。 … **V2.2(已启动)**:`PullAssistForce` 加 `gate_category="prone"` —— 达标判据 只统计 prone 类 env”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §34 V2.1 判决(2026-08-10)
  • The fallen-state reset was designed, not sampled from SO(3) - fixed category shares with jitter, a low drop that settles physically, equal left/right shares for mirror augmentation, and a numeric check before trainingfallen-pose-reset-distribution
    Observed oncerecoverytraining-runcurriculumdomain-randomizationprocess

    Build a fallen-start distribution from named, physically plausible categories with jitter and a settle phase, keep mirrored categories at equal probability, and check the realized shares and penetration numerically before spending a training run on it.

    Symptom

    A get-up policy can only learn from the fallen states its resets produce; uniformly random orientations produce ground-penetrating and limit-jammed states the robot can never be in.

    Context

    R0 reset_root_fallen: supine 30%, prone 30%, side_l 15%, side_r 15%, mid (random axis 50-125 deg) 10%, +/-15 deg jitter, full yaw, dropped from 0.28-0.40 m and left to settle under physics, joints uniform inside the soft limits with a 5% margin plus small random velocities. Random SO(3) was rejected (the advisor agreed). side_l and side_r must have equal probability because mirror augmentation turns a left fall into a right fall. The advisor had proposed supine and prone only for R0; the spec included side and mid because the feasibility accounts showed physical solutions for all of them, and wrote "narrow back to supine+prone" down as the first fallback. With no display on the training box the reset was checked numerically instead of by eye.

    Change

    Category mix as above; realized shares, settle height and penetration measured over 512 envs before the first run. A fallen-state bank (real falls, settled and stored) was pre-registered for R2.

    Outcome

    Realized shares 29.3/31.6/16.4/17.8% against the config, settle +0.262 m, final penetration 0/512 (a 0.10 m peak at the write instant, ankle links only, pushed out within 80 ms because the 0.28 m drop floor is shorter than a fully extended leg). R0's failure was a reward basin, not a reset artifact. The fallen-state bank stayed unbuilt through V3.1 (checklist item open); R0.3 later re-sliced the prone share into roll_l/roll_r bands, which is what forced the acceptance distribution to be frozen separately.

    Mechanism

    A category-structured, physically settled start distribution keeps training on states the robot can actually occupy, and equal mirrored shares keep mirror augmentation a pure doubling of data rather than a bias.

    Applies when

    • designing reset distributions for get-up, recovery or multi-contact skills
    • mirror/symmetry augmentation is on and the task has chiral start states
    • no viewport is available to inspect resets on the training machine
    “角度 jitter ±15°、yaw 全域、0.28~0.40 m 低空放下由物理沉降,关节软限位内 均匀(留 5% 余量)+ 小随机速度。**不用 random SO(3)**(会采出穿地/极限卡死 等现实不可能状态,顾问同判) … side_l/side_r **概率必须相等**(镜像增强的样本同分布前提)”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §3 R0 任务定义 / §9 核查单
  • Two simulators disagreed 1.8x on one policy's torque demand - four explanations were eliminated with numbers, the surviving suspect was never tested, and neither reading was allowed to cancel the othertorque-disagreement-between-simulators-unresolved
    Hypothesisrecoverysim2sim-gatesim2simactuator-modelingattribution

    When two simulators disagree on a safety-relevant quantity, eliminate the measurement explanations one at a time with numbers, name the survivor as a hypothesis, keep the pessimistic reading binding, and run the direct test (replay one action sequence open-loop through both plants) before the quantity is used to pass a gate for hardware.

    Symptom

    For R3.0, Isaac read hip_pitch torque demand at 75-80% of the limit (PASS); MuJoCo read the same policy at 148-152% (1.5x over the limit).

    Context

    Candidates were eliminated one by one: PD gains (identical, RS06 kp 30 / kd 1.5), the settle window (both skip the first 0.5 s), the statistic (worse-of-two-legs vs per-joint - explains ~15%), self-collision (the no-contact subset still read 152%), sampling (500 Hz peak vs 50 Hz sample - ~9%). About 1.8x remained. The prime suspect - Isaac's implicit actuator (PD solved inside the PhysX integration, chosen to mimic kHz firmware PD) against MuJoCo's 500 Hz explicit PD - was named but untested.

    Change

    The Isaac PASS was recorded as valid for the Isaac plant only; the prescribed decider was an open-loop replay of one action sequence through both plants, compared step by step, needing no training. It was upgraded to a precondition of R3.1.

    Outcome

    R3.1 ran in parallel with, not after, the replay, and the replay never appears as done. By §40 it was "the oldest open account" on the line, to be fed by real-robot logs - which the spec also never records. Meanwhile the 50 Hz sample under-read the peak by 34% as smoothing narrowed the spikes, so Isaac-side readings grew more optimistic exactly when the comparison mattered more.

    Mechanism

    Each measurement artifact explained a slice of the gap; what remained is a plant difference, and an untested plant hypothesis cannot license discarding the pessimistic simulator on a safety-relevant quantity.

    Conflicts

    The spec prescribes the open-loop replay (§23), upgrades it to an R3.1 precondition, then records that R3.1 ran in parallel without it (§24), and lists it as the oldest open account before hardware (§40); nothing through §50 (2026-08-14) records a result. Every later "torque gate PASS" on this line is an Isaac-plant reading plus a MuJoCo check, never a reconciled one.

    Applies when

    • Isaac and MuJoCo (or sim and hardware) report different torques, contacts or slips
    • implicit vs explicit actuator models are in play
    • a gate passes on one simulator only
    “**同一个策略,一侧判 PASS 一侧判超限 1.5 倍。** … 口径项全部扣掉后仍剩 **~1.8×** 没有解释。 … 但**这条没有验证,不能拿它当结论去抵消 MuJoCo 的读数**。 … (做法:拿同一条 动作序列在两边开环回放,逐拍比 τ —— 不需要重训)”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §23 MuJoCo 复核门:成功率更稳,但 τ 与 Isaac 差 ~1.8× 没收敛
  • The advisor's "runtime five-stage state machine + per-stage reference poses + RL residual" appeared in none of the three papers it cited - reading the originals changed the plan and downgraded two widely repeated industry claimsadvisor-paraphrase-vs-paper
    Replicatedrecoveryprocessprocessattribution

    Read the primary source behind any piece of advice before adopting its architecture; record where the paraphrase and the original differ, downgrade claims the originals do not support to speculation, and adopt what the verified sources actually share.

    Symptom

    After the violent first real run, an advisor proposed re-architecting recovery as a runtime staged state machine with reference poses and an RL residual, citing HoST, HumanUP and StableMimic.

    Context

    The three papers were read in full on 2026-08-09 and tabulated (deployment form, what the "stages" really are, hard constraints, references). HoST: one end-to-end policy, height-gated rewards in training, action anchored as q + beta*a with a beta curriculum. HumanUP: two training stages with the same observation/action; the vendor's three-stage state machine is the baseline it beats (41.7% vs 78.3%); Stage II tracks an 8x slowed Stage I trajectory (4x too violent, 10x does not converge). StableMimic: a learned soft gate. On 08-10 more sources were checked the same way: the Agility page does not say Digit's self-righting was learned in simulation (only step recovery is stated as RL), so the claim was downgraded to speculation; HoST's support for the 12-DoF armless Mini Pi exists in its code repository, not in the paper text.

    Change

    Adopted only what the originals share: a hard action bound or anchored action space, strong smoothing including a second-difference term, a slowed hidden reference, heavy DR and real fallen states. The staged runtime state machine was not adopted.

    Outcome

    The next re-rooting candidates came straight from the verified material, and the beta-anchored action space (HoST, with the Mini Pi configuration as the nearest real-robot precedent) became V2, the lineage that later stood up on hardware.

    Mechanism

    A paraphrase compresses a paper into the advisor's own architecture; only the original shows what was actually deployed, what was a baseline, and which numbers came with which ablation.

    Applies when

    • an advisor, agent or summary proposes an architecture with citations
    • an industry claim ("X learned it in sim") is about to justify a design
    • several papers are cited for one combined recipe
    “⇒ **顾问的核心形态"runtime 五阶段状态机 + 每阶段参考姿态 + RL residual"在三篇引文 里均不存在**,其中 HumanUP 还点名 state machine 是局限。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 三篇引文精读判决(2026-08-09 全文核对;顾问转述与原文有出入)
  • The trainer read a stale USD after the URDF mass update - regenerate derived assets and gate on an automated equality instrumentderived-asset-staleness-check
    Mechanism understoodinfraplant-calibrationplant-calibrationprocesssim2sim

    For every derived plant artifact (USD from URDF, generated value files), pair the generation step with an automated source-vs-derived equality instrument, prove the instrument can fail, gate training on its PASS, and re-run physical audits after every regeneration.

    Symptom

    Measured link masses had been committed to URDF/MJCF (total 9.58 -> 9.792 kg, weighed values), but Isaac reads the derived USD asset - which still carried the old masses: a silent 2.2% mass fork between the training plant and the evaluation plant.

    Context

    The v12 checklist made USD regeneration a hard precondition ("硬性 前置") and, crucially, backed it with an instrument: check_usd_mass.py compares USD vs URDF per-link mass AND inertia trace, validated by showing it FAILs the stale asset naming 9 offending links, then PASSes after re-conversion (13/13 links consistent, total 9.7920). The self-collision filter audit was re-run after the regeneration (three poses, 0.00 N) because a regenerated asset invalidates physical audits done on the old one. This milestone was also where the three plants first aligned: "三边 plant 首次对齐(armature+摩擦+ 实称质量)就在这一代".

    Change

    convert_urdf re-run on the training machine, regenerated asset committed, check_usd_mass.py PASS required as an acceptance gate for the generation; dependent audits repeated post-regeneration.

    Outcome

    The 2.2% plant fork was closed before it could distort a generation's acceptance numbers; the staleness class of bug now has a permanent detector instead of a memory.

    Mechanism

    Source-of-truth edits do not propagate to derived binary assets by themselves; any consumer reading the derivative silently trains or evaluates on the old plant. An automated equality check between source and derivative - proven able to fail - turns an invisible staleness into a red gate, and regeneration invalidates every audit performed on the old artifact.

    Applies when

    • editing masses/inertia/geometry in URDF or MJCF sources
    • a trainer or evaluator consumes converted/derived assets
    • plant numbers differ between simulators for no visible reason
    “25ba997 把连杆质量更新为实称值(总重 9.58→9.792 kg,URDF/MJCF 已改),但 Isaac 读的是 train/assets/laika_v2.usd —— 仍是旧质量。… 否则 Isaac(9.58)与 MuJoCo(9.79)质量分叉 2.2%,v12 验收数字失真。验收门:python tools/check_usd_mass.py 必须 PASS … 对旧资产实测 FAIL/9 连杆点名,仪器已验证”
    train/WALK_V12_SPEC.md § 7. 核查单 (⚠️ 先重转 USD)
  • Curriculum-gate a penalty to the phase where its disease occurs - early on it only taxes explorationgate-penalties-to-the-disease-phase
    Replicatedwalkcurriculumcurriculumreward-shapingaction-rate

    For penalties aimed at late-stage pathologies (freezing, saturation, degenerate attractors), ramp the weight in only after exploration noise has decayed; anchor the terminal weight to measured healthy-vs-sick raw values, and shift all related tripwires to after the ramp completes.

    Symptom

    The action_saturation penalty, applied from iteration 0 in v8a, taxed exploration itself: with init_noise_std 1.2 the sampled actions paid ~-2.45/step before any policy had formed - while the disease it targets (clamp freezing) is a LATE pathology (v9 froze at iteration ~2624).

    Context

    v10 re-introduced the same penalty behind a curriculum gate: weight 0 until iter 1000, ramping linearly to -1.0 by iter 2000 - present only when the disease can occur, absent while exploration noise dominates. The trust argument was evidence, not hope: in v8a the term, while active, had pulled joint_pos_ref from 0.041 up to 0.155 and climbing - proof it can extract a policy from the frozen pit. Weight magnitudes were anchored to measured raw values (healthy v5 0.310 / v6 0.106 vs frozen v7 1.145 / v9 1.22 per step: at -1.0 healthy pays 6-18% of tracking, frozen pays 60%+, standing ~0). v10c then isolated the gated term as THE anti-freeze mechanism by single variable, upgraded to untouchable status in v11: "S 的门控机制(v10c 单变量铁案:任何情况下 不许撤,只许调终值)" - and v11 dared to relax other penalties only because S stood guard.

    Change

    action_saturation gated 0 -> -1.0 over iters 1000-2000 (later terminal value tuned -1.0 -> -0.5 with the gate mechanism itself frozen); tripwires adjusted to respect the gate's timing (freeze check moved to iter 2500-3000 to give the ramped term its effect window).

    Outcome

    Freezing stopped recurring while early training kept full exploration; the mechanism graduated from experiment to invariant within two versions.

    Mechanism

    A penalty's incidence depends on who occupies its support: early in training that is exploration noise (whose suppression starves learning), late it is the converged pathology. Time-gating aligns the penalty's presence with its target's presence, buying the constraint without the exploration tax - and tripwire timing must then be computed from the gate schedule, not from ungated precedents.

    Applies when

    • a structural penalty punishes exploration in early training
    • a late-onset pathology (freeze/saturation) needs a standing guard
    • deciding when a curriculum ramp should engage
    “v8a 实锤它的病根是"罚在采样动作上"——init_noise_std 1.2 的早期等于罚探索(~−2.45/步);而冻结是晚期病(v9 速率 2624 才死平)… 门控让它只在病发期在场。… v8a 里它在场时 joint_pos_ref 从 0.041 爬到 0.155 且仍在升——有从低谷爬出的实证。”
    train/WALK_V10_SPEC.md § 2. S 保险 —— action_saturation 课程门控
  • Set torque limits per joint from measured gait peaks - a uniform percentage is the wrong shape, and training must use the deployed numberstorque-limit-shape-by-measured-peaks
    Mechanism understoodwalkactuator-modelingactuator-modelinghardwareplant-calibration

    Measure per-joint torque peaks in the actual gait and set each limit as measured-peak x margin capped at rating; then propagate the same numbers into training and add an automated deploy-time consistency check - never derate by a uniform percentage, never let training assume torque deployment will not grant.

    Symptom

    A uniform 50% torque derating (18/8.5/7) had piled safety margin on the joints that never use it while cutting the busiest joint below half its measured demand.

    Context

    Per-joint gait peaks were measured (walk_v5 at cmd 0.3/0.6): RS06 (hip_pitch/knee) uses 5.5-5.9 N*m = 15-16% of its 36 N*m rating - cutting it to 12 is a free safety win; RS02's ankle_pitch runs at 16.2 N*m = 95% of its 17 N*m rating - "它是速度的硬件瓶颈", no room to cut; RS00 measured 36-44%, capped at 11. The resulting shape 12/17/11 replaced the uniform percentage. Sweeps across several limit sets (rated / 50% / 14-17-11 / 12-17-11) produced identical speed, lift, and landing force - within this range the limits do not shape the gait; what matters is consistency: "关键是训练和硬件必须是同一个数", because the exporter fills effort_limit from tau_limit, and a policy trained at rated 36/17/14 "会假设有三倍力矩可用" while deployed at 12/17/11 (exactly the v5 cross-generation inconsistency later suspected in its wild kicking).

    Change

    robot.yaml tau_limit set to the measured-shape 12/17/11, firmware written to match, and train/isaac_values.py regenerated so training sees the same limits; the deploy tool self-checks limits against robot.yaml on every run.

    Outcome

    Free safety margin captured where demand is low, the real bottleneck joint left at rating, and the train/deploy torque worlds unified with an automated consistency check.

    Mechanism

    Torque demand is grossly unequal across joints in a gait (15% vs 95% of rating here); a uniform percentage misallocates the safety budget by construction. And since the trainer treats effort_limit as a plant truth, any train/deploy mismatch is an invisible plant gap of exactly the mismatch ratio.

    Applies when

    • choosing safety torque limits for a legged platform
    • training-vs-deployment actuator limit audit
    • one joint runs near rating while others idle
    “曾用统一 50%(18/8.5/7)是错的形状: 把余量堆在用不到的 RS06 上, 却把 ankle_pitch 砍到需求的 52%。… RS02 在 0.6 m/s 已用到额定 95%, 它是速度的硬件瓶颈 … 实测多组限幅 … 完全一致 —— 限幅在这个范围对步态零影响, 关键是训练和硬件必须是同一个数。… 若训练仍按额定 36/17/14, 学出的策略会假设有三倍力矩可用。”
    train/WALK_V6_MINIMAL.md § 3. 训练侧必须同步的一件事
  • The hip_roll (l+r) asymmetry scalar predicted real-robot lateral drift - promote validated sim scalars into the gatehip-roll-sum-predicts-lateral-drift
    Replicatedomnireal-acceptancereal-acceptancegate-batterysim2sim

    Hunt for cheap sim scalars that predict real-robot behaviors, validate them on direction AND ordering across multiple policies, then promote them into the acceptance battery; treat later violations as debt to justify in writing, not noise to ignore.

    Symptom

    A persistent hip_roll left/right asymmetry row in the sim2sim symmetry table had been dismissed as "calibration or mechanical asymmetry" noise; meanwhile real deployments drifted sideways by policy-dependent amounts.

    Context

    Forward-kinematics analysis reframed the scalar: both hip_rolls move the feet in +y for positive angle, so a same-signed (l+r) sum IS a lateral translation mode - the scalar is a direct lateral-drift bias estimate. Checked against real deployments: s1e (l+r = -0.0178, smallest magnitude) was the steadiest with least drift; 700 (+0.0253) drifted mildly left; A800 (+0.0267) drifted clearly left with the largest tilt 12.9 deg. Direction correct 3/3, ordering correct 3/3 (the log's heading calls it "四枚四中", four-for-four).

    Change

    The scalar was promoted into the acceptance battery as a posture-class criterion alongside tilt-max median: "hip_roll 左右不对称 |l+r| 不得比父代大" - doubling as a heat proxy (error ~ torque ~ heating).

    Outcome

    Used at every later gate; when the C4 product exceeded it by +0.005 rad (~+0.3 deg vs parent), the criterion was not silently waived - it was booked as explicit debt with a mechanism argument (the increment is task-required, far smaller than the sidewalk amplitude +/-2.2 deg) plus a related account (stand saturation 32.4% -> 37.2%).

    Mechanism

    A policy's static joint-angle bias in a translation-producing mode integrates into real-world drift; sim can measure that bias precisely and cheaply. A sim scalar earns gate status exactly when its predictions are validated against hardware in both direction and ordering - and a validated gate may only be exceeded with a written mechanism-level justification, never silently.

    Conflicts

    The log's heading says "四枚四中" (4/4) but the evidence table lists three policies and the text says "方向 3/3、排序 3/3"; the fourth instance is not shown in this file.

    Applies when

    • a real robot drifts or leans in a policy-dependent way
    • deciding which sim measurements deserve gate status
    • a validated gate criterion is marginally exceeded by a new product
    “s1e | −0.0178(绝对值最小)| 微右、最不飘 | 三者中最稳、飘最小 ✓ … A800 | +0.0267 | 左、最飘 | 明显左飘、倾角最大 12.9° ✓ 方向 3/3、排序 3/3。 → 正式纳入验收表(与「倾角 max 中位」并列为姿态类判据)。”
    train/C_LADDER_RUN.md § 3e. 顺带:hip_roll 左右不对称 (l+r) 就是横移偏置 —— 四枚四中
  • action_rate weight is the sim2real bandwidth knob - re-tune it whenever a rate limiter is removedaction-rate-weight-vs-bandwidth
    Observed oncewalkreward-shapingaction-ratereward-shapingactuator-modeling

    Set action_rate weight relative to real actuator bandwidth, and re-tune it any time another smoothing/limiting element (filter, slew limiter, gain) changes - reward weights are load-bearing parts of the actuator model.

    Symptom

    With a low action_rate_l2 weight the policy learns fast actions; the unmodeled part of the actuator response is then excited hardest, and sim2real "直接崩" (collapses outright). With too high a weight, actions become so slow the robot cannot maintain balance.

    Context

    The reference developer called action_rate_l2 the single most important reward for transfer, with side-by-side video evidence that the high-penalty, slower policy is clearly better on hardware. Lucen context: the team had just removed the SOFT_SPD=1.0 velocity limiter, which had been an implicit actuator-bandwidth constraint - leaving action_rate as the only remaining constraint on action speed.

    Change

    Decision recorded: after removing SOFT_SPD, re-evaluate the action_rate weight rather than keep the old value, since its effective role changed from "additional smoother" to "sole bandwidth constraint".

    Outcome

    Logged as a priority follow-up ("重新评估 action_rate 权重 - 拆掉 SOFT_SPD 之后这一项的作用变了"); the failure mode it guards against is training high-frequency actions the real actuators cannot track.

    Mechanism

    Slower actions stay inside the frequency band where the ideal-PD sim actuator and the real actuator agree; fast actions probe the band where unmodeled delay, inductance, and bandwidth limits dominate, so model error is amplified in exact proportion to action speed. Any removed external rate limit transfers that constraint's entire job onto the action_rate penalty.

    Conflicts

    The low/high tradeoff evidence is the external developer's report (with video); the Lucen-side entry is a pre-registered risk and decision, not yet an on-robot A/B at the time of writing.

    Applies when

    • removing or adding an action filter, slew limiter, or low-level speed cap
    • real robot shows high-frequency chatter or overheating absent in sim
    • tuning smoothness rewards before a hardware deployment
    “权重低 → 动作快 → 执行器模型不准的部分被放大,sim2real 直接崩 / 权重高 → 动作慢 → 好迁移,但可能慢到无法维持平衡 … 我们刚拆掉 SOFT_SPD=1.0 的限速器,等于把执行器带宽约束整个移除了。action_rate 惩罚现在是唯一还在约束动作速率的东西,需要重新评估权重”
    Experience.md § action_rate_l2 是他认为最关键的 reward (lines 61-70)
  • Close a question with an audit, then freeze the wording - later symptoms may not reopen it without new hard evidencefrozen-verdicts-semantic-boundaries
    Mechanism understoodomniprocessprocessattributionplant-calibration

    When an audit closes a hardware-vs-policy question, record the closing evidence, freeze a citable wording for future recurrences, and set the reopening bar explicitly; separate robustness perturbations from plant-truth questions so a DR rung's failure can never silently reopen a closed measurement.

    Symptom

    Recurring directional bias on the robot kept re-suggesting "maybe the hardware/COM/mechanics are asymmetric", threatening to re-litigate questions that audits had already closed - burning attention each time a descendant policy leaned or drifted.

    Context

    Two boundary decisions were written as permanent: (1) semantic separation - "S2④ COM ±20mm = 纯鲁棒性扰动,不再承担「解释真机后仰」任务" - if the COM-DR rung degrades, the ONLY allowed conclusion is "policy insufficiently robust to COM uncertainty"; reopening "is the CAD COM wrong" is forbidden because the mass audit was completed and closed (@63f9212). (2) a frozen wording for chirality, to be quoted verbatim whenever left/right bias appears in later rungs: observed directional bias = policy-level spontaneous symmetry breaking; plant asymmetry = no supporting evidence after the mass + model symmetry audit; mitigation candidate pi_sym queued, not blocking. The evidential basis was quantitative: the root policy was perfectly symmetric under +/-6 N*s pushes (40/40) while descendants broke (17/40, 13/40) - "手性是 S2 训练中获得的, 根没有; 机械侧已双 PASS 关案, 不重开".

    Change

    Closed questions carry (a) the audit commit that closed them, (b) a frozen citable wording for recurrences, and (c) an explicit evidence bar for reopening ("无新硬证据不得重开").

    Outcome

    Later chirality observations (C2's 15 pp turn gap, hip_roll drift bias) were handled as policy-lineage properties with policy-side mitigations, without a single hardware re-audit cycle.

    Mechanism

    Symptom classes recur under different guises; without a frozen verdict each recurrence re-runs the same expensive investigation and risks a different (worse-informed) conclusion. Freezing verdict plus wording converts recurring symptoms into citations, while the evidence bar keeps the closure honest rather than dogmatic - the root/descendant symmetry comparison is what makes "it's the training, not the machine" checkable at any time.

    Applies when

    • a recurring symptom keeps suggesting an already-audited hardware cause
    • writing conclusions for a completed calibration/audit
    • a DR rung's degradation invites re-measuring the plant
    “若 S2④ 退化,结论只能是「当前 policy 对 COM 不确定性不够鲁棒」,不得重开「CAD COM 是不是错了」… 手性冻结表述 … Plant asymmetry: no supporting evidence after mass + model symmetry audit … 无新硬证据不得重开机械不对称”
    train/OMNI_V0_SPEC.md § 4. 语义分界与手性冻结表述(2026-08-07 用户定,永久)
  • Knee swing collapsed because it directly trades against the slip penalty - price the conflict explicitly and clamp what reward cannot holdknee-swing-vs-slip-pricing
    Mechanism understoodwalkreward-shapingreward-shapingattribution

    When a behavior collapses as another metric improves, look for the term pair trading them and set their price ratio deliberately (with escalation and reverse tripwires pre-registered); where the policy actively spends action budget to undo your target, stop paying more reward and clamp the target space structurally.

    Symptom

    Knee peak-to-peak swing collapsed across generations - v5 33 deg, v10 26-30, v10b 7-8, v11 6.5-8.6 - and rolling the clock back did not recover it, acquitting the clock; the collapse tracked the gated slip penalty instead: "屈膝与不打滑在当前奖励里直接对抗" - v10b's excellent 93 deg slip was purchased with knee amplitude.

    Context

    Reward-side flexion fixes had failed three times: raising reference amplitude backfired twice (v9/v11), and v11's deep-squat default was actively fought by the policy - it spent 0.68 of action budget pulling the squat straight ("被策略花 0.68 动作拉直反杀"). v12's design accepted the conflict as real and attacked on two tracks: (1) ECONOMICS - a direct knee_swing_amplitude reward (+0.3, target 0.55 rad, capped at 0.6/step = 55% of tracking), explicitly opposed to the slip penalty by design ("显式对立——这正是设计:v12 就是这场对抗的定价实验"), with an escalation ladder (K +0.3 -> +0.5, then slip -0.5 -> -0.3, one layer at a time) and a reverse tripwire (slip telemetry back at v10 levels -> slip weight to -0.8, accept ~20 deg knee compromise); (2) STRUCTURE - knee target bounds [0.2, 0.9] rad so full straightening is physically impossible (straightest 11.5 deg) and the 0.68 fighting budget is released. A bonus falsifiable prediction was attached: phase-lock strength tracks amplitude (v9_probe 48 deg locked 2.5 Hz; v11 low-amplitude 1.36 Hz unlocked), so if K works, hardware phase-lock should return - one change, two verdicts.

    Change

    knee_swing_amplitude reward + knee target clamp + pre-registered escalation/reverse levers; the failed reward-side-only approach retired.

    Outcome

    The lineage was frozen before v12 trained (strategic reset), but the diagnosis stands as the walk line's clearest example of two reward terms trading a behavior between them, with the pricing experiment and structural clamp fully designed and calibrated.

    Mechanism

    When two terms price opposite aspects of one motion (swing amplitude creates yaw momentum that becomes slip), the optimizer settles wherever the price ratio puts it - patching one side moves the equilibrium, not the conflict. Explicit pricing makes the trade a designed quantity; structural clamps remove the regions where the policy spends budget fighting the designer.

    Conflicts

    The pricing experiment (K vs slip) was designed and calibrated but never trained - the 2026-08-05 reset suspended v12; the collapse attribution table and the 0.68-action counterattack are measured, the remedy's效果 is untested.

    Applies when

    • one gait quality degrades in lockstep with another's improvement
    • the policy visibly fights a default pose or reference
    • repeated reward-side fixes for the same behavior have failed
    “膝摆塌在 v10→v10b,头号嫌疑是门控滑移罚(四代实测膝 p2p:v5 33° / v10 26~30° / v10b 7~8° / v11 6.5~8.6°;退时钟没救回 → 非时钟)——"屈膝"与"不打滑"在当前奖励里直接对抗 … 奖励侧修屈膝已三败 … v11 深蹲 default 被策略花 0.68 动作拉直反杀”
    train/WALK_V12_SPEC.md § 0. 定位 / 2. K —— 膝摆经济(与滑移罚的对偶)
  • The get-up kept getting faster because standing earlier paid more every step - lowering torque authority barely slowed it, and only zeroing the standing income for the first 3 s moved the pace into the design bandper-step-income-drives-speed-time-gate
    Mechanism understoodrecoveryreward-shapingreward-shapingcurriculum

    When a skill is too fast, find the term that pays for finishing early and gate that income by time; keep the "get into position" term ungated so the policy does not learn to wait, use a ramp instead of a cliff, and confirm with a paired same-level experiment that the drift is motivational before changing it.

    Symptom

    The user judged the get-up too fast (Isaac medians about 0.7-1.6 s) and suspected path dependence: the policy seemed to get faster the longer it trained.

    Context

    Lowering the beta authority 0.40 -> 0.30 cut impact but moved supine only 1.70 -> 2.00 s: coordination-limited, not torque-limited. A paired experiment inside one beta level (checkpoint 15,600 vs 18,499, +2,900 iterations, same ruler) measured the drift: get-up medians -7 to -10%. The spec concluded the motive, not the path, was the cause - per-step standing income pays for every early step, and any lineage (even one from scratch) races toward the fastest solution inside its constraints.

    Change

    V2.5: the standing income (base_height, stand_pose, still, feet_on_ground) multiplied by w(t) = clamp(t/3 s, 0, 1); upright deliberately NOT gated, so righting and sitting up early still pay and the policy is not taught to lie flat and wait; a ramp, not a step. V2.5b: zero before t0 = 3 s, then a 1 s ramp.

    Outcome

    V2.5: Isaac 100%, get-up +17-43% slower, MuJoCo 100/100/98/100% (the best cross-simulator reading yet), still short of the 3.5-4.5 s design band - a linear ramp only discounts early income. V2.5b: MuJoCo supine 2.04 -> 4.10 s and prone 3.18 -> 4.04 s, inside the band; the Isaac pace barely moved (a lineage habit on a gradient-free plateau). Later the zero gate proved harmful when trained from scratch (curriculum-history-is-part-of-the-product) and in the V3.1 lineage (time-gate-vs-wide-stance-retire-the-fix).

    Mechanism

    Constraints on authority or velocity change how the fastest solution looks; the time structure of the task income decides how fast the fastest solution is.

    Applies when

    • a policy is faster or more aggressive than wanted and constraints do not slow it
    • progress-style rewards pay every step spent at the goal
    • performance drifts faster with more training at fixed settings
    “**V2.5 机制(唯一)**:站立收入(base_height/stand_pose/still/feet_on_ground) 乘时间斜坡 w(t)=clamp(t/T_gate,0,1),T_gate=3.0 s;**upright 刻意不门控** (翻正/坐直早期照常拿钱,防"躺平等门开" … **用户假设量化 证实:逐步计酬动机在 β 包络内持续压缩时间,约束挡不住动机 —— V2.5 动机层 修法为正解。**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §39 V2.5 预注册 / §39 补 配对实验
  • An exponential kernel on instantaneous velocity punishes gait oscillation - track the cycle averagecycle-average-tracking-for-gait-quantities
    Mechanism understoodomnireward-shapingreward-shapingattribution

    Reward velocity tracking on gait-cycle averages (or filtered values), not instantaneous samples, whenever the desired behavior oscillates at stride frequency; widening the kernel does not fix a variance penalty.

    Symptom

    Even while the robot genuinely sidewalked (verified after the metric fix), the Isaac-side tracking reward sat on the ignore-floor: true sidewalk scored 0.178 vs 0.189 for ignoring the command - the reward was mildly punishing the desired behavior.

    Context

    Sidewalking is inherently oscillatory: per-frame vy std was 0.177 while the tracking kernel width was sigma = 0.15, applied to the instantaneous value. A kernel-width scan showed widening sigma 0.15 -> 0.50 still loses (-0.14 -> -0.07): "指数核惩罚的是方差,而侧步天生带方差" (the exponential kernel penalizes variance, and side-stepping inherently carries variance). Modeling with measured parameters: replacing instantaneous vy with the mean over one gait cycle (0.5 s) flips the margin decisively (true sidewalk 1.888 vs ignore 1.281, +0.607), half-cycle is neutral (+0.006), two cycles adds nothing more. Explicitly flagged as extrapolation pending Isaac-side implementation. This also vindicated a previously dismissed external note (sigma too small) - right conclusion, different mechanism than claimed (variance, not gradient).

    Change

    Proposed fix recorded: change the tracked quantity from instantaneous vy to a one-gait-cycle running average; widening sigma alone rejected by the scan.

    Outcome

    Diagnosis complete and quantified; the C4 product shipped via feed-forward before the reward change was implemented, so the cycle-average fix remained a verified-by-model, not-yet-trained change.

    Mechanism

    E[exp(-(v-c)^2/sigma^2)] decreases with Var(v) even when E[v] = c exactly; a gait's phase-locked oscillation guarantees variance at the stride frequency, so instantaneous tracking rewards structurally prefer standing still at the command mean. Averaging over exactly one cycle removes stride-frequency variance while preserving command-following error.

    Conflicts

    The cycle-average fix itself is model-extrapolated ("⚠️ 这一条是外推,须在 Isaac 侧实装并复量后才能当结论") - the diagnosis is measured, the remedy untested in training at the time of writing.

    Applies when

    • tracking rewards for lateral/turn/any oscillation-carrying velocity
    • a verified behavior scores below the ignore-floor
    • choosing sigma for exp-kernel tracking terms
    “侧走时 vy 的逐帧摆幅 std = 0.177,而 track_lin_vel_y_exp 核宽 σ = 0.15,且作用在瞬时值上 … 真侧走(均值 66%,振荡 ±0.18)0.178 | 完全无视指令 0.189 … 真侧走的得分比无视指令还低。… σ 从 0.15 放到 0.50,侧走仍然吃亏 … 把跟踪目标从瞬时 vy 换成一个步态周期(0.5 s)的平均 vy:… 1.888 vs 1.281”
    train/C_LADDER_RUN.md § 3n. 二/三 Isaac 训练奖励为何一直坐在「无视底分」/ 修法不是放宽 σ
  • Ungated phase shaping made standing 42x more expensive than stepping - and the stepping was cooking the hip motorsmoving-gate-42x-stand-tax
    Mechanism understoodomnireward-shapingreward-shapinghardwarereal-acceptanceprocess

    Gate every phase/clock-driven shaping term on the command that justifies motion, price the cmd=0 case explicitly during design - and when a reward flaw is sim-only, book it with trigger conditions instead of operating immediately on a working lineage.

    Symptom

    At cmd = 0 the sim policy never stood still - it stepped in place and crept 0.98 m per 20 s; on the real robot the same lineage's stepping made hip_roll motors run 20 degC hotter than every other joint (43-48 degC vs 25-28).

    Context

    The arithmetic closed it: the three phase-shaping terms (joint_pos_ref 1.6, feet_contact_number 1.2, feet_clearance_swing 1.6) are driven by the gait clock with NO command gating, so standing at cmd=0 forfeits 3.32/step of shaping while honest standing earns only 0.078 of tracking - stepping wins 42x. The heat chain: perpetual stepping = perpetual single support = one hip_roll stalled at ~4.3 N*m (25% of torque limit) carrying the torso's frontal-plane moment - two hip_rolls = 90% of whole-machine steady-state I2R; measured stand-vs-step comparison: total heat -71% when actually standing. The suspended test acquitted the actuator (0.21 rad sag -> 0.0008 rad in air) and mechanics vetoed the easy fix ("降 kp 救不了热" - equilibrium torque equals the external load regardless of kp). The fix (moving_gate: hard-gate the three shaping terms on |cmd| > eps) was designed - then DEFERRED by the user because the real robot at the time stood fine: "真机不表现该问题, 为真机不存在的病改奖励表不划算", with written trigger conditions (C-ladder stand row persistently failing, or real robot starting to step/drift at cmd=0) and the known hazard tag (this is exactly the reward-change class that triggers the B-arm signature). When the real robot later DID step and cook, the booked trigger fired and moving_gate moved from debt to to-do with its benefit re-priced: "停止空烧 hip_roll,稳态发热降 ~71%".

    Change

    moving_gate designed with the gate_by_cmd convention; deferral, triggers, and expected heat recovery all pre-registered instead of patching the reward for a then-sim-only symptom.

    Outcome

    The cmd=0 stepping went from mystery to closed arithmetic; the thermal measurement (hip_roll +20 degC) quantitatively confirmed the 90%-of-heat prediction; the reward change waited for real-world justification instead of spending a risky revision early.

    Mechanism

    Clock-driven shaping terms define a perpetual-motion bounty unless gated by command; the resulting idle gait is not a training bug but the table's optimum. Its cost surfaces on hardware as stall-torque heating set by statics (mass x lateral offset), which no gain change can remove - only removing the motion (gating) or widening the stance can.

    Applies when

    • the policy steps in place or creeps at zero command
    • specific joints run hot in idle behaviors
    • deciding when a known reward flaw justifies a risky mid-lineage fix
    “塑形合计 −3.32/步 … 净: 站定亏 42 倍 … 两颗 hip_roll 4.29 / 4.09 N·m(各占限扭 25%),占全机稳态 I²R 的 90% … 吊挂实测 hip_roll 跟踪误差 0.21 rad → 0.0008 rad … 降 kp 救不了热 … 真机不表现该问题, 为真机不存在的病改奖励表不划算”
    train/README.md § C1 FAIL 节 (cmd=0 的 sim/真机分歧记账) / C2 真机 A/B 三b 发热定性
  • The real robot's right-leg kicking was over-trained-delay times loop gain - irreducible pipeline latency is plant, model it fully from day onepipeline-latency-is-plant-not-dr
    Mechanism understoodomniactuator-modelingactuator-modelingreal-acceptanceattributionsim2sim

    Measure the end-to-end action pipeline delay and build it into the nominal plant and every acceptance gate from day one; treat power/scale deratings that "fix" oscillation as gain-reduction crutches flagging an unmodeled delay, and expect higher-feedback-gain policies to be MORE delay-fragile.

    Symptom

    On hardware, s1c/s1d at action scale 1.0 always kicked wildly with the right leg (s1c only ran as SOTA at power 0.8; s1d only at 0.7) - while sim showed nothing under default evaluation.

    Context

    Sim reproduced the incident item by item once the real pipeline delay was injected: s1d@1.0 with --delay 1 fell at 10.2 s, --delay 2 at 5.2 s; s1c@1.0 stressed (r_hip_roll saturation 5 -> 16%; "右脚" = the policy's chirality makes the right leg its high-gain leg); and the combos that worked on hardware (s1c@0.8+delay2, s1d@0.7+delay2) all survived in sim. Mechanism: the real pipeline is ~1-2 ticks (BusWorker next-cycle pickup + CAN round trip) but S1.1 trained only to 1 tick - "超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定" (delay beyond training x full loop gain = oscillation; the power derating had been buying stability by compressing loop gain). s1d was MORE fragile than s1c because its yaw 3-layer stack had learned higher feedback gain - higher gain, lower delay tolerance. Three changes: latency DR widened to cover reality; acceptance gates and smoke runs moved permanently to --delay 2 ("门必须在真机条件下预测 真机"); and the doctrine written twice-paid: "不可约的管线属性(延迟、 限速)不是'随机化选项',是 plant 本体,第一天就该全额建模" - S1's nominal-then-robust staging falsified by hardware for the second time. The later s1e hardware run at power 1.0 (no kicking, normal force) closed the loop: "0.8 = 旧代拐杖" - the derating had been a crutch for the under-modeled delay, not a real requirement.

    Change

    Latency modeled as plant from day one of any lineage (measured 1-2 ticks covered, bridge-layer rate limits likewise modeled by default); every gate and smoke evaluation issued under --delay 2.

    Outcome

    Kicking reproduced, explained, and eliminated in the s1e generation at full scale and full power; the deploy-side crutches (0.7/0.8) retired for the new lineage.

    Mechanism

    Feedback oscillation onset is a product of loop gain and phase lag; a policy trained below the real delay learns gains that sit past the real stability margin, and any output derating masks it by scaling gain down. Since pipeline delay is deterministic hardware property - not an uncertainty - it belongs in the nominal plant, and every evaluation must include it or the gate predicts a robot that does not exist.

    Applies when

    • hardware oscillation/kicking that sim only reproduces with added delay
    • a policy only runs on hardware at reduced power/scale
    • defining what belongs in the nominal plant vs the DR list
    “真实链路延迟 ~1~2 拍 … S1.1 只训到 1 拍——超训延迟 × 全环路增益 = 振荡;衰减 = 压环路增益换稳定。s1d 比 s1c 更脆 = yaw 三层栈学出更高反馈增益,增益越高延迟容忍越低。… 教训入账:S1「先标称后鲁棒」第二次被真机证伪——不可约的管线属性(延迟、限速)不是"随机化选项",是 plant 本体,第一天就该全额建模。”
    train/OMNI_V0_SPEC.md § 3. S1.4(真机右脚乱踢事故强制)
  • Training-log reward values and fixed-command eval values live on different distributions - comparing them once claimed a 44% improvement that was really 6-10%same-distribution-reward-comparison
    Mechanism understoodwalksim-evalmeasurementattributionprocess

    Quote reward-term values only with their distribution attached (command range, DR on/off, environment), and compare across runs only when those match; re-measure in a common environment before claiming any improvement percentage.

    Symptom

    A v6-era analysis concluded slip had dropped 44% by comparing the training log's Episode_Reward against values calibrated in a fixed-command play environment; a same-condition re-measurement showed the true improvement was 6-10%.

    Context

    The training log's reward is an expectation over the training command distribution (vx 0.15-0.5, yaw +/-0.6, with pushes and domain randomization); play-environment calibrations are taken at a single fixed command with DR off. Subtracting one from the other compares apples to oranges - the warning was written into the v7 pre-flight: "奖励数值只能在同一指令分布下比较 … 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次)".

    Change

    Rule adopted: any before/after reward-term comparison must hold the command distribution, DR state, and evaluation environment fixed; training-log values compare only against training-log values of runs with identical command/DR configs.

    Outcome

    The phantom 44% improvement was retracted; later term-level accounting (e.g. the C4 ignore-floor work) consistently specified its distribution before quoting numbers.

    Mechanism

    A reward term's expectation depends on the visited-state distribution as much as on the policy; changing the command distribution or DR moves every term's baseline. Cross-distribution differences therefore measure the distributions, not the policy change.

    Applies when

    • comparing reward telemetry across training runs or vs play evals
    • claiming improvement percentages from training logs
    • term-level reward accounting for diagnosis
    “奖励数值只能在同一指令分布下比较。训练日志的 Episode_Reward 是在训练指令分布上算的(vx 0.15~0.5 / 偏航 ±0.6 / 带推力与域随机化), 拿它和固定 cmd 的 play 环境标定值相减会得出错误结论(v6 那轮已经栽过一次: 据此以为滑移降了 44%, 同条件对拍只有 6~10%)。”
    train/WALK_V7_SPEC.md § 3. 开训自查 ⚠️
  • Cross-simulator gate (Isaac Lab to MuJoCo) comes before any hardware attemptsim2sim-gate-before-sim2real
    Replicatedwalksim2sim-gatesim2simprocessgate-battery

    Gate every policy through a second simulator with an independently built plant before hardware; treat sim2sim failure as a contract or overfitting bug, and sim2real failure after a sim2sim pass as a plant/actuator gap.

    Symptom

    A policy that only ever ran in its training simulator carries untested dependencies on that simulator's solver, contact model, and defaults; the first place those dependencies surface should not be the real robot.

    Context

    Standing order of operations for the whole Lucen program, recorded as the opening line of the experience log: train in Isaac Lab, gate in MuJoCo, only then go to hardware. The MuJoCo side is the same plant used for evaluation batteries, so a sim2sim pass also validates the exported policy + contract (obs ordering, scales, defaults) outside the training stack.

    Change

    Pipeline rule adopted - every checkpoint must pass the MuJoCo evaluation battery (sim2sim) before it is considered for real deployment (sim2real).

    Outcome

    Used generation after generation as the cheap filter; hardware sessions only ever received policies that had already survived a second simulator.

    Mechanism

    Two simulators disagree exactly where a policy is overfit to simulator-specific artifacts (contact softness, integrator, default parameters, obs conventions); a cross-sim transfer catches contract bugs and solver overfitting at zero hardware risk, so real-robot failures that remain are attributable to genuine plant/actuator gaps.

    Applies when

    • planning the path from training to first hardware trial
    • exported policy behaves differently outside the training framework
    • triaging whether a real-robot failure is contract vs plant
    “先sim2sim - 从isaaclab 到mujoco / 再sim2real”
    Experience.md § opening lines (1-2)
  • A reward on a quantity the actor cannot observe teaches "produce less of it", never "correct it" - closed-loop correction needs an outer loopreward-observability-limit
    Mechanism understoodomniobservation-designobservation-honestyreward-shapingcontract-freeze

    Before adding a reward, check the actor can observe (or infer) the quantity: unobservable-error rewards buy only average suppression - route correction tasks to an outer loop whose commands stay in distribution, and do not break a frozen contract to add an observation a deploy-side loop can supply.

    Symptom

    Heading kept drifting despite world-frame yaw rewards, and a reviewer proposed heading-error rewards - raising the question of what yaw shaping can even teach this actor.

    Context

    The adopted architectural verdict: the actor's 45-dim base observation cannot see accumulated heading at all - projected_gravity is invariant to rotation about the gravity axis, and omega_z is a rate, not an angle. World-frame yaw-rate rewards are therefore privileged shaping that can only teach "少产生旋转" (generate less rotation), never "偏了以后拉回原线" (pull back to the line after drifting) - the policy cannot represent the error it would need to correct. The S1 gate (<=5 deg / 10 s) demands exactly the former, so the stack is right for its gate; active heading correction is assigned to the deployment outer loop (--heading P-loop converting heading error into in-distribution wz commands) plus small-wz training - and the 215-dim contract is explicitly NOT extended with a heading observation ("契约不加 heading 观测,冻结不动"). The reviewer's companion bias hypothesis was adjudicated with data: drift is bimodal - a basin mechanism decides whether you leave (seeds vary +/-16-46 deg vs -385 to -391 deg), and once out, rotation direction is constant (weight chirality; candidate root: the phase clock always swings left first).

    Change

    Yaw shaping kept as rate-tracking (three-layer stack); heading correction owned by the deploy outer loop; contract frozen; the "which behaviors need an outer loop" question settled by observability analysis rather than reward tuning.

    Outcome

    Stopped a contract change and a futile reward direction; drift work split correctly into rate-suppression (trainable) and error correction (outer loop), consistent with the earlier measured 10x drift reduction from the deploy-side loop.

    Mechanism

    A policy can only condition on its observation sigma-algebra; rewards on functions outside it shift the marginal action distribution (open-loop average effects) but cannot create feedback on the unobserved variable. Whether to add an observation, an outer loop, or accept average-shaping is decided by the task's gate: suppression gates need shaping, correction gates need the variable in some loop's view.

    Applies when

    • adding rewards on accumulated/世界-frame quantities (heading, position)
    • deciding between a new observation, an outer loop, and shaping
    • a drift symptom persists across reward-weight changes
    “actor 的 45 维基座观测不到累计航向(projected_gravity 对绕重力轴旋转不变,ωz 是速率不是角度)——世界系 yaw 奖励是特权塑形,只能教「少产生旋转」,不能教「偏了以后拉回原线」。… 主动纠偏闭环 = S3 把小 wz 进分布 + deploy --heading 外环 … 215 契约不加 heading 观测,冻结不动。”
    train/OMNI_V0_SPEC.md § 3. 评审④判决(2026-08-06,S1.3 开训前)
  • Teleop fed the sidewalk axis a command beyond its training band - feet clipped; give each axis its own speed settingteleop-command-band-per-axis
    Mechanism understoodomnireal-deployreal-acceptancehardwareattribution

    Give every command axis its own teleop scale, clamped to that axis's training band, and reproduce any hardware incident in sim with the exact deployed command values before touching training.

    Symptom

    Robot stepped on its own foot when sidewalking left under teleop - and only when going left.

    Context

    The teleop tool used one speed setting for all axes: --teleop-speed 0.20 applied to A/D sent cmd_vy = 0.20, above the training band's top (0.08-0.18) where foot-spacing margin is thinnest. Sim reproduction of the incident (product policy, pw0.8, 5 seeds x 20 s, true collision threshold = single foot width 104 mm): at vy 0.20 the minimum foot distance was 111-115 mm - 7-11 mm from self-collision - vs 147 mm at vy 0.10. Left was 4x more dangerous than right (25% vs 6% of time inside the 160 mm soft wall at vy 0.10), matching the left-only symptom; the margin did not degrade over time (pressing more just lengthened exposure).

    Change

    deploy_policy gained --teleop-side (default 0.10), separating the lateral speed from the forward speed so each axis's teleop command sits inside its own trained band.

    Outcome

    Command now inside the band with 43 mm margin at default; the incident became a quantified, reproduced, closed account rather than a mystery.

    Mechanism

    The policy's competence envelope is the training command distribution per axis; teleop mappings that share one scalar across axes silently command out-of-band inputs on the weakest axis. Asymmetric risk (left vs right) came from the policy's own chirality bias, so a symmetric command produced an asymmetric hazard.

    Applies when

    • wiring a joystick/teleop layer over a learned policy
    • a hardware incident occurs on one command direction only
    • training bands differ across command axes
    “A/D 一直与 W/S 共用速度档,所以按 A 下发的是 vy = 0.20 —— 既超训练带(0.08~0.18)上沿 … 0.20(遥控实际值)| 111~115 mm | 7~11 mm … 且左比右危险 4 倍 … 处置:deploy_policy 新增 --teleop-side(默认 0.10),侧移与前进档分开。”
    train/C_LADDER_RUN.md § 3p. 一 向左走踩到自己 → --teleop-speed 0.20 同时喂给了 vy
  • Symmetrizing the config made the gait MORE asymmetric - the asymmetry lived in the policy weightsasymmetry-in-weights-not-config
    Mechanism understoodwalkattributionattributionreward-shapingcurriculum

    Localize a persistent asymmetry by intervening at the config layer first: if the symptom survives (or worsens), it is in the weights - fix it with symmetry-constrained training, not with trims or offsets.

    Symptom

    walk_v1 on hardware: straight-line command curved 149 deg in 15 s (9.9 deg/s) with 3.06 m lateral runout; turn gain +31% one way vs +129% the other (75% difference); knee asymmetry 4.4 deg in sim, 9.6 deg on the robot.

    Context

    The obvious suspect was the asymmetric default pose in the config. The decisive test: symmetrize standing_pose and run the SAME policy in sim - the asymmetry got LARGER (hip_pitch 6.8 -> 9.3 deg). Root cause therefore not in the config but baked into the policy weights: PPO without a symmetry constraint commonly converges one-sided, because splitting the work 50/50 and loading one side yield the same return, and the gradient falls randomly into one of the equivalent optima.

    Change

    Fix redirected from config trimming to retraining with mirror data augmentation (walk_v2 spec) - a weights-level fix for a weights-level disease.

    Outcome

    With augmentation (and the symmetric-default precondition), stand_v1 reached 0.0 deg asymmetry on all six joint pairs (from 4.4-7.7 deg), height fluctuation 7 mm -> 1 mm, mean |action| down 33%.

    Mechanism

    Reward-equivalent solution families (who carries the load) leave the symmetric solution unpreferred; SGD picks an arbitrary member and entrenches it. Config changes move the coordinate frame around the entrenched asymmetric function - they cannot move the function. The counterfactual test (change config, watch symptom) localizes the layer the disease lives in.

    Applies when

    • a robot veers or loads one side despite a symmetric-looking config
    • deciding between config trims and retraining for an asymmetry
    • mirrored-turn gains differ by tens of percent
    “根因不在配置里:把 standing_pose 对称化后在 sim 里跑同一策略,不对称反而变大(hip_pitch 6.8°→9.3°)—— 说明不对称烙在策略权重里。这是无对称约束的 PPO 的常见收敛结果(左右各担一半与一边多担的回报相同,梯度会随机落进其中一个)。”
    train/RETRAIN_v2.md § 1. 为什么是对称增强(证据)
  • Calibrate a wall penalty by measuring healthy and sick policies - healthy pays ~0, the disease pays a wallcalibrate-threshold-between-healthy-and-sick
    Mechanism understoodwalkreward-shapingreward-shapinggate-battery

    Calibrate every threshold penalty by evaluating its exact formula on replays of at least one healthy and one sick policy: place the threshold between their distributions, size the weight so the sick policy pays a decisive fraction of tracking while the healthy one pays ~0, and pre-compute neighboring thresholds for cheap adjustment.

    Symptom

    Real walk_v7 occasionally clipped its own legs (stance narrowed to 133 mm mean vs nominal 214.5); a foot-distance penalty was needed, but an uncalibrated threshold/weight risked either doing nothing or becoming a reverse barrier.

    Context

    The term (relu(d_min - lateral foot distance), measured in the base yaw frame because world-frame y is meaningless after turning) was calibrated before training by replaying three known policies through the exact reward formula at the acceptance operating point: healthy v5 (183 mm) pays 0.6% of tracking - effectively free; narrowed v7 (133 mm) pays 22% - effective widening pressure; collapsed v8 (93 mm) pays 55% - a wall. d_min 0.16 was placed deliberately between healthy and sick, with alternative thresholds (0.14/0.18) pre-computed in the tool output for later adjustment. The shape self-check was named as a standing question: "先问'零代价的选项是什么'" - the zero-cost region must be exactly the desired behavior.

    Change

    feet_lateral_distance added at d_min 0.16 / weight -10, with telemetry expectation pre-registered (should decay toward 0 as stance learns >160 mm; if bow-legged over-widening >214 appears, only then discuss an upper bound).

    Outcome

    v9_probe on hardware: no leg contact ("没碰腿(N2 兑现)"), stance min 145/126 mm green; the term's zero-cost design left healthy gait untaxed.

    Mechanism

    A relu threshold penalty defines a free region and a priced region; its correctness is entirely in where the boundary sits relative to the healthy and pathological distributions. Replaying known-good and known-bad policies through the exact formula measures both distributions in the term's own currency, making the threshold and weight a placement decision instead of a guess.

    Applies when

    • adding any relu/threshold-style wall penalty
    • a safety margin (foot distance, joint limit, clearance) needs enforcement without taxing normal behavior
    • choosing between candidate thresholds for a new term
    “形状自检 (v6 横向组/v8-B 的教训 —— 先问"零代价的选项是什么"): 标称站距付 0, 健康步态付 ~0, 收窄才付费 … v5(健康) 183 mm … 0.6% ≈ 免费 | v7(收窄) 133 mm … 22% —— 有效推宽 | v8(塌陷) 93 mm … 55% —— 墙 … d_min=0.16 恰在 v5(183)与 v7(133)之间”
    train/WALK_V9_SPEC.md § 2. N2 —— 脚距惩罚(已完成权重预标定)
  • An auto-curriculum ratchet capped out at iter 248 and never engaged - stage difficulty manually or verify engagementauto-curriculum-engagement-check
    Observed oncewalkcurriculumcurriculumdomain-randomizationprocess

    Prefer manually staged difficulty with gated transitions; if you use an automatic curriculum, instrument its internal state and alarm when it stops engaging - a saturated curriculum is constant DR wearing a curriculum's name.

    Symptom

    A curriculum mechanism intended to grow difficulty adaptively (s1f's ratchet) hit its cap at iteration 248 and never bit again - for 96% of the run its effect was equivalent to constant DR, i.e. the curriculum existed in name only.

    Context

    When external advice suggested graded wz bands (start ±0.15, then ±0.30), the team agreed with grading but explicitly rejected automatic curriculum, citing the s1f episode. The same logic had already been paid for with push grading: ±0.6 failed twice, ±0.3 was feasible - grading matters, but the grade transitions were made by hand at verified checkpoints.

    Change

    Ladder policy: difficulty staged manually, one band per rung, each transition gated by the acceptance battery; automatic ratchets not used unless their engagement is monitored and demonstrated.

    Outcome

    Every C-ladder band change (wz ±0.12-0.25 first, wider later) was an explicit, attributable rung; no silent constant-DR-in-disguise runs recurred.

    Mechanism

    Adaptive curricula couple their own state machine to noisy training metrics; a ratchet that saturates early stops adapting but keeps its name, so the operator believes difficulty is progressing when it is frozen. Manual staging costs more decisions but each decision is observable and reversible.

    Applies when

    • choosing between auto-curriculum and staged bands for a new skill
    • a curriculum's difficulty parameter plateaus early in training
    • post-hoc attribution of what difficulty a lineage actually saw
    “C2 的 wz 分级(±0.15 → ±0.30,不要一上来 ±0.6)—— 与我们付过学费的 push 分级同形(±0.6 两轮 FAIL,±0.3 才可行)。但必须手动分级,不做自动课程 —— s1f 的课程化棘轮 iter 248 封顶未咬合,96% 时长等价常量 DR”
    train/C_LADDER_RUN.md § 1. 采纳 3 条 (C2 的 wz 分级)
  • Audit rewards by realized contribution (weight x achieved value) - a weight of 2.0 was really paying 0.04realized-contribution-audit
    Mechanism understoodwalkreward-shapingreward-shapingattribution

    Evaluate a reward table by each term's realized per-step contribution under the current policy, never by its weight column; if a term's realized value is ~0, escalating its weight is a no-op - change the term's structure instead.

    Symptom

    Foot dragging persisted through repeated weight escalation: feet_air_time had been raised 0.25 -> 1.0 -> 2.0 across versions with no behavioral change, and the training-side comment even recorded the fact ("几乎没有 单支撑相, 是拖着脚蹭") without the fix changing form.

    Context

    Computing realized per-term contributions in the trained state exposed the economy: feet_air_time contributed weight 2.0 x achieved 0.019 = 0.038 per step, against tracking's +1.20 - lifting the leg earned 3% of what tracking earned, so dragging was the rational optimum no matter the weight escalation. The same table acquitted the energy penalties (total -0.30 negative vs +1.74 positive) that a naive read of weights (-5.0 orientation!) would have blamed.

    Change

    Fix redirected from "raise the weight again" to "add a term whose realized contribution changes the optimum": a clearance penalty sized so its realized magnitude (~0.018/foot when dragging) is comparable to feet_air_time's, enough to flip the optimum without drowning tracking.

    Outcome

    With the term economy corrected (plus posture/range fixes), swing height reached 34 mm and tracking 87% by v6; weight escalation of the old term was abandoned.

    Mechanism

    A reward weight is only a multiplier on whatever the policy currently achieves on that term; when the achieved value is near zero (behavior absent), escalating the weight multiplies near-zero. Optimizer behavior is governed by realized per-step magnitudes, so audits must be conducted in that currency.

    Applies when

    • a behavior persists despite repeated weight increases
    • auditing whether penalties are "too strong" or rewards "too weak"
    • sizing a new reward term against existing ones
    “把 feet_air_time 权重从 0.25 → 1.0 → 2.0 一路加,但没有加高度项。量级算下来:feet_air_time 权重 2.0 × 实得 0.019 = 0.038,而跟踪奖励是 1.2。抬腿的边际收益只有跟踪的 3%,拖地当然是最优解。… 正项 +1.74,负项 −0.30。能量惩罚不是瓶颈,抬腿没收益才是。”
    train/WALK_DIAGNOSIS.md § 决定性证据(#1) / 各项奖励的实际量级
  • Training the final recipe from scratch in one run - every mechanism the lineage had accumulated - produced 0% and a seated robot; the order in which the lineage acquired those mechanisms was part of why it workedcurriculum-history-is-part-of-the-product
    Observed oncerecoverytraining-runcurriculumprocess

    A recipe that ends a lineage is not a recipe for a from-scratch run: consolidate it as ordered curriculum phases matching how the lineage acquired its mechanisms, check that every curriculum criterion is reachable from the starting policy, and read the run's raw term values, not the total reward, before calling it green.

    Symptom

    V3.0 trained the lineage's whole final recipe from scratch in one 9,000 iteration run - full beta curriculum, prone-conditioned pull assist, friction DR, the 3 s zero gate on standing income, flat_feet - testing the proposition "the product is defined by its configuration, not by its training history". Training looked all green (reward 26.33, episode length 500, 100% time-outs).

    Context

    Read in raw units against v2_6c at the same weights, the green was a seated equilibrium: base_height 0.377 vs 0.640, stand_pose 0.205 vs 0.516, flat_feet 0.0000 (zero because it sits outside its height gate, not because the feet were flat). Acceptance: 0.0% in Isaac at the deployed authority, 0.0% on every MuJoCo friction level, and still 0.0% at the training-time authority (100% seated at 0.222 m, upright and still).

    Change

    The full stdout (316k lines) was read: the beta curriculum's criterion (standing share over 0.35) was met zero times, so beta never left the wide setting and the policy had no experience at the deployed authority; the pull curriculum was stuck on the same criterion. The lineage had escaped the seated basin with immediate income (V2.0-V2.2) and only then added the zero gate to cure rushing (V2.5); from scratch, the zero gate removed the early "stand fast, earn more" gradient needed to escape. Verdict "the curriculum history is part of the product", limited to n = 1. v2_6c stayed the product.

    Outcome

    V3.1 kept the order as explicit phases: P1 from scratch with immediate income (zero gate off) until the curricula advance, P2 adding the zero gate. P1b/P1c escaped the seated basin and passed; P2 was later judged net negative and dropped (time-gate-vs-wide-stance-retire-the-fix).

    Mechanism

    Mechanisms that refine a competent policy (time gates, tight authority) can delete the gradient a naive policy needs, and a curriculum whose advancement criterion the naive policy never meets freezes at its first level.

    Conflicts

    The spec limits the falsification to "this recipe + this curriculum criterion" (n = 1, no seed sweep, no criterion tuning). V3.1's phased run succeeding is consistent with the ordering reading but changed other terms too.

    Applies when

    • consolidating a long lineage of continuation fixes into one clean recipe
    • a from-scratch run with all mechanisms enabled plateaus early
    • curriculum state is not logged or never advances
    “命题:产物由配置定义,而非训练史定义。 … 训练 log(完整 stdout 316k 行)`[beta_anchor]` 仅初始 1 行,**达标 0 次** … V2.0~V2.2 靠**即时计酬**爬出坐姿盆地(§33),站立巩固后 V2.5 才装归零门 治"过快"(§39)。**课程史是产品的一部分。**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §42 V3.0 判决(2026-08-11):从零单 run 全机制 FAIL 于坐姿盆地
  • Compare the achieved reward to the computed ignore-floor to tell "never learned" from "learned but unprofitable"ignore-floor-diagnosis
    Mechanism understoodomniattributionattributionreward-shaping

    For any skill that trains flat, compute the reward the null policy would earn on that term; achieved==floor means the behavior never paid out (find why: exploration, reward observability, or feasibility) - do not tune weights first.

    Symptom

    C4 sidewalk failed on both arms; the question was whether the policy had found sidewalk and rejected it as unprofitable, or never found it at all - two diagnoses with opposite fixes.

    Context

    The theoretical value of the sidewalk tracking term for a policy that completely ignores the command was computable from the command distribution: 0.189. Trained final values landed at 0.187 (arm A) and 0.204 (arm B) - sitting exactly on the ignore-floor - while the honest balance at the optimum actually favored sidewalking (net +0.88/step inside the side bucket). Later the same arithmetic closed the whole saga: for the true reward landscape, doing real sidewalk scored 0.153 vs 0.944 for ignoring - the policy's refusal "是理性最优,不是探索失败" (rational optimum, not exploration failure) under one hypothesis, and under the final measurement-corrected account the policy had "每一次都在 做理性选择" (made the rational choice every time).

    Change

    Diagnostic rule adopted: compute the ignore-floor for the new term; if the achieved value sits on it, the behavior was never expressed in useful volume (or the reward cannot distinguish it - check both); if the achieved value is above floor but the behavior is absent at deployment, the policy sampled it and priced it out - then the reward balance, not exploration, is the lever.

    Outcome

    Correctly identified that PPO had not merely under-valued sidewalk; each subsequent hypothesis (waveform sign, regularization cage, exploration form, reward kernel) was tested against this floor arithmetic, which kept the search honest through three reversals.

    Mechanism

    Every reward term has a computable value under the null behavior; the achieved-vs-floor gap is a one-number audit of whether the optimizer ever monetized the target behavior. It converts "training failed" into one of two mechanistically distinct states with different fixes.

    Applies when

    • a new skill's tracking reward plateaus early
    • deciding between exploration fixes and reward-weight fixes
    • post-mortem of a failed curriculum rung
    “track_lin_vel_y_exp 训练终值恰好坐在「完全无视指令」的底分上(A 0.187 / B 0.204,理论值 0.189),而终点 balance 明明有利(side 桶内净 +0.88/步)。不是学会了不划算,是根本没学到。”
    train/C_LADDER_RUN.md § 3e. 为什么首战 FAIL
  • A 10 s acceptance episode left 6-7 s of standing to observe - a narrow stance held for that window and split on hardware; a gate cannot see instability slower than its own horizonepisode-length-bounds-what-a-gate-sees
    Observed oncerecoverysim-evalgate-batteryreal-acceptance

    Size the standing phase of an acceptance episode, and its disturbances and floor friction, to what deployment will impose; a pass on a short static window certifies only that window.

    Symptom

    v2_6 passed every simulation gate (success 99.6%, re-falls 0-1%) and then, on the real robot, stood up and slid into the splits several times; prone starts stood and then fell backwards.

    Context

    Acceptance ran 10 s episodes; a ~1-2 s get-up left roughly 6-7 s of static standing on a nominal floor with no disturbance. The narrow stance's lateral margin and the straight-knee stance's lack of any flex buffer are both failure modes that need time, disturbance or lower friction to show.

    Change

    The gap was booked as a known blind spot of the gate ("long-duration standing stability") alongside the task-space stance criterion; later rungs added MuJoCo friction sweeps at mu 0.4 to every checkpoint scan.

    Outcome

    The spec through §50 records the blind spot but no longer standing window or disturbance row in the recovery acceptance itself.

    Mechanism

    An acceptance episode observes only the dynamics that unfold within its horizon under its conditions; slow drifts and disturbance-triggered failures are outside it by construction.

    Applies when

    • a policy passes sim gates and fails on hardware after a delay
    • acceptance episodes are short relative to deployment use
    • stability is judged without pushes or friction variation
    “**sim 门为什么没逮住**:10 s episode 起身后只站 ~6-7 s,静态窗口内窄站距 撑得住;真机站立时长/扰动谱在门口径之外 —— 长时站立稳定性记为口径缺口。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §43 判读:sim 门为什么没逮住
  • When training fails repeatedly, inject the target behavior open-loop - stop tuning rewards for an unverified behavioropen-loop-probe-before-reward-tuning
    Mechanism understoodomniattributionattributionprocesscurriculum

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop on the real plant/sim first, and only resume training once you hold a measured, safe, sign-verified target trajectory.

    Symptom

    Three sidewalk training rounds failed; hypotheses multiplied (exploration failure / wrong reference waveform / insufficient authority) with no way to pick between them by running more training.

    Context

    Instead of a fourth reward guess, the team wrote probe_side_ref.py: the candidate reference is injected open-loop on top of a frozen policy's output (bypassing PPO entirely), directly measuring "what happens if the robot literally does this waveform" - separating all three hypotheses in one experiment (5 seeds x 8 s per condition, several waveform families and gains). The probe immediately eliminated the authority hypothesis (full-amplitude execution, 5/5 survival) and localized the problem to the waveform/measurement side. The closing principle was written down after the saga: without a verified target behavior, tuning rewards is "在黑暗里试钥匙" (trying keys in the dark).

    Change

    Standing method: before opening another training rung on a failing skill, build an open-loop (or task-space) generator of the intended behavior, measure whether the physical system can express it and what it looks like - then train toward a verified, quantified target.

    Outcome

    The probe chain produced the verified waveform (reversed-sign triangle, half gain), quantified safe amplitude (tilt 8.2 deg at band top, foot distance clear of the wall), exposed the metric bug when probe and training disagreed, and ultimately supplied the feed-forward that made C4 pass in +100 iters.

    Mechanism

    Training couples exploration, reward design, and feasibility into one opaque outcome; open-loop injection cuts the loop and tests feasibility and waveform alone. A behavior demonstrated open-loop converts the remaining failure into a pure credit-assignment/reward question - and its measured trajectory becomes the reference itself.

    Applies when

    • repeated training failures on one skill with multiple live hypotheses
    • uncertainty whether the platform can physically express the behavior
    • a reference trajectory's shape/sign/amplitude is guessed, not measured
    “三轮 FAIL 之后不再猜,写 train/probe_side_ref.py 把参考开环注入到策略输出之上(绕过 PPO),直接量「照这个波形做会怎样」,一次分开三个假说:甲 探索 / 乙 波形 / 丙 权限。”
    train/C_LADDER_RUN.md § 3f. C4 真因定谳(开环探针) / 3k. 建议的下一步
  • Low-friction robustness traced to kd DR bandwidth, not friction training - by digging resolved params across 8 lineages, 3840 cellskd-bandwidth-mu-law-attribution
    Mechanism understoodomniattributionattributiondomain-randomizationprocess

    Attribute capability differences by tabulating every lineage's resolved training params and eliminating zero-variance and non-aligned columns first; never let an eval-side override knob serve as the explanation axis, and never write a mechanism into a law before it survives a targeted test.

    Symptom

    Lineages differed wildly in low-ground-friction survival, and the intuitive explanation - "some trained ground friction, some didn't" - was about to steer the ladder toward a ground-mu training rung.

    Context

    The attribution ran as a full parameter-vs-result cross: 8 lineages x 4 eval kd levels x 6 mu levels x 20 seeds = 3840 cells, with each lineage's RESOLVED training params dug out and compared item by item. First kill: all 8 lineages had ground mu pinned at (1.0,1.0) - zero variance - so low-mu differences cannot come from friction training at all. The only training parameter aligned with the mu score was kd DR bandwidth: narrow (<=0.24) lineages scored 19.9/19.5/19.5, wide (>=0.40) scored 17.1/15.2/14.6/14.2/12.8 - the two groups completely non-overlapping. Every rival was excluded item by item (kd center no; kp band no; COM small-beneficial non-driving; friction rung a clean double null 19.5->19.5 and 15.2->14.6; iteration count non-monotonic), and the one clean single-variable causal link confirmed it: the s2e-3 kd surgery (0.7,1.3)->(1.08,1.32) moved the score 17.1->19.5. Counter-proof against "each best at its own operating point": the narrow-band lineage evaluated OUT of band (18.2) still beat the wide-band lineage at its own band center (9.2). Two axes were ordered never to be conflated (the first attribution's own error): training kd bandwidth is a parameter axis / lineage property; the eval-side --kd-scale knob is a plant axis (more damping physically helps on slippery floors for ALL policies) - "plant 轴只能当部署缓解,不能当 归因". A tempting mechanism story ("drag vs step attractor") was tested and falsified, and explicitly kept OUT of the law: "机制未定, 不入定律".

    Change

    The planned ground-mu training rung was recommended closed ("建议 不开") in favor of a kd band-narrowing rung (0.8,1.2)->(0.9,1.1) centered on the deployed value - with a pre-registered risk that the law demands "bandwidth = measured dispersion" and the real robot's kd dispersion was not yet measured; if it exceeds +/-10%, narrowing sacrifices real coverage and the rung must yield.

    Outcome

    A whole training rung was deleted from the ladder by attribution alone (the second S2 pass dropped mu and push, 5 rungs -> 3); floor material became a deployment-selection input (mu <~0.6 -> deploy the kd1.2 gain profile) rather than a training target.

    Mechanism

    Cross-lineage performance differences must be attributed over the actual training-parameter table, not over eval knobs or plausible stories: eval knobs act on the plant for every policy (a physical effect), while lineage properties come only from training-time parameters. Zero-variance columns are free eliminations, and one clean single-variable rung is worth more than any correlation.

    Applies when

    • explaining why lineages differ on a robustness axis
    • an eval-side knob (gain scale, power) changes results and invites misattribution
    • deciding whether to open a DR rung for an axis never actually varied in training
    “8 血统地面 μ 训练带全部钉 (1.0,1.0) 零方差,低 μ 差异与「训没训地面摩擦」无关,是 kd DR 带宽的副产物 … 宽 ≤0.24 → 19.9/19.5/19.5;宽 ≥0.40 → 17.1/15.2/14.6/14.2/12.8, 两组完全不重叠。… 训练 kd 带宽 = 参数轴/血统属性;评测部署 --kd-scale = plant 轴 … plant 轴只能当部署缓解, 不能当归因。… 机制未定, 不入定律。”
    train/OMNI_V0_SPEC.md § 4. 地面 μ 鲁棒性 = kd DR 带宽的副产物 (2026-08-08)

Next page