Skip to content

Training Coach

The coach reads a training run and returns a diagnosis, in-whitelist proposals and an experiment plan — and every claim cites one of the cards below. This is that corpus: 22 methodology rules distilled from a real sim2real programme, and 171 episode cards behind them, each with the sentence in the war history it came from.

Doctrine

A report may cite any of these as doctrine-N.

  1. doctrine-1Contract freeze and fingerprint discipline

    The policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.

    Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).

    Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.

    contract-freeze-and-checkerlegacy-profile-pinningderived-asset-staleness-checkgain-profile-belongs-in-the-stamprecovery-two-policies-and-a-state-machinewalk-recovery-fsm-handoff

  2. doctrine-2Attribution by resolved training params - never eval-override knobs

    Capability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.

    Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).

    Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.

    kd-bandwidth-mu-law-attributiondeploy-knob-attribution-before-retrainingcycle-time-override-is-oodresolved-config-is-source-of-truth

  3. doctrine-3PASS gates become constraints; FAIL gates become objectives

    Once a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.

    Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.

    Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".

    fixed-acceptance-matrix-per-rungpreregistered-stop-criteria-per-rung

  4. doctrine-4One variable per ladder rung - counted against what the checkpoint saw

    A rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.

    Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.

    Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.

    resume-state-dr-auditorthogonal-batch-with-ablation-orderfill-the-missing-factorial-cell

  5. doctrine-5Pre-register risks, readings, and stop criteria before the ladder

    Before a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.

    Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).

    Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.

    preregister-risks-and-fork-readingspreregistered-stop-criteria-per-rungpreregistered-real-expectationsfeasibility-accounts-lock-design-point

  6. doctrine-6Plant parameters are measured, never invented

    Every plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.

    Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).

    Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.

    friction-measured-not-guessedarmature-n2-rotor-inertiatorque-limit-shape-by-measured-peakslatency-lerp-reverse-extrapolationfeasibility-accounts-lock-design-pointsim-api-friction-columnsget-up-feasibility-accounts-before-trainingsingle-support-gain-authority-probe

  7. doctrine-7Sim2sim gate before sim2real - under deployment conditions

    Every checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).

    Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).

    Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.

    sim2sim-gate-before-sim2realeval-plant-honesty-contact-paramspipeline-latency-is-plant-not-drbody-frame-velocity-api-audittorque-penalty-bought-by-leg-bracingtorque-disagreement-between-simulators-unresolved

  8. doctrine-8Observation honesty - the actor's inputs are a hardware contract

    The actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.

    Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).

    Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.

    observation-honesty-critic-onlyhistory-obs-needs-plant-variationreward-observability-limitdeploy-heading-loop-and-align-training

  9. doctrine-9Reward economics are audited in realized currency

    Reward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.

    Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).

    Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.

    realized-contribution-auditreward-cost-of-ignoring-auditgate-new-reward-terms-by-commandignore-floor-diagnosiscalibrate-threshold-between-healthy-and-sickcalibration-threshold-with-withdrawal-clauseinert-reward-term-auditseated-basin-dead-exp-kerneltail-torque-needs-hinge-on-computed-demand

  10. doctrine-10The zero-cost option must be the desired behavior

    For every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.

    Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).

    Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.

    penalize-the-slip-not-the-jointsaturation-cheating-zero-rate-costmoving-gate-42x-stand-taxtermination-closes-degenerate-basinpenalty-gate-is-an-escape-hatchsoft-limit-penalty-charges-nominal-poseunpriced-foot-attitude-is-a-free-variableenumerate-cheapest-cheats-before-trainingbinary-band-reward-fake-touchdown

  11. doctrine-11Measurement discipline: independent referees, signs, distributions

    A disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.

    Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).

    Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.

    independent-referee-for-metric-disputesbody-frame-velocity-api-auditsame-sign-response-is-yaw-biasmedian-hides-bimodal-distributionratio-metrics-need-absolute-checkheading-integral-not-body-ratesame-distribution-reward-comparisonmultiseed-sign-test-for-driftsingle-impulse-recovery-is-chaotic

  12. doctrine-12The deployment pipeline is plant

    Irreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.

    Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).

    Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.

    pipeline-latency-is-plant-not-drpower-scale-hurts-nonforward-axesdeploy-scaling-not-training-equivalentteleop-command-band-per-axislatency-dr-covers-measured-pipelinedeploy-rate-limiter-windupslew-anchor-is-an-integratorbeta-anchored-action-targetpower-derating-cuts-full-range-contract

  13. doctrine-13DR budget is finite; its distribution is the measured support

    Robustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.

    Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).

    Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.

    push-dr-conditional-budget-conservationdr-tail-plant-continuationconstant-value-dr-overfits-margintask-shaping-before-plant-hardeningcom-randomization-forces-leg-spreadcom-dr-rollback-on-symptomthin-dr-judged-by-channel-coveragefriction-priority-re-measured-after-plant-change

  14. doctrine-14Gates measure what hardware feels: posture, margins, stripped assists

    Acceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.

    Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).

    Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.

    task-metrics-vs-posture-metricsstand-gate-posture-not-survivalconstant-value-dr-overfits-marginsuspension-probe-removes-free-stabilizerhip-roll-sum-predicts-lateral-driftchirality-scored-separatelylow-speed-commands-reward-draggingend-state-confusion-matrixfrozen-acceptance-distribution-and-pinned-seedvideo-as-acceptance-recordepisode-length-bounds-what-a-gate-seesremoved-wall-returns-on-hardware

  15. doctrine-15Fork and root selection: recoverability, maturity, frozen rewards

    Choose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.

    Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).

    Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.

    fork-root-recoverable-shortfallroot-maturity-vs-product-qualityfine-tune-reward-change-falsified

  16. doctrine-16Curricula: verified engagement, lineage counters, disease-phase gating

    Automatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.

    Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).

    Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.

    auto-curriculum-engagement-checkcurriculum-counter-lineage-stepsgate-penalties-to-the-disease-phaseaggregate-metrics-mask-subgroup-failurebucket-share-is-not-a-gradient-levercurriculum-criterion-conditioned-on-lagging-categoryper-step-income-drives-speed-time-gatetime-gate-vs-wide-stance-retire-the-fix

  17. doctrine-17Probe before training: feasibility first, hypotheses in tables

    After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.

    Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).

    Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).

    open-loop-probe-before-reward-tuninghypothesis-table-code-auditperiod-doubling-evidence-racesuspended-test-isolates-actuator-authorityconfiguration-probe-wall-not-slopeprone-dead-end-is-foot-placementdof-vel-penalty-is-not-a-pacing-knobamplitude-cut-falsified-yaw-fix

  18. doctrine-18External advice is recomputed locally; values transfer as ratios

    Every external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.

    Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).

    Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.

    external-advice-audit-against-own-arithmetictransfer-ratios-not-absoluteslatency-dr-covers-measured-pipelinereference-structure-fk-amplitude-divisioncycle-average-tracking-for-gait-quantitiesadvisor-paraphrase-vs-paper

  19. doctrine-19Hardware sessions are scripted experiments, not tuning sessions

    Real-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.

    Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).

    Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.

    risk-ordered-real-deploymentbattery-bracketed-real-abknow-zero-command-behaviorpush-test-chirality-protocolno-field-tuning-protocolreversible-single-variable-field-experimentssim-veto-needs-real-confirmationfirst-real-get-up-violent-stage-one-policystaged-hang-mat-floor-for-get-uppower-cycle-preflighttwo-machine-config-disciplinefall-guard-becomes-a-statehardware-log-is-the-attribution-input

  20. doctrine-20Close questions in writing; restart when the debt is structural

    Audited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.

    Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).

    Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.

    frozen-verdicts-semantic-boundariesstale-verdicts-under-old-stackzero-offset-calibration-shifts-envelopeplant-swap-invariants-vs-shiftsfreeze-lineage-fix-structure-restartminimal-reward-table-with-provenancewrite-hardware-verdicts-back

  21. doctrine-21Name the quantity in the space it lives in

    A goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.

    Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).

    Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.

    joint-space-proxy-for-task-space-quantitycontact-detector-single-signal-lieszero-partial-credit-is-not-an-iteration-problemheading-integral-not-body-rate

  22. doctrine-22Continuation needs a live gradient; a release is chosen by a scan

    Continue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.

    Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).

    Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.

    converged-continuation-is-poisoncheckpoint-choice-is-a-full-gate-scanstance-decided-by-get-up-pathstop-stacking-roll-back-and-auditcurriculum-history-is-part-of-the-productcontinuation-budget-not-from-zeroroot-maturity-vs-product-quality

Experience cards

56 cards tagged reward-shaping.

  • action_rate weight is the sim2real bandwidth knob - re-tune it whenever a rate limiter is removedaction-rate-weight-vs-bandwidth
    Observed oncewalkreward-shapingaction-ratereward-shapingactuator-modeling

    Set action_rate weight relative to real actuator bandwidth, and re-tune it any time another smoothing/limiting element (filter, slew limiter, gain) changes - reward weights are load-bearing parts of the actuator model.

    Symptom

    With a low action_rate_l2 weight the policy learns fast actions; the unmodeled part of the actuator response is then excited hardest, and sim2real "直接崩" (collapses outright). With too high a weight, actions become so slow the robot cannot maintain balance.

    Context

    The reference developer called action_rate_l2 the single most important reward for transfer, with side-by-side video evidence that the high-penalty, slower policy is clearly better on hardware. Lucen context: the team had just removed the SOFT_SPD=1.0 velocity limiter, which had been an implicit actuator-bandwidth constraint - leaving action_rate as the only remaining constraint on action speed.

    Change

    Decision recorded: after removing SOFT_SPD, re-evaluate the action_rate weight rather than keep the old value, since its effective role changed from "additional smoother" to "sole bandwidth constraint".

    Outcome

    Logged as a priority follow-up ("重新评估 action_rate 权重 - 拆掉 SOFT_SPD 之后这一项的作用变了"); the failure mode it guards against is training high-frequency actions the real actuators cannot track.

    Mechanism

    Slower actions stay inside the frequency band where the ideal-PD sim actuator and the real actuator agree; fast actions probe the band where unmodeled delay, inductance, and bandwidth limits dominate, so model error is amplified in exact proportion to action speed. Any removed external rate limit transfers that constraint's entire job onto the action_rate penalty.

    Conflicts

    The low/high tradeoff evidence is the external developer's report (with video); the Lucen-side entry is a pre-registered risk and decision, not yet an on-robot A/B at the time of writing.

    Applies when

    • removing or adding an action filter, slew limiter, or low-level speed cap
    • real robot shows high-frequency chatter or overheating absent in sim
    • tuning smoothness rewards before a hardware deployment
    “权重低 → 动作快 → 执行器模型不准的部分被放大,sim2real 直接崩 / 权重高 → 动作慢 → 好迁移,但可能慢到无法维持平衡 … 我们刚拆掉 SOFT_SPD=1.0 的限速器,等于把执行器带宽约束整个移除了。action_rate 惩罚现在是唯一还在约束动作速率的东西,需要重新评估权重”
    Experience.md § action_rate_l2 是他认为最关键的 reward (lines 61-70)
  • Cutting swing amplitude 40% raised yaw-momentum demand 53% - the falsified fix is recorded so nobody walks that road againamplitude-cut-falsified-yaw-fix
    Mechanism understoodwalksim-evalmeasurementattributionreward-shaping

    Test gait fixes against the quantity the ground must actually supply (torque/force rates vs friction ceilings), not against kinematic proxies; record falsified fixes with their mechanism so the search space shrinks permanently.

    Symptom

    Support-foot yaw slip stayed at ~223 deg (vs 225 deg) after walk_v6 cut joint swing amplitudes by ~40% (hip_pitch 20.6->10.4 deg, knee 31.3->18.8 deg) - the change built on the theory "smaller swing = less yaw momentum to dump into the ground".

    Context

    Direct measurement inverted the theory: yaw-momentum amplitude ROSE 16% (+/-0.1000 -> +/-0.1159) and its rate of change rose 53% (1.86 -> 2.84 N*m demanded from the ground), pinned exactly at the foot's supply ceiling (2.8-3.3 N*m at mu 0.6-0.7) - so slip could not drop. The extra demand lives in higher harmonics: v6's crisper foot placement (better clearance 34 mm, lower landing force) shortens the momentum- exchange window, concentrating the same exchange into less time. The verdict was written as a closed road: "下一轮不要再走这个方向". An honest residue was also booked: WHY amplitude down but momentum up 16% remained unresolved, with the named next step (per-rigid-body decomposition of H_z, since per-joint RMS sensitivity ignores phase correlations).

    Change

    The "reduce amplitude to reduce yaw momentum" lever was removed from the planning space; future yaw-slip work redirected toward the supply side (friction) and momentum-rate mechanics.

    Outcome

    Slip unchanged (225 -> 223 deg); the falsification and its mechanism became a permanent constraint on the fix search space.

    Mechanism

    Ground yaw torque demand scales with the rate of change of angular momentum, not its amplitude; kinematic amplitude cuts that also sharpen contact timing can raise dH/dt while lowering range. When demand exceeds the friction-limited supply ceiling, slip is set by the ceiling, so demand-side changes below the ceiling do nothing visible.

    Applies when

    • attacking foot slip or yaw drift via gait shape changes
    • a fix targets an amplitude while the constraint is a rate
    • documenting a failed intervention after a version comparison
    “walk_v6 | ±0.1159 (+16%) | 2.50 Hz | 2.84 N·m (+53%) … 脚的供给上限 2.8~3.3 N·m(μ 0.6~0.7)—— v6 正好顶在天花板上, 所以滑移一点没降。… 结论: "减小摆动幅度以降低偏航动量"这条被 v6 证伪 —— 砍 39% 幅度, 需求反升 53%。下一轮不要再走这个方向。”
    train/WALK_DIAGNOSIS.md § ② 未生效的机理: 减小摆动幅度反而让偏航需求上升
  • Symmetrizing the config made the gait MORE asymmetric - the asymmetry lived in the policy weightsasymmetry-in-weights-not-config
    Mechanism understoodwalkattributionattributionreward-shapingcurriculum

    Localize a persistent asymmetry by intervening at the config layer first: if the symptom survives (or worsens), it is in the weights - fix it with symmetry-constrained training, not with trims or offsets.

    Symptom

    walk_v1 on hardware: straight-line command curved 149 deg in 15 s (9.9 deg/s) with 3.06 m lateral runout; turn gain +31% one way vs +129% the other (75% difference); knee asymmetry 4.4 deg in sim, 9.6 deg on the robot.

    Context

    The obvious suspect was the asymmetric default pose in the config. The decisive test: symmetrize standing_pose and run the SAME policy in sim - the asymmetry got LARGER (hip_pitch 6.8 -> 9.3 deg). Root cause therefore not in the config but baked into the policy weights: PPO without a symmetry constraint commonly converges one-sided, because splitting the work 50/50 and loading one side yield the same return, and the gradient falls randomly into one of the equivalent optima.

    Change

    Fix redirected from config trimming to retraining with mirror data augmentation (walk_v2 spec) - a weights-level fix for a weights-level disease.

    Outcome

    With augmentation (and the symmetric-default precondition), stand_v1 reached 0.0 deg asymmetry on all six joint pairs (from 4.4-7.7 deg), height fluctuation 7 mm -> 1 mm, mean |action| down 33%.

    Mechanism

    Reward-equivalent solution families (who carries the load) leave the symmetric solution unpreferred; SGD picks an arbitrary member and entrenches it. Config changes move the coordinate frame around the entrenched asymmetric function - they cannot move the function. The counterfactual test (change config, watch symptom) localizes the layer the disease lives in.

    Applies when

    • a robot veers or loads one side despite a symmetric-looking config
    • deciding between config trims and retraining for an asymmetry
    • mirrored-turn gains differ by tens of percent
    “根因不在配置里:把 standing_pose 对称化后在 sim 里跑同一策略,不对称反而变大(hip_pitch 6.8°→9.3°)—— 说明不对称烙在策略权重里。这是无对称约束的 PPO 的常见收敛结果(左右各担一半与一边多担的回报相同,梯度会随机落进其中一个)。”
    train/RETRAIN_v2.md § 1. 为什么是对称增强(证据)
  • A binary reward band on the swing knee had zero gradient everywhere below it, so the one-leg policy parked in an unloaded "fake touchdown" that Isaac's 5 N threshold scored as success and MuJoCo showed as real pressing - a capped constant-gradient ramp, retrained from scratch, passed 40/40binary-band-reward-fake-touchdown
    Mechanism understoodonelegreward-shapingreward-shapingsim2sim

    Shape approach-to-target rewards as capped ramps with gradient from the starting posture, never as bands or indicators; and compare contact-based terms across simulators, because a policy riding just under a force threshold looks perfect in one and wrong in the other.

    Symptom

    At iteration 1,000 of the first one-leg run the swing foot never lifted: the policy stood with the "raised" foot resting lightly on the ground. In Isaac the contact-match term paid 96% of full marks; the same policy in MuJoCo pressed that foot on the ground for 450 frames.

    Context

    The swing-leg goal was "shank folded fully back" (knee 1.5-1.95 rad), rewarded as a binary band: +0.8 inside [1.5, 1.95], zero elsewhere. From knee 0.05 to 1.5 rad the term was flat. Contact is judged at a 5 N force threshold, so a foot carrying less than 5 N counts as lifted. The walk line had hit the same disease with a binary indicator (v4) and fixed it with a capped ramp (knee_swing_amplitude).

    Change

    swing_knee_fold changed from the binary band to a ramp clamp(|q|/1.5, 0, 1) - a constant gradient capped near 86 deg - and the policy was retrained from scratch (V0r1). After the first real-robot try showed the fold still too low, its weight went 0.8 -> 2.0 (V0.1).

    Outcome

    V0r1 model_2300 passed the full acceptance 40/40 (swing knee 1.72 rad, about 98.5 deg) and was stamped as oneleg_v0.onnx; the cross-simulator disagreement is recorded as the thing that caught the cheat.

    Mechanism

    A reward that is flat until the target is reached gives no gradient to approach it, so the policy settles for the nearest state other terms reward - here, a foot that satisfies the contact threshold without lifting; a second simulator with different contact force resolution exposes such threshold-riding.

    Applies when

    • rewarding a posture target with an in-band / out-of-band indicator
    • a contact threshold decides whether a foot counts as lifted
    • trainer-side contact terms are near full marks while the video looks wrong
    “初版二值带 [1.5,1.95] 在膝 0.05→1.5 全程零梯度,策略停在"卸力虚点地"(Isaac 5N 阈下 contact_match 96% 满分 / MuJoCo 同策略 450 帧实压——跨仿真器互证抓作弊);v4 二值指示同型病,按 knee_swing_amplitude 判例改常数梯度封顶 ramp,从零重训 … **oneleg_v0.onnx = V0r1 model_2300, 40/40 PASS**”
    git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §4 奖励表 swing_knee_fold 行 / §8 核查单 5
  • Raising a command bucket's share does not strengthen its per-state gradient - it only starves the other modesbucket-share-is-not-a-gradient-lever
    Mechanism understoodomnicurriculumcurriculumreward-shapingdomain-randomization

    When a skill is not learning, first prove its per-state signal is nonzero (ignore-floor and probe checks); only rebalance sampling shares to fix genuine sample starvation, and account the regression risk to the diluted modes before doing it.

    Symptom

    Sidewalk was not learning, and the reflex proposal was to give the side bucket a larger share of sampled commands.

    Context

    The C4-redo3 rung explicitly kept the 20/40/20/20 bucket (stand/forward/turn/side) with the reasoning written out: PPO computes advantages per state, so bucket proportion does not change the per-state gradient of side states; at 4096 envs x 20% x 24 steps the rollout already contained ~19.7k sidewalk states - sample count was not the bottleneck. And the cost side was already measured: cutting forward from 60% to 40% had made vx+0.30 die at +400 in an earlier run - more cuts would only collapse it sooner.

    Change

    Bucket proportions held constant across the entire C4 redo series; the actual bottlenecks (metric frame bug, reward variance penalty, exploration form) were pursued instead.

    Outcome

    Sidewalk was eventually fixed with zero bucket changes (feed-forward delivery, +100 iters); forward/turn skills never suffered starvation-induced regressions during the redo series.

    Mechanism

    Policy-gradient credit is assigned per visited state; oversampling a mode multiplies its states in the batch but not the informativeness of each, so if the per-state gradient is ~0 (behavior unreachable or reward indifferent), N times zero is still zero - while the displaced modes genuinely lose data and regress.

    Applies when

    • proposing to oversample a failing task/command mode
    • a majority mode regresses after share rebalancing
    • budgeting env count vs mode share for a multi-skill policy
    “比例不动:PPO 逐状态算优势,桶占比不改变单状态梯度;4096 env × 20% × 24 = 每 rollout 已有 1.97 万个侧走状态,样本数不是瓶颈;而 forward 60%→40% 已实测让 f30 在 +400 处死掉,再加码只会更早塌。”
    train/C_LADDER_RUN.md § 3i. 桶 20/40/20/20 不动(比例不动)
  • Calibrate a wall penalty by measuring healthy and sick policies - healthy pays ~0, the disease pays a wallcalibrate-threshold-between-healthy-and-sick
    Mechanism understoodwalkreward-shapingreward-shapinggate-battery

    Calibrate every threshold penalty by evaluating its exact formula on replays of at least one healthy and one sick policy: place the threshold between their distributions, size the weight so the sick policy pays a decisive fraction of tracking while the healthy one pays ~0, and pre-compute neighboring thresholds for cheap adjustment.

    Symptom

    Real walk_v7 occasionally clipped its own legs (stance narrowed to 133 mm mean vs nominal 214.5); a foot-distance penalty was needed, but an uncalibrated threshold/weight risked either doing nothing or becoming a reverse barrier.

    Context

    The term (relu(d_min - lateral foot distance), measured in the base yaw frame because world-frame y is meaningless after turning) was calibrated before training by replaying three known policies through the exact reward formula at the acceptance operating point: healthy v5 (183 mm) pays 0.6% of tracking - effectively free; narrowed v7 (133 mm) pays 22% - effective widening pressure; collapsed v8 (93 mm) pays 55% - a wall. d_min 0.16 was placed deliberately between healthy and sick, with alternative thresholds (0.14/0.18) pre-computed in the tool output for later adjustment. The shape self-check was named as a standing question: "先问'零代价的选项是什么'" - the zero-cost region must be exactly the desired behavior.

    Change

    feet_lateral_distance added at d_min 0.16 / weight -10, with telemetry expectation pre-registered (should decay toward 0 as stance learns >160 mm; if bow-legged over-widening >214 appears, only then discuss an upper bound).

    Outcome

    v9_probe on hardware: no leg contact ("没碰腿(N2 兑现)"), stance min 145/126 mm green; the term's zero-cost design left healthy gait untaxed.

    Mechanism

    A relu threshold penalty defines a free region and a priced region; its correctness is entirely in where the boundary sits relative to the healthy and pathological distributions. Replaying known-good and known-bad policies through the exact formula measures both distributions in the term's own currency, making the threshold and weight a placement decision instead of a guess.

    Applies when

    • adding any relu/threshold-style wall penalty
    • a safety margin (foot distance, joint limit, clearance) needs enforcement without taxing normal behavior
    • choosing between candidate thresholds for a new term
    “形状自检 (v6 横向组/v8-B 的教训 —— 先问"零代价的选项是什么"): 标称站距付 0, 健康步态付 ~0, 收窄才付费 … v5(健康) 183 mm … 0.6% ≈ 免费 | v7(收窄) 133 mm … 22% —— 有效推宽 | v8(塌陷) 93 mm … 55% —— 墙 … d_min=0.16 恰在 v5(183)与 v7(133)之间”
    train/WALK_V9_SPEC.md § 2. N2 —— 脚距惩罚(已完成权重预标定)
  • Every new penalty ships with a pre-registered withdrawal clause - if healthy gait must pay above the cap, the term stands downcalibration-threshold-with-withdrawal-clause
    Replicatedwalkreward-shapingreward-shapingprocess

    Introduce every new penalty with: the zero-cost-option audit, a replay-calibrated weight formula (healthy pays a fixed small fraction of tracking), and a pre-registered withdrawal condition - and let the clause fire without argument when the calibration says the term cannot be afforded.

    Symptom

    Three same-shaped crashes had established a failure archetype: v4's clearance, v8a's landing window (weight off by 58x uncalibrated), and v6a's bare hip_yaw suppression all combined a zero-cost "don't move" option with a fee on any motion - a reverse barrier that pushes policies toward standing still.

    Context

    The v11 landing-window penalty was therefore introduced under a calibration-threshold protocol: (1) shape chosen with the window tightened (h_gate 0.03 -> 0.02, because 0.03 equaled the clearance target and priced the entire descent); (2) weight from a FORMULA, not judgment: measure the term's raw value on healthy replays (v5/v10b), set w = -(0.10-0.15 x tracking reward) / raw_healthy; (3) withdrawal clause pre-registered: if healthy gait must pay >15% of tracking no matter the tuning, the term is withdrawn to the next version rather than forced in - "不硬上". The companion hip_yaw quieting term ran the same protocol (calibrate on replays, healthy pays <=5%) and was later retired entirely when a structural fix (zero action scale) made its shaping tax unnecessary.

    Change

    Penalty introduction protocol: shape audit (what is the zero-cost option?), replay-based weight formula, healthy-pay cap with a written stand-down condition - all before training.

    Outcome

    The landing term was in fact withdrawn under its clause (v12 records "P5 落地窗口罚 已撤 … 维持撤下"), demonstrating the protocol firing as designed instead of the fourth same-type crash.

    Mechanism

    A penalty's damage mode is mispricing healthy behavior; since the healthy price is measurable in advance on replays, both the weight and the go/no-go decision can be computed rather than discovered by a ruined training run. The withdrawal clause converts "make it work" pressure into a clean deferral.

    Applies when

    • adding any motion-taxing penalty to a working gait
    • a proposed term's weight has no measurement behind it
    • a previous same-shaped term crashed training
    “权重公式而非拍脑袋:先在 v5/v10b 回放上量 h_gate=0.02 的原始值,w = −(0.10~0.15 × 跟踪奖励) / raw_健康;标定门槛:若健康步态无论如何要付 >15% 跟踪,本项撤下留 v12,不硬上 (v4 clearance/v8a-B/v6a 三次同型翻车的教训:代价为零的"不动"选项 + 一动就收费 = 反向壁垒)。”
    train/WALK_V11_SPEC.md § 6. P5 —— 落地窗口罚(三代欠账,标定门槛制)
  • Oversized lateral COM randomization (+/-5 cm) deliberately forces leg spreadcom-randomization-forces-leg-spread
    Observed oncewalkdr-tuningdomain-randomizationreward-shaping

    DR ranges can be behavior-shaping tools, not just robustness padding: oversize a randomization axis to force a strategy the reward struggles to express - and expect a compensating behavior to appear as the cost.

    Symptom

    Feet drift toward the centerline and even collide; policy has no incentive to keep a lateral support base.

    Context

    COM randomization ranges were chosen asymmetrically by axis: lateral +/-5 cm ("比常规大,故意的" - larger than usual, on purpose), fore-aft +/-2 cm, vertical +/-2 cm. The oversized lateral range is not robustness padding but a behavioral forcing function. Lucen logged it as directly relevant to its own roll-channel / sideways leg-kick symptom.

    Change

    Set COM randomization to lateral +/-5 cm, fore-aft +/-2 cm, vertical +/-2 cm, with the lateral band intentionally oversized to make narrow stances fail during training.

    Outcome

    Effective at separating the feet on the reference robot; side effect - the base began swaying left-right, which then required a foot-centerline distance penalty (see reward-chain-foot-height-landing-spacing).

    Mechanism

    Randomizing COM laterally makes narrow-stance policies fall for some draws, so PPO discovers wide stances as the only strategy robust across the band - DR used as an implicit reward. The sway side effect appears because the policy hedges against unknown COM by active lateral correction.

    Applies when

    • feet too close / self-collision in a learned gait
    • roll-axis instability suspected to come from narrow stance
    • choosing COM or mass-offset DR ranges
    “两脚太近甚至互撞 → 先试质心横向随机化 ±5 cm,逼迫策略把脚分开;有效但引发新问题——基座开始左右摇摆 … 横向 ±5 cm(比常规大,故意的,用来逼出分腿)/ 前后 ±2 cm / 垂直 ±2 cm”
    Experience.md § 质心随机化范围 (lines 75, 84-86)
  • An exponential kernel on instantaneous velocity punishes gait oscillation - track the cycle averagecycle-average-tracking-for-gait-quantities
    Mechanism understoodomnireward-shapingreward-shapingattribution

    Reward velocity tracking on gait-cycle averages (or filtered values), not instantaneous samples, whenever the desired behavior oscillates at stride frequency; widening the kernel does not fix a variance penalty.

    Symptom

    Even while the robot genuinely sidewalked (verified after the metric fix), the Isaac-side tracking reward sat on the ignore-floor: true sidewalk scored 0.178 vs 0.189 for ignoring the command - the reward was mildly punishing the desired behavior.

    Context

    Sidewalking is inherently oscillatory: per-frame vy std was 0.177 while the tracking kernel width was sigma = 0.15, applied to the instantaneous value. A kernel-width scan showed widening sigma 0.15 -> 0.50 still loses (-0.14 -> -0.07): "指数核惩罚的是方差,而侧步天生带方差" (the exponential kernel penalizes variance, and side-stepping inherently carries variance). Modeling with measured parameters: replacing instantaneous vy with the mean over one gait cycle (0.5 s) flips the margin decisively (true sidewalk 1.888 vs ignore 1.281, +0.607), half-cycle is neutral (+0.006), two cycles adds nothing more. Explicitly flagged as extrapolation pending Isaac-side implementation. This also vindicated a previously dismissed external note (sigma too small) - right conclusion, different mechanism than claimed (variance, not gradient).

    Change

    Proposed fix recorded: change the tracked quantity from instantaneous vy to a one-gait-cycle running average; widening sigma alone rejected by the scan.

    Outcome

    Diagnosis complete and quantified; the C4 product shipped via feed-forward before the reward change was implemented, so the cycle-average fix remained a verified-by-model, not-yet-trained change.

    Mechanism

    E[exp(-(v-c)^2/sigma^2)] decreases with Var(v) even when E[v] = c exactly; a gait's phase-locked oscillation guarantees variance at the stride frequency, so instantaneous tracking rewards structurally prefer standing still at the command mean. Averaging over exactly one cycle removes stride-frequency variance while preserving command-following error.

    Conflicts

    The cycle-average fix itself is model-extrapolated ("⚠️ 这一条是外推,须在 Isaac 侧实装并复量后才能当结论") - the diagnosis is measured, the remedy untested in training at the time of writing.

    Applies when

    • tracking rewards for lateral/turn/any oscillation-carrying velocity
    • a verified behavior scores below the ignore-floor
    • choosing sigma for exp-kernel tracking terms
    “侧走时 vy 的逐帧摆幅 std = 0.177,而 track_lin_vel_y_exp 核宽 σ = 0.15,且作用在瞬时值上 … 真侧走(均值 66%,振荡 ±0.18)0.178 | 完全无视指令 0.189 … 真侧走的得分比无视指令还低。… σ 从 0.15 放到 0.50,侧走仍然吃亏 … 把跟踪目标从瞬时 vy 换成一个步态周期(0.5 s)的平均 vy:… 1.888 vs 1.281”
    train/C_LADDER_RUN.md § 3n. 二/三 Isaac 训练奖励为何一直坐在「无视底分」/ 修法不是放宽 σ
  • Changing the gait clock silently flipped a hardwired threshold's meaning - write derived constants as expressionsderived-constants-must-track-their-base
    Mechanism understoodwalkprocessprocessreward-shapingcontract-freeze

    Before changing any base parameter (clock, control rate, scale), enumerate every constant derived from it and every constant that must NOT change; convert derived literals into expressions of the base so the next change cannot silently flip a term's meaning.

    Symptom

    Slowing the clock 0.40 -> 0.50 s would have silently inverted the feet_air_time threshold's semantics: the 0.25 s threshold was hardwired, so at ct 0.40 the swing window (~0.20 s) sat below it (constant pressure to lengthen strides), while at ct 0.50 the window (~0.25 s) equals it - the term's meaning flips from "push longer" to "neutral" with no code error anywhere.

    Context

    The clock change audit walked every dependent quantity: most followed automatically (joint_pos_ref / clearance / contact_number cycle_time params, gait_phase observation, deploy/sim2sim/policy_io, export) - wiring confirmed, zero hand edits; the air_time threshold was the one hardwired constant, fixed by preserving the RATIO: 0.25 -> 0.3125 = 0.625 x ct, with the recommendation to commit it as the expression 0.625*ct "一劳永逸" (solved once and forever). The same audit also listed what must NOT follow the clock (50 Hz control rate, physics dt/decimation, 47-dim contract, action_latency absolute seconds, PD/torque limits) - the change's blast radius stated in both directions.

    Change

    feet_air_time threshold re-expressed as a fraction of cycle_time; auto-following vs must-not-change lists written into the spec for the clock migration.

    Outcome

    The clock migration (v10, repeated in v11) carried no silent semantic flips; the expression form removed the trap for every future clock change.

    Mechanism

    Constants derived from a base parameter encode a ratio at their birth; storing the evaluated number severs the dependency, so changing the base leaves stale semantics with no failing test. Expressions preserve the intent; and an explicit both-directions dependency list (follows / must-not-follow) is what makes a base-parameter change reviewable.

    Applies when

    • changing gait clock, control frequency, or units
    • a reward threshold interacts with a phase/window duration
    • config audit finds literals that encode ratios
    “feet_air_time 阈值 0.25 是写死的,不跟 ct 走——0.40 时摆动窗 ~0.20s<0.25(恒拉长压力),0.50 时摆动窗 ~0.25s≈阈值(语义翻转)。按比例保原压力:0.25 → 0.3125(=0.625×ct;建议直接写成 0.625 * ct 表达式,一劳永逸)。”
    train/WALK_V10_SPEC.md § 3. T —— 慢时钟 (训练侧必做一件)
  • A joint-velocity penalty meant to slow the get-up cut joint speed 16% and left the get-up time unchanged - the knob never moved the variable, so the idea it was meant to test stayed untesteddof-vel-penalty-is-not-a-pacing-knob
    Mechanism understoodrecoveryreward-shapingreward-shapingattribution

    Before reading a result as a test of an idea, check that the knob actually moved the independent variable; velocity regularizers smooth a schedule they do not set, and a schedule driven by per-step task income moves only when that income's time structure does.

    Symptom

    The get-up took 0.6-0.9 s in Isaac with large torque demand; the user proposed getting up more slowly so less torque would be needed.

    Context

    The idea had support in the accounts: the acceptance bound is an upper bound of 5 s (5-8x margin), the quasi-static squat path peaks at 25% of the limits, and rolling over needs no momentum. R3.2 raised dof_vel from -1e-3 to -5e-3 as the single variable.

    Change

    dof_vel -1e-3 -> -5e-3 (child-run from R3.1).

    Outcome

    Get-up medians moved +0.02-0.04 s (noise); raw joint velocity -16%; torque demand median got worse (hip_pitch 46-48% -> 63-67%) as the new term competed with torque_headroom on the same joints; MuJoCo 98 -> 96%. Verdict FAIL on the knob, not on the idea, and the rung was not adopted. When pace was later attacked through the income's time structure (V2.5/V2.5b), the MuJoCo get-up moved into the 3.5-4.5 s design band.

    Mechanism

    The pace was set by base_height_progress paying for every step spent high (stand earlier, earn more); a velocity regularizer only smooths motion along the same schedule and does not change when the robot stands up.

    Applies when

    • trying to make a skill slower or gentler with smoothness penalties
    • an experiment's primary metric did not move and a verdict is being written
    • two penalties act on the same joints
    “**关键判读:`dof_vel` 罚只把关节速度压了 16%,而起身用时一点没变。** 也就是说**这一级根本没有把"慢下来"这个自变量推动起来** —— 所以它**不构成对 用户假说的检验** … 起身节奏由 `base_height_progress` 的逐步计酬决定(早站起来就多 拿),速度正则只在同一条时间轨迹上把动作抹匀,不改变何时站起来。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §25 R3.2(dof_vel −1e-3→−5e-3,慢一点起身)
  • Before training a one-leg stand the spec named the cheapest cheats - hopping on the support foot, a raised foot resting unloaded, a leg tripod - and gave each a countermeasure and a gate; one still appeared and was caught by exactly those gatesenumerate-cheapest-cheats-before-training
    Observed onceonelegreward-shapingreward-shapinggate-battery

    Before training, list the cheapest behaviours that would satisfy each reward term without doing the task, give each a countermeasure in the reward and a gate in acceptance, and prove the intended behaviour is reachable with a probe - then treat any gate the policy games as evidence about the reward, not the gate.

    Symptom

    The literature's single-leg benchmark reports eight state-of-the-art general policies holding a clean one-leg stand 0 times out of 90 - they survive by sneaking steps and hops - so the task's first adversary was the policy's own cheating.

    Context

    The spec's shape self-check ("what is the zero-cost option?") listed, for the one-foot bucket: the cheapest cheat, a foot resting on the ground without load, countered by a 5 N contact threshold plus positive swing income; the second cheapest, small hops on the support foot to reset balance, countered by a continuous support-air penalty plus a gate of zero support-foot flight segments. The probe that preceded training had already seen a third: early low-lift postures "survived" by pressing the swing foot at 78-95 N, a leg tripod, removed by folding the shank back. The two-foot bucket was checked too: its zero-cost behaviour is ordinary standing, with no odd base state.

    Change

    Countermeasures and gates written before training: swing-contact and support-air penalties, gate 2 (zero swing-foot contact frames above 5 N), gate 3 (zero support-foot flight segments).

    Outcome

    The first run still found the unloaded-foot cheat (a binary reward band gave it no gradient to lift) - and it was caught, by the contact gates and the cross-simulator comparison, not discovered on hardware. The retrained V0 passed all gates 40/40, including zero support-foot flight after the flight detector was corrected.

    Mechanism

    A policy optimizes the reward, not the intent; the cheapest behaviours that satisfy the reward are predictable from the reward's structure, and a gate written for each before training turns a silent cheat into a failed row.

    Applies when

    • designing rewards for balance, contact or "hold still" tasks
    • benchmark policies are known to cheat the task
    • writing acceptance gates for a new skill
    “文献里 8 个 SOTA 通用策略在单脚站基准上 0/90 干净保持, 全靠偷步偷跳活命,这是本任务的第一反作弊对象 … 单脚桶下最便宜的作弊是"脚虚放地上不受力"——接触判定 >5 N 力阈(沿用),配 swing_height_band 正收入拉开。 … 第二便宜是"支撑脚小跳重置"——support_air_penalty 连续罚 + 验收门支撑脚腾空段=0 双保险。”
    git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §0 目标口径 / §5 形状自检(零成本选项是什么)
  • Exponential tracking kernels go flat exactly when the error is largest - pair them with an L2 term for the far fieldexp-kernel-needs-l2-far-field
    Mechanism understoodwalkreward-shapingreward-shaping

    Never let an exp/Gaussian kernel be the only tracking pressure on a quantity that can drift far from target: pair it with an unbounded (L2) term sized as the "don't diverge" floor, and check which frame the kernel reads.

    Symptom

    With only an exp-type yaw tracking term (exp(-err/std^2), std 0.25), a robot whose heading had drifted badly received almost no corrective gradient: at error 0.6 rad/s the term evaluates to exp(-0.36/0.0625) = 0.003 - near zero AND flat.

    Context

    The exp kernel is excellent for fine tracking near zero error but its gradient vanishes at large error - precisely when correction matters most. Fix: add track_ang_vel_z_err_l2 (-0.5), a plain quadratic on the same quantity: "exp 管精细跟踪、L2 管'别发散', 互补". Both terms deliberately read WORLD-frame wz (matching the exp term's source), because this torso sways enough that body-frame wz means are systematically off (measured -0.039 while actually turning +0.152). The same far-field-gradient argument reappears in the v8 risk list: frozen joints could not climb back because their huge error put them on the exp plateau ("远端梯度消失是冻结自锁的帮凶").

    Change

    Added the L2 companion term at -0.5 alongside the existing exp term (a term that had been in an earlier draft and was lost in a rewrite - itself worth noticing).

    Outcome

    Corrective pressure restored across the whole error range; the exp+L2 pairing became the house pattern for tracking terms.

    Mechanism

    d/de[exp(-e^2/s^2)] -> 0 as e grows: the kernel saturates and cannot distinguish bad from terrible. A quadratic's gradient grows with error, covering the far field; summing the two yields monotone corrective pressure with fine shaping near the target.

    Applies when

    • tracking rewards use exp/Gaussian kernels alone
    • a drifted or frozen state fails to recover during training
    • designing tracking terms for quantities with large transient errors
    “exp 在误差大时梯度趋零, 恰好在最需要纠正的时候失灵。… 误差 0.6 → exp(-0.36/0.0625) = 0.003, 接近零且平坦。… exp 管精细跟踪、L2 管"别发散", 互补。”
    train/WALK_V7_SPEC.md § ② track_ang_vel_z_err_l2 −0.5 —— 补 exp 的梯度洞
  • Sort external training advice into adopt / already-have / modify / would-trap by recomputing it on your own configexternal-advice-audit-against-own-arithmetic
    Replicatedomniprocessprocessattributionreward-shaping

    Never apply external tuning advice directly: recompute each claim on your own reward table and probe data, classify it adopt / have / modify / trap, and record why - and verify external citations actually exist.

    Symptom

    External AI/literature advice for the omni ladder arrived plausible-sounding but was written without knowledge of this robot's actual reward table, contract, and history; following it blindly would have broken single-variable discipline and, in one case, made sidewalk unlearnable.

    Context

    Before the C ladder, every external suggestion was audited: 3 adopted (ellipsoid command sampling; staged wz bands; command-switch acceptance), 3 already present (unified reward; frame history - the frozen 168-dim 5-frame window; per-100-iter acceptance), 2 modified (stand share kept at 20% to avoid a second variable; back share NOT raised because the probe showed backward works untrained 20/20@67%, so oversampling would only crowd out forward), and 1 flagged as a trap: "start vy very small (0.06-0.15)" - on THIS reward table vy was only an L2 tax, so ignoring a vy=0.06 command costs 0.4% of the vx tracking scale, 28-180x cheaper than ignoring forward, with quadratic shrinkage making small commands weaker still. A separate retrieval-reliability note: two search agents returned fabricated verbatim quotes from arXiv PDFs (2 papers, verified fake and discarded); only HTML/abstract/source-verifiable material was used.

    Change

    Advice classified only after recomputing each claim with local numbers; the "start small" trap was replaced by adding a gated lateral tracking term (the ladder's only true reward surgery) instead of shrinking the command.

    Outcome

    The adopted items (ellipsoid modes, staged wz, transition acceptance) entered the ladder; the trap was avoided; one external factual error (calling s1g the mainline start - it was falsified 0/20) was caught. Later, one initially-dismissed item (sigma=0.15 too narrow) turned out right for a different reason than claimed - see cycle-average-tracking-for-gait-quantities.

    Mechanism

    External advice encodes the advisor's reward table and robot, not yours; the transfer-validity test is whether the claim survives recomputation under your own arithmetic (reward margins, probe baselines, contract freeze). Items that survive become experiments; items that don't become documented traps.

    Applies when

    • incorporating LLM or literature advice into a training plan
    • advice conflicts with locally measured baselines
    • an external claim depends on reward-table details the advisor cannot know
    “其建议 C4「先很小,vy = ±0.06~0.15」—— 在我们这张奖励表下会让侧走学不起来 … 忽略侧走比忽略前进便宜 28~180 倍,且指令越小激励越弱(平方缩放)—— "先很小"在稳定性上对、在梯度上正好把信号缩没了。”
    train/C_LADDER_RUN.md § 1. 外部 AI 训练建议的评估(采纳 / 已有 / 要改 / 会踩坑)
  • PPO's Gaussian noise cannot compose phase-locked oscillations - deliver them as feed-forward and let the policy learn the residualfeedforward-for-phase-locked-skills
    Mechanism understoodomnicurriculumcurriculumreward-shapingcontract-freeze

    If a skill needs a temporally coherent (phase-locked) action component, do not expect step-wise exploration to find it: inject a verified feed-forward and train the policy as a residual stabilizer, keeping the feed-forward inside the deployment contract.

    Symptom

    Four different reward arrangements (no reference / wrong-sign reference / correct-sign reference / cage released) all failed to elicit sidewalk, while open-loop probes proved the behavior existed and was safe on the same platform with the same policy as base.

    Context

    Producing lateral velocity requires a phase-locked hip_roll oscillation synchronized to the gait clock. PPO's exploration is per-step, zero-mean, uncorrelated Gaussian noise - it can never compose a sustained phase-locked component, so the behavior is unreachable by exploration regardless of how it is rewarded. The fix changed the delivery channel: target = default + scale*action + lat_ff(cmd_vy, phi). The policy's action becomes a residual on top of the feed-forward, retaining full balance authority (it can even cancel the feed-forward); the feed-forward supplies exactly the component exploration cannot. This mirrors why the sagittal joint_pos_ref worked (it also delivered phase structure), just via a different channel.

    Change

    Contract-level change, done cleanly: new profile omni_ff (= omni + lat_ff_gain -0.5), existing omni profile bit-identical; feed-forward applied after the action delay stage; missing cmd/phase raises instead of silently dropping; deployment must use the same phi as build_obs (recomputing gives a one-tick phase misalignment).

    Outcome

    From C2-700, +100 iterations sufficed: product omni_c4_ff800 scored vy +120%/+125% (from +4%/-1%), 260/260 cells at 20/20 survival, zero old-skill regression, left/right gap 5 pp - the entire C4 saga resolved by changing the delivery mechanism, not the reward.

    Mechanism

    Exploration noise spans only the subspace its correlation structure can express; skills requiring coherent oscillation lie outside the span of i.i.d. per-step noise. Feed-forward moves the required structure into the action pipeline where it needs zero probability mass to appear, reducing the learning problem to stabilizing around a demonstrated behavior - which PPO does well.

    Applies when

    • a periodic/oscillatory skill trains flat under every reward variant
    • open-loop injection of the behavior already works
    • considering GRU/curriculum/exploration tricks for a rhythmic skill
    “病因不在奖励,在探索形式:产生侧向速度需要相位锁定的 hip_roll 振荡,PPO 的逐步高斯噪声零均值无相关,合不出相位锁定分量。… target = default + scale·a + lat_ff(cmd_vy, φ)。策略动作因此是前馈之上的残差,保留全部平衡权限”
    train/C_LADDER_RUN.md § 3j. C4-redo4:唯一变量 = 侧步参考改为前馈注入(契约级)
  • Design the next run to complete the 2x2 - either outcome then convicts or acquits a factor cleanlyfill-the-missing-factorial-cell
    Mechanism understoodwalkprocessprocessattributionreward-shaping

    When two config factors are jointly suspected, lay out the factorial of existing evidence, spend one run on the missing cell with both interpretations and an early-abort tripwire written in advance - and treat either outcome as a verdict, not a disappointment.

    Symptom

    Hip joints froze at the action clamp in v7, but the history could not say whether the culprit was the raised action_rate (-0.2) or the halved reference amplitude (scale 0.15): existing versions covered only three corners of the (rate x scale) space - v5 (-0.03, 0.30) healthy, v6 (-0.03, 0.15) healthy, v7 (-0.2, 0.15) frozen.

    Context

    v9 was designed explicitly as the missing cell (-0.2, 0.30), with the readings pre-registered: v9 not frozen -> the real anti-freeze force was always the reference amplitude and -0.2 may stay; v9 frozen -> -0.2 is convicted beyond appeal (freezes at both amplitudes) and the next version goes straight to a structural fix. "两个结局都是干净的信息" - both endings are clean information.

    Change

    One training run allocated purely to complete the factorial, with freeze tripwires (joint_pos_ref telemetry <0.1 at iter 1000-1500 -> abort, do not run to 6000) so a conviction costs the minimum compute.

    Outcome

    v9 froze - the rate weight was convicted at both amplitudes ("−0.2 铁案定罪"), and v10 moved to the structural saturation fix with the weight question closed instead of re-litigated.

    Mechanism

    Three corners of a 2x2 leave the two factors confounded in the failure corner; the fourth observation makes each factor's marginal effect identifiable. Pre-registering both readings turns the run into a guaranteed-informative experiment regardless of outcome.

    Applies when

    • two config changes are confounded in a failure
    • version history already covers some corners of a factor grid
    • deciding what single experiment buys the most attribution
    “这恰好补齐一个 2×2 实验矩阵的缺格 … v9 不冻 → 真正的抗冻结主力一直是参考摆幅,−0.2 可以留;v9 仍冻 → −0.2 铁案定罪(两种摆幅下都冻),v10 直接上结构修复 … 两个结局都是干净的信息。”
    train/WALK_V9_SPEC.md § 0. 设计原则 (2×2 实验矩阵)
  • Fine-tuning through a reward-table change was falsified (scatter, half-recover, collapse) - continuation training is legal only with the reward frozenfine-tune-reward-change-falsified
    Replicatedomnicurriculumcurriculumfork-selectionreward-shapingprocess

    Never fine-tune through a reward-table change - retrain from zero; reserve checkpoint continuation for frozen-reward plant/DR widening, reset noise_std when branching, and watch for the scatter/half-recover/collapse signature as the abort trigger.

    Symptom

    The s1c A/B experiment: arm B fine-tuned from an existing checkpoint under the revised reward (same contract, same network, changed reward + small DR) and failed with a characteristic signature - scatter, half-recover, fall back ("打散→半恢复→摔回"); arm A trained from zero under the same config won decisively (full shaping lifted swing to 21.6 mm within 500 iters; shipped at 5500).

    Context

    Verdict recorded: "从零 + 强塑形是本机唯一验证过的发育路径" (from-zero plus strong shaping is this machine's only validated development path). The signature became a standing stop criterion in every later rung that touched a reward ("s1c B 臂签名,出现即停"). Crucially the boundary of the law was drawn explicitly when S2 continuation training was proposed: "当年证伪的是「奖励表中途改版的 fine-tune」… S2 续训奖励表全程冻结、只逐级加宽 plant/DR,属另一类" - continuing a checkpoint with the reward FROZEN while widening plant/DR one rung at a time is a different class and was allowed (and then worked, powering the whole S2/C lineage) - with the honest fallback that if frozen-reward continuation ever collapses, that rung retrains from zero and the doctrine gets re-examined with data. Fine-tune arms also need mechanical care: reset the checkpoint's collapsed noise_std (terminal 0.033 "会杀死探索") and account for iteration counters re-zeroing (curriculum gates fire immediately).

    Change

    Reward changes and lineage continuation permanently separated: reward revisions -> from-zero retrain; plant/DR widening -> frozen-reward continuation with per-rung gates; the B-arm signature promoted to a universal tripwire.

    Outcome

    No later reward revision was attempted by fine-tune; frozen-reward continuation carried S2 (PD/COM/friction rungs) and the C command ladder successfully from the s1e root.

    Mechanism

    A trained policy sits in an optimum of its reward's geometry; changing the reward moves the optimum but leaves the policy's exploration noise near-zero and its value function calibrated to the old returns - it disassembles the old solution faster than it can assemble the new one. Widening DR under a frozen reward instead keeps the optimum's identity and asks only for local robustification.

    Applies when

    • proposing to fine-tune an existing policy under a revised reward
    • planning a robustification ladder from a validated checkpoint
    • a continued run scatters then partially recovers then collapses
    “B 臂 fine-tune 证伪(打散→半恢复→摔回——从零 + 强塑形是本机唯一验证过的发育路径)。… 当年证伪的是「奖励表中途改版的 fine-tune」(B 臂,塑形突变致终盘摔回);S2 续训奖励表全程冻结、只逐级加宽 plant/DR,属另一类;若 s2_lag1 续训本身塌方,回退方案 = 该级从零重训,续训教义再议(拿数据说话)。”
    train/OMNI_V0_SPEC.md § 3. S1.3 / 4. 与 s1c fine-tune 证伪的关系
  • Gate a new reward term by its command so all old modes score pointwise identicalgate-new-reward-terms-by-command
    Mechanism understoodomnireward-shapingreward-shapingprocess

    When a reward term must be added mid-lineage, gate it on the condition that defines the new task so every pre-existing situation scores exactly as before - and still watch for value-rescale pathologies inside the new mode.

    Symptom

    Adding a lateral tracking reward (track_lin_vel_y_exp) ungated would have paid 0-2.0 per step even in modes with cmd_vy = 0 (healthy gait sway of vy ~0.1 already earns 1.28), shifting the whole reward table by a large bias and rescaling the value function - no longer "just adding one mode".

    Context

    C4 was the C ladder's only true reward surgery. Single-variable discipline required that the change be invisible to every existing mode. The chosen construction: gate_by_cmd=True - the term pays only when |cmd_vy| > 0.02, so for all modes with cmd_vy == 0 the term is pointwise zero, i.e. the reward is pointwise identical to before the change. The same trick appeared earlier in C1: replacing the vy L2 tax with a command-error version that is "对 cmd_vy≡0 逐点同值" (pointwise equal when cmd_vy is 0), explicitly classified as not-a-reward-change.

    Change

    track_lin_vel_y_exp added with gate_by_cmd=True (weight +2.0, std 0.15); the residual acknowledged honestly - inside the side bucket the values DO change, so the rung still watched the known reward-reshuffle pathology signature (s1c B-arm: scatter -> half-recover -> collapse) as a stop criterion.

    Outcome

    Old modes provably unaffected (pointwise-equal argument); attribution for any change in old-skill metrics stayed clean through the C4 redo series.

    Mechanism

    PPO's critic normalizes to the reward scale it sees; an ungated additive term shifts returns in every state and re-scales advantages globally, entangling the new skill with all old ones. Command-gating confines the new term's support to the new mode's state distribution, making "pointwise identical elsewhere" a provable property rather than a hope.

    Applies when

    • adding a tracking/shaping term for a new command or skill to a lineage that must not regress
    • reward change proposed while other skills are still being gated
    • reviewing whether a config diff counts as a reward change
    “只在 |cmd_vy| > 0.02 时付。不门控的话它对 cmd_vy≡0 的老模式也给 0~2.0 分(健康摇摆 vy≈0.1 → 1.28),等于给整张奖励表加一个大偏置、值函数尺度全变 … 门控后老模式逐点得 0 = 与加项前逐点同值,单变量纪律成立。… 但 side 桶内的值确实变了 —— 这仍是奖励表改版,开级盯 s1c B 臂签名”
    train/C_LADDER_RUN.md § 3d. gate_by_cmd=True(重要)
  • Curriculum-gate a penalty to the phase where its disease occurs - early on it only taxes explorationgate-penalties-to-the-disease-phase
    Replicatedwalkcurriculumcurriculumreward-shapingaction-rate

    For penalties aimed at late-stage pathologies (freezing, saturation, degenerate attractors), ramp the weight in only after exploration noise has decayed; anchor the terminal weight to measured healthy-vs-sick raw values, and shift all related tripwires to after the ramp completes.

    Symptom

    The action_saturation penalty, applied from iteration 0 in v8a, taxed exploration itself: with init_noise_std 1.2 the sampled actions paid ~-2.45/step before any policy had formed - while the disease it targets (clamp freezing) is a LATE pathology (v9 froze at iteration ~2624).

    Context

    v10 re-introduced the same penalty behind a curriculum gate: weight 0 until iter 1000, ramping linearly to -1.0 by iter 2000 - present only when the disease can occur, absent while exploration noise dominates. The trust argument was evidence, not hope: in v8a the term, while active, had pulled joint_pos_ref from 0.041 up to 0.155 and climbing - proof it can extract a policy from the frozen pit. Weight magnitudes were anchored to measured raw values (healthy v5 0.310 / v6 0.106 vs frozen v7 1.145 / v9 1.22 per step: at -1.0 healthy pays 6-18% of tracking, frozen pays 60%+, standing ~0). v10c then isolated the gated term as THE anti-freeze mechanism by single variable, upgraded to untouchable status in v11: "S 的门控机制(v10c 单变量铁案:任何情况下 不许撤,只许调终值)" - and v11 dared to relax other penalties only because S stood guard.

    Change

    action_saturation gated 0 -> -1.0 over iters 1000-2000 (later terminal value tuned -1.0 -> -0.5 with the gate mechanism itself frozen); tripwires adjusted to respect the gate's timing (freeze check moved to iter 2500-3000 to give the ramped term its effect window).

    Outcome

    Freezing stopped recurring while early training kept full exploration; the mechanism graduated from experiment to invariant within two versions.

    Mechanism

    A penalty's incidence depends on who occupies its support: early in training that is exploration noise (whose suppression starves learning), late it is the converged pathology. Time-gating aligns the penalty's presence with its target's presence, buying the constraint without the exploration tax - and tripwire timing must then be computed from the gate schedule, not from ungated precedents.

    Applies when

    • a structural penalty punishes exploration in early training
    • a late-onset pathology (freeze/saturation) needs a standing guard
    • deciding when a curriculum ramp should engage
    “v8a 实锤它的病根是"罚在采样动作上"——init_noise_std 1.2 的早期等于罚探索(~−2.45/步);而冻结是晚期病(v9 速率 2624 才死平)… 门控让它只在病发期在场。… v8a 里它在场时 joint_pos_ref 从 0.041 爬到 0.155 且仍在升——有从低谷爬出的实证。”
    train/WALK_V10_SPEC.md § 2. S 保险 —— action_saturation 课程门控
  • Diagnose a behavior failure by enumerating hypotheses and auditing each against the actual config, cheapest firsthypothesis-table-code-audit
    Replicatedwalkattributionattributionprocessreward-shaping

    Before changing anything, write the full hypothesis list for the symptom and audit each against the resolved config and measured magnitudes, cheapest check first; train only on the survivors.

    Symptom

    Real robot leaned forward "wanting to walk" but dragged its feet instead of lifting them - a symptom with many plausible causes and no obvious single fix.

    Context

    Seven hypotheses were listed and each checked against the actual training config files (velocity_env_cfg.py, isaac_values.py), ordered by check cost: missing foot clearance term (CONFIRMED, primary - feet_air_time existed but no swing-height term at all); energy penalties dominating (REJECTED - energy terms total -0.19 vs tracking +1.2, 16%); command range too narrow (CONFIRMED - (0.15,0.35)); nominal pose too crouched / action scale too small (HALF - knee 0.5 rad = 28.6 deg deep, scale fine); mixed PD across motor types (REJECTED - already grouped); missing base-height reward (REJECTED - present at -5.0); height-drop termination (REJECTED - none exists, which itself became finding #4 of the fix list).

    Change

    The audit produced a ranked fix list (add clearance penalty; widen speed range; reduce nominal crouch) with each rejected hypothesis documented so it would not be re-litigated.

    Outcome

    Three confirmed causes fixed over v5/v6: swing height went 22-23 mm -> 34 mm, tracking 81% -> 87%; the rejected hypotheses stayed rejected (no wasted rungs on energy weights or PD grouping).

    Mechanism

    Multi-cause symptoms invite guess-and-train loops; a written hypothesis table forces each candidate to be confirmed or rejected against actual values (not impressions), and cost-ordering the checks means most hypotheses die for the price of reading a config.

    Applies when

    • a real or sim behavior failure has multiple plausible causes
    • the team is about to "try a fix" without an audit
    • post-mortems keep re-proposing already-rejected causes
    “真机现象:躯干前倾像要走,脚抬不起来(拖着蹭)。按成本从低到高逐条核查 … | 1 | 缺 foot clearance | ✅ 成立,首要 | 有 feet_air_time,无任何摆动足高度项 | | 2 | 能量惩罚压过跟踪 | ❌ 不成立 | 能量类合计 −0.19,跟踪 +1.2,只占 16% |”
    train/WALK_DIAGNOSIS.md § walk 拖地问题 — 七条假设的代码核查结果
  • Compare the achieved reward to the computed ignore-floor to tell "never learned" from "learned but unprofitable"ignore-floor-diagnosis
    Mechanism understoodomniattributionattributionreward-shaping

    For any skill that trains flat, compute the reward the null policy would earn on that term; achieved==floor means the behavior never paid out (find why: exploration, reward observability, or feasibility) - do not tune weights first.

    Symptom

    C4 sidewalk failed on both arms; the question was whether the policy had found sidewalk and rejected it as unprofitable, or never found it at all - two diagnoses with opposite fixes.

    Context

    The theoretical value of the sidewalk tracking term for a policy that completely ignores the command was computable from the command distribution: 0.189. Trained final values landed at 0.187 (arm A) and 0.204 (arm B) - sitting exactly on the ignore-floor - while the honest balance at the optimum actually favored sidewalking (net +0.88/step inside the side bucket). Later the same arithmetic closed the whole saga: for the true reward landscape, doing real sidewalk scored 0.153 vs 0.944 for ignoring - the policy's refusal "是理性最优,不是探索失败" (rational optimum, not exploration failure) under one hypothesis, and under the final measurement-corrected account the policy had "每一次都在 做理性选择" (made the rational choice every time).

    Change

    Diagnostic rule adopted: compute the ignore-floor for the new term; if the achieved value sits on it, the behavior was never expressed in useful volume (or the reward cannot distinguish it - check both); if the achieved value is above floor but the behavior is absent at deployment, the policy sampled it and priced it out - then the reward balance, not exploration, is the lever.

    Outcome

    Correctly identified that PPO had not merely under-valued sidewalk; each subsequent hypothesis (waveform sign, regularization cage, exploration form, reward kernel) was tested against this floor arithmetic, which kept the search honest through three reversals.

    Mechanism

    Every reward term has a computable value under the null behavior; the achieved-vs-floor gap is a one-number audit of whether the optimizer ever monetized the target behavior. It converts "training failed" into one of two mechanistically distinct states with different fixes.

    Applies when

    • a new skill's tracking reward plateaus early
    • deciding between exploration fixes and reward-weight fixes
    • post-mortem of a failed curriculum rung
    “track_lin_vel_y_exp 训练终值恰好坐在「完全无视指令」的底分上(A 0.187 / B 0.204,理论值 0.189),而终点 balance 明明有利(side 桶内净 +0.88/步)。不是学会了不划算,是根本没学到。”
    train/C_LADDER_RUN.md § 3e. 为什么首战 FAIL
  • Audit which joints your imitation term constrains - a task that needs deviation is fighting the referenceimitation-term-scope-audit
    Mechanism understoodomnireward-shapingreward-shapingcurriculum

    List which joints your imitation/deviation terms actually constrain and check the new skill's required motion against that list; for balance-coupled joints deliver references as feed-forward residuals, not absolute-position targets - and never assume "reference = 0" is neutral.

    Symptom

    Sidewalk would not learn despite a dedicated tracking reward; meanwhile the gait-shaping imitation term (joint_pos_ref) computed its error norm over ALL 12 joints while its reference covered only the 6 sagittal joints - roll/yaw reference was constantly 0.

    Context

    Two prior generations had shown the forward gait itself was taught by joint_pos_ref, not discovered by PPO (v6 halved the shaping and swing height collapsed 35 mm -> 4 mm). So the reference is load-bearing - but sidewalk requires hip_roll to deviate from nominal, and the all-joints norm punished exactly that deviation: "一边悬赏一边罚过程" (posting a bounty while punishing the process). A follow-up experiment (C4-redo3, free_roll=True releasing the 4 roll joints from the norm) raised the regularization headroom 6x -> 27x yet sidewalk stayed flat and released hip_roll wandered, killing other skills - net negative, withdrawn. A --roll-absolute probe showed the converse failure: pinning roll to a clock-driven absolute trajectory drove tilt 6.9 -> 13.7 deg. Conclusion recorded: absolute-position imitation cannot teach actions that must be superimposed on state feedback.

    Change

    The audit reframed the problem: neither punishing roll deviation nor freeing roll nor absolute roll tracking works; the reference for a balance-coupled joint must be delivered as feed-forward under the policy's residual control (see feedforward-for-phase-locked-skills).

    Outcome

    free_roll rung: joint_pos_ref term rose 0.887 -> 1.104 (release confirmed effective) but vy stayed flat; regularization hypothesis eliminated by experiment.

    Mechanism

    An imitation error norm defines a cage: joints inside it are pulled to the reference in absolute position, so any skill requiring systematic deviation is taxed per step; but joints carrying active balance cannot follow absolute references either, since their correct position depends on state. The scope and the delivery mechanism of the reference are therefore design decisions per joint, not defaults.

    Applies when

    • adding a skill that moves joints your reference sets to zero/nominal
    • an imitation or deviation penalty coexists with a new tracking reward
    • considering releasing joints from a shaping term mid-lineage
    “前进步态也不是 PPO 自己发现的,是 joint_pos_ref 教出来的(v6 砍半塑形 → 抬脚 35 mm 塌到 4 mm…)。而 ref_joint_offset 原本只写 6 个矢状面关节,roll/yaw 参考恒 0 —— 侧走既没被教,roll 一偏离 nominal 反被 joint_pos_ref 扣分。一边悬赏一边罚过程。”
    train/C_LADDER_RUN.md § 3e. 为什么首战 FAIL / 3i. 解锁笼子
  • A quadratic clearance reward stalled for 3500 iters near target - switch to an indicator on accumulated heightindicator-reward-avoids-gradient-decay
    Mechanism understoodwalkreward-shapingreward-shaping

    When a shaped term plateaus near its target, check the gradient profile: replace vanishing-gradient forms with threshold/indicator forms for the final approach, and prefer delta-accumulation over absolute positions to immunize against frame offsets.

    Symptom

    The quadratic-error clearance term froze at -0.002 from iteration 2000 to 5500 - thousands of iterations with no progress on foot lift.

    Context

    Diagnosis: a quadratic penalty's gradient vanishes as the error approaches target, so exactly where the last millimeters must be earned the incentive fades to nothing. The replacement (Humanoid-Gym form): accumulate the swing-phase height climb per foot, reward a BINARY indicator |accumulated - target| < 0.01 masked to the planned swing window, reset on contact, weight +1.6 as a positive reward. Two properties: the indicator's incentive is constant until the threshold is crossed (no decay zone), and accumulating height DELTAS makes any constant sole-frame offset cancel automatically - which structurally sidesteps the earlier 0.0585 m zero-point bug ("顺带绕开我先前那个'忘了减 0.0585 导致惩罚恒为 0'的坑").

    Change

    Clearance reformulated from quadratic penalty on instantaneous height to indicator on per-swing accumulated climb (target 0.03 m by leg-length scaling, weight +1.6).

    Outcome

    Part of the v5 package under which lift finally moved (v5 29 mm, v6 34 mm vs the stalled 18-24 mm era); the offset-cancellation property removed one whole bug class from the term.

    Mechanism

    Policy-gradient learning follows the reward's local slope; quadratic shaping concentrates slope far from target and starves it near target, so convergence stalls precisely at the finish line. An indicator pays a constant bounty until the goal is met; formulating on deltas rather than absolutes removes sensitivity to reference- frame constants.

    Applies when

    • a reward term's value freezes short of target for thousands of iters
    • designing clearance/height/precision terms
    • reward code depends on absolute link positions
    “现行二次型在接近 target 时梯度趋零 —— 这正是 clearance 从 iter 2000 到 5500 卡在 −0.002 不动的原因。… 二值指示在跨过阈值前梯度恒定,没有衰减区 … 累积 delta 让 SOLE_OFFSET 自动抵消”
    train/WALK_V5_SPEC.md § 3. clearance 改峰值型(去掉二次型的梯度衰减)
  • Prove a new penalty actually fires - two ways a clearance term silently did nothinginert-reward-term-audit
    Mechanism understoodwalkreward-shapingreward-shapingplant-calibration

    Before training with a new reward term, log its realized per-step value under the current policy and confirm it is nonzero where intended - check coordinate zero-points against FK and check who occupies the term's gate; and never weaken the term that creates the states your new term needs.

    Symptom

    A newly designed swing-height clearance penalty could have trained as a no-op twice over, and the companion advice to lower feet_air_time actively backfired when tried.

    Context

    Instance 1 (zero-point offset): the proposed code used body_pos_w of the foot link, but that is the ankle_roll_link frame origin, which sits 0.0585 m above the ground even with the foot flat on it - so (0.03 - 0.0585) is always negative and the penalty is永远 0; the 0.0585 offset must be subtracted (verified identical in MuJoCo FK and Isaac). Instance 2 (gate occupancy): the clearance penalty fires only in swing phase; a dragging policy keeps both feet in contact, so the penalty is constantly 0 for exactly the policy it was meant to fix - and worse, any slight lift immediately incurs it, a reverse threshold. Lowering feet_air_time to 0.5 on that advice measurably collapsed air time to 0.0002 (below v3). Corrected understanding: "clearance 是把已有的摆动相抬高, 造出摆动相仍要靠 air_time" - air_time creates the swing phase, clearance raises it.

    Change

    Fixed the height zero-point; kept feet_air_time as the swing-phase creator with clearance layered on top; both errors documented as corrections to the team's own earlier advice.

    Outcome

    With both fixed, swing height rose from 22-23 mm (v2) to 29 mm (v5) to 34 mm (v6); the inert-term failure class entered the standing checklist.

    Mechanism

    A penalty's gradient exists only where its gate is occupied and its argument crosses its threshold; frame offsets shift the threshold out of reach, and phase gates can have zero occupancy under exactly the policy being treated. Terms interact as an ecology - one term must create the states in which another can act.

    Applies when

    • adding any gated or thresholded penalty (clearance, impact, slip)
    • a new term produces no behavioral change at any weight
    • body-frame positions are used in reward code
    “body_pos_w 是 ankle_roll_link 坐标系原点,平放触地时仍高出地面 0.0585 m。… (0.03 − 0.0585) 恒为负 → 惩罚永远是 0 … clearance 惩罚只在摆动相生效,拖地时两脚始终触地 → 惩罚恒 0;而一旦轻微抬脚就立刻扣分,对正在拖地的策略是反向门槛。… 正确认识:clearance 是"把已有的摆动相抬高",造出摆动相仍要靠 air_time。”
    train/WALK_DIAGNOSIS.md § walk_v4 独立验收 — 本文档给的两处代码/建议是错的
  • Three times a joint-angle stand-in for a foot-level quantity was gamed or lied - the absolute ankle roll sold stance width to buy flat feet, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apartjoint-space-proxy-for-task-space-quantity
    Replicatedrecoveryreward-shapingreward-shapingmeasurementreal-acceptance

    Express foot-level (task-space) goals and acceptance criteria in task space - link attitude and lateral spacing from world poses - never through joint angles that assume other joints are at zero, and measure the task-space value before trusting a joint-space estimate of it.

    Symptom

    The line's first real-robot get-up (v2_6, 2026-08-11) was "fairly stable", but after standing the feet were too close and the robot slid into the splits and fell several times; later reports added that it also got up through a split posture.

    Context

    (1) flat_feet penalized sum |q_ankle_roll|, but a flat foot is ankle roll compensating hip roll; the proxy taxed the compensated wide solution, and the untaxed combination was hips straight plus ankles at zero - flat and narrow. On hardware the lateral support shrank to hip width plus centimetres, lateral balance rested on 11 N*m ankle motors, and the feet slid apart. (2) The next rung's acceptance criterion "hip roll >= 20 deg" assumed zero hip yaw; at 50 deg of yaw the lateral contribution is x cos 50 ~ 0.64 - the same substitution again, inside a criterion. (3) A new task-space diagnostic reading the foot links' world poses measured v2_6c's stance at 0.159 m where the kinematic audit from joint angles had said 0.271 m (0.271 x cos 47 ~ 0.17).

    Change

    Rule written into the spec: task-space quantities are never expressed through joint-space proxies. V3.1's stance terms were all task-space: flat_feet_task from the foot links' world orientation, lateral foot spacing in metres, stand_pose stripped of both roll joints.

    Outcome

    From scratch with task-space terms (V3.1 P1c): lateral stance 0.355 m, foot residual tilt median 0 deg / P75 2.0 deg, all six criteria passing, mu 1.0-0.4 all 100%.

    Mechanism

    A joint proxy bundles the goal with everything else those joints do; the optimizer finds the combination the proxy does not tax, and a joint-based criterion silently assumes the other joints sit at their nominal.

    Applies when

    • rewarding flat feet, stance width, foot placement or end-effector pose
    • an acceptance criterion is written in joint angles for a geometric goal
    • joints with large yaw or coupled axes are involved
    “**病根 = 关节空间代理**:`flat_feet` 罚 Σ|q_ankle_roll|(§41 取的简易口径)。 "脚掌平"的运动学正解是 **踝滚补偿髋滚**(q_ankle_roll ≈ −q_hip_roll) … 代理把"脚平"和"站距"绑死在一起卖了。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §43 真机首试入账(2026-08-11)判读:flat_feet 的代理口径错误
  • Knee swing collapsed because it directly trades against the slip penalty - price the conflict explicitly and clamp what reward cannot holdknee-swing-vs-slip-pricing
    Mechanism understoodwalkreward-shapingreward-shapingattribution

    When a behavior collapses as another metric improves, look for the term pair trading them and set their price ratio deliberately (with escalation and reverse tripwires pre-registered); where the policy actively spends action budget to undo your target, stop paying more reward and clamp the target space structurally.

    Symptom

    Knee peak-to-peak swing collapsed across generations - v5 33 deg, v10 26-30, v10b 7-8, v11 6.5-8.6 - and rolling the clock back did not recover it, acquitting the clock; the collapse tracked the gated slip penalty instead: "屈膝与不打滑在当前奖励里直接对抗" - v10b's excellent 93 deg slip was purchased with knee amplitude.

    Context

    Reward-side flexion fixes had failed three times: raising reference amplitude backfired twice (v9/v11), and v11's deep-squat default was actively fought by the policy - it spent 0.68 of action budget pulling the squat straight ("被策略花 0.68 动作拉直反杀"). v12's design accepted the conflict as real and attacked on two tracks: (1) ECONOMICS - a direct knee_swing_amplitude reward (+0.3, target 0.55 rad, capped at 0.6/step = 55% of tracking), explicitly opposed to the slip penalty by design ("显式对立——这正是设计:v12 就是这场对抗的定价实验"), with an escalation ladder (K +0.3 -> +0.5, then slip -0.5 -> -0.3, one layer at a time) and a reverse tripwire (slip telemetry back at v10 levels -> slip weight to -0.8, accept ~20 deg knee compromise); (2) STRUCTURE - knee target bounds [0.2, 0.9] rad so full straightening is physically impossible (straightest 11.5 deg) and the 0.68 fighting budget is released. A bonus falsifiable prediction was attached: phase-lock strength tracks amplitude (v9_probe 48 deg locked 2.5 Hz; v11 low-amplitude 1.36 Hz unlocked), so if K works, hardware phase-lock should return - one change, two verdicts.

    Change

    knee_swing_amplitude reward + knee target clamp + pre-registered escalation/reverse levers; the failed reward-side-only approach retired.

    Outcome

    The lineage was frozen before v12 trained (strategic reset), but the diagnosis stands as the walk line's clearest example of two reward terms trading a behavior between them, with the pricing experiment and structural clamp fully designed and calibrated.

    Mechanism

    When two terms price opposite aspects of one motion (swing amplitude creates yaw momentum that becomes slip), the optimizer settles wherever the price ratio puts it - patching one side moves the equilibrium, not the conflict. Explicit pricing makes the trade a designed quantity; structural clamps remove the regions where the policy spends budget fighting the designer.

    Conflicts

    The pricing experiment (K vs slip) was designed and calibrated but never trained - the 2026-08-05 reset suspended v12; the collapse attribution table and the 0.68-action counterattack are measured, the remedy's效果 is untested.

    Applies when

    • one gait quality degrades in lockstep with another's improvement
    • the policy visibly fights a default pose or reference
    • repeated reward-side fixes for the same behavior have failed
    “膝摆塌在 v10→v10b,头号嫌疑是门控滑移罚(四代实测膝 p2p:v5 33° / v10 26~30° / v10b 7~8° / v11 6.5~8.6°;退时钟没救回 → 非时钟)——"屈膝"与"不打滑"在当前奖励里直接对抗 … 奖励侧修屈膝已三败 … v11 深蹲 default 被策略花 0.68 动作拉直反杀”
    train/WALK_V12_SPEC.md § 0. 定位 / 2. K —— 膝摆经济(与滑移罚的对偶)
  • Torque caps cannot soften footfalls - impact is falling-mass momentum, only the reward can treat itlanding-impact-not-fixed-by-torque-caps
    Mechanism understoodwalkreward-shapingreward-shapinghardwareactuator-modeling

    Classify each hardware symptom by the physics that sets it: quantities fixed by ballistic momentum at contact must be treated through the policy's trajectory (reward terms on approach velocity/force), never through actuator caps - and size such penalty weights against your own tracking reward, not a lighter robot's.

    Symptom

    Footfalls slammed at 1.78x body weight in sim baseline (human walking: 1.2-1.5x); the tempting hardware-side fix was cutting actuator torque limits.

    Context

    Measured directly: scaling torque limits from x1.0 down to x0.4 left peak landing force essentially unchanged (1.75 -> 1.78x body weight) - the impact force comes from the momentum of the falling mass at touchdown, not from motor effort. The fix has to change the trajectory, i.e. the policy, i.e. the reward: feet_contact_forces penalty above a threshold of 113 N (= 1.2x the 9.58 kg robot's weight), clipped, weight -0.005. The weight was sized locally, not copied: the reference robot's -0.001 would amount to 0.9% of tracking reward on this robot ("策略不会理它" - the policy would ignore it); -0.005 gives 4.4%.

    Change

    Added threshold-type contact-force penalty (-0.005, threshold 1.2x body weight) as one of v6-minimal's three changes; hardware torque cuts explicitly rejected as a footfall treatment.

    Outcome

    Landing force 1.72x -> 1.55x by v6 (target <1.5x, missed by 3% - progress booked honestly); the torque-cap dead end was documented so it would not be retried.

    Mechanism

    At touchdown the ground stops a ballistic mass; the impulse is set by approach velocity and effective inertia, which motors can no longer influence in the final instant. Only earlier trajectory choices (approach velocity, timing) reduce it - and those are selected by the reward, not by actuator limits.

    Applies when

    • footfall impact or landing noise on hardware
    • proposals to derate torque as a softness fix
    • importing contact-force penalty weights from another robot
    “⚠️ 硬件限扭降不了落脚力 —— 砸地力来自下落质量的动量: 实测 tau ×1.0→×0.4, 落脚力 1.75→1.78× 体重纹丝不动。只有这条奖励能治。… ⚠️ 权重不能用 Pi 的 −0.001 —— 实测在我们身上只占跟踪奖励的 0.9%, 策略不会理它 (Pi 6.94 kg 更轻)。−0.005 给到 4.4%。”
    train/WALK_V6_MINIMAL.md § ③ 新增 feet_contact_forces
  • Narrowing the speed range to stop high-speed falls entrenched crouch-shuffling - judge gait quality at the speed that demands a gaitlow-speed-commands-reward-dragging
    Mechanism understoodwalkcurriculumcurriculumreward-shapinggate-battery

    Set command ranges to include speeds that physically demand the target behavior, and evaluate behavior-quality gates at those speeds; when a restriction is added to suppress a failure, check what new optimum it creates at the remaining commands.

    Symptom

    After the command range was narrowed to (0.15, 0.35) m/s (to treat walk_v1's 134% overspeed and 8.3 s fall at 0.5), the policy settled into crouched foot-dragging; tracking rose monotonically with speed (63% at cmd 0.2, 76% at 0.3, 87% at 0.45), showing low speeds were where the degenerate gait was optimal.

    Context

    The narrowing advice was the author's own and is retracted in the file: it treated the symptom (falls at speed) while reinforcing the root cause (at 0.15-0.35 m/s, shuffling in a crouch is globally optimal - the Froude number is so low that even humans would not lift their feet). A zero-cost experiment confirmed the flip side: at cmd 0.5 the same policy met BOTH tracking (81%) and clearance (23.0/23.2 mm) standards.

    Change

    Speed range widened back toward (0.15, 0.5) - upper bound deliberately slightly above the mechanically feasible ~0.44 m/s so the policy finds the boundary itself; acceptance re-pointed: gait-quality criteria (tracking, clearance) judged at 0.45-0.5 m/s, low speed kept only as a survival check.

    Outcome

    v4 -> v6 progression under the widened range delivered 87% tracking with 34 mm clearance; the "low command = drag" account was confirmed by the monotone tracking-vs-speed curve.

    Mechanism

    Command distribution is part of the reward: physics prices gaits per speed, and at very low speed the energetic optimum is no swing phase at all. Restricting training to that regime makes the degenerate gait the correct answer to the posed problem - and grading a gait at a speed that does not require stepping measures nothing.

    Applies when

    • a gait degenerates after a command-range restriction
    • quality metrics improve monotonically toward the range boundary
    • writing acceptance criteria for gait quality vs survival
    “现在看那个建议可能起了反作用:0.15~0.35 m/s 下蹲着蹭就是全局最优,抬腿反而亏。收窄治的是"高速摔倒"的症状,却强化了拖地的病根。… 验收标准里的 cmd 0.2 本身就是拖地速度(Froude 数极低,人在那个速度下也不抬脚)。accept_v2 应把速度跟踪与 clearance 的判定点改到 0.45~0.5 m/s”
    train/WALK_DIAGNOSIS.md § ② 放宽速度区间 / ① 零成本实验
  • The restart reward table lists a reason for every term AND a lesson for every exclusion - absent terms are removed, not zero-weightedminimal-reward-table-with-provenance
    Mechanism understoodomnireward-shapingreward-shapingprocesscontract-freeze

    Maintain the reward table as an evidence ledger: every term cites the episode that justifies it, every excluded term cites the episode that convicted it (including the development stage it is valid at), and retired terms are deleted from the config, never left at weight zero.

    Symptom

    Seven walk generations had accumulated an entangled reward table where nobody could say which term earned its place; the restart needed a table that could be audited line by line.

    Context

    The minimal table v2 was built under three written principles: "一项管 一件事、结构性反抬脚的项一个不留、塑形只留一套相位逻辑;不在表里的一律 不加" (one term per job; zero structurally-anti-lift terms; exactly one phase-shaping logic; nothing outside the table gets added). Every row carries its provenance (e.g. base_height target = standing height cites the v4 crouch lesson; split x/y tracking cites the merged-exp gradient hole; world-frame yaw cites the v1/v3 body-frame lesson). Every EXCLUSION carries its same-type precedent: feet_landing_vel out because it is poison while the gait is unformed (drag pays 0, lifting pays - the v4-clearance / v8a-B "reverse threshold" family) though it was fine in v7 when the gait already existed - term validity depends on developmental stage; feet_air_time out with its measured non-lever evidence (v6 had it at 2.0 and still lifted 4 mm); and replaced terms are REMOVED from the config ("置 None,不是权重 0 挂着") so audits see truth, not dormant weight.

    Change

    Reward table rebuilt as ~20 rows each with weight + provenance column; exclusion list maintained alongside with the falsifying episode for each; dormant terms deleted rather than zeroed.

    Outcome

    Later revisions (S1.2/S1.3) modified the table by citing and updating specific rows' evidence rather than re-arguing the whole design; the table doubled as the lineage's reward-lesson index.

    Mechanism

    A reward table is a set of standing hypotheses; attaching each row's evidence makes revisions targeted and reversible, and recording why a term is absent prevents the cycle of re-adding known poisons. Deleting vs zero-weighting matters because config audits and DR interactions see the term either way - a zero-weight term is dormant complexity waiting to be flipped on wrongly.

    Applies when

    • designing a reward table for a restart or new task
    • someone proposes re-adding a previously removed term
    • auditing which reward rows still earn their place
    “原则:一项管一件事、结构性反抬脚的项一个不留、塑形只留一套相位逻辑;不在表里的一律不加 … feet_landing_vel(评审 #4):拖地时代价恒 0、抬脚才收费——与 v4-clearance/v8a-B 同属「反抬脚门槛」家族,步态未成形时是毒;v7④ 加它时步态已存在。… 已从 cfg 移除(置 None),不是权重 0 挂着。”
    train/OMNI_V0_SPEC.md § 3. 最小奖励表 v2 / 明确不带
  • Ungated phase shaping made standing 42x more expensive than stepping - and the stepping was cooking the hip motorsmoving-gate-42x-stand-tax
    Mechanism understoodomnireward-shapingreward-shapinghardwarereal-acceptanceprocess

    Gate every phase/clock-driven shaping term on the command that justifies motion, price the cmd=0 case explicitly during design - and when a reward flaw is sim-only, book it with trigger conditions instead of operating immediately on a working lineage.

    Symptom

    At cmd = 0 the sim policy never stood still - it stepped in place and crept 0.98 m per 20 s; on the real robot the same lineage's stepping made hip_roll motors run 20 degC hotter than every other joint (43-48 degC vs 25-28).

    Context

    The arithmetic closed it: the three phase-shaping terms (joint_pos_ref 1.6, feet_contact_number 1.2, feet_clearance_swing 1.6) are driven by the gait clock with NO command gating, so standing at cmd=0 forfeits 3.32/step of shaping while honest standing earns only 0.078 of tracking - stepping wins 42x. The heat chain: perpetual stepping = perpetual single support = one hip_roll stalled at ~4.3 N*m (25% of torque limit) carrying the torso's frontal-plane moment - two hip_rolls = 90% of whole-machine steady-state I2R; measured stand-vs-step comparison: total heat -71% when actually standing. The suspended test acquitted the actuator (0.21 rad sag -> 0.0008 rad in air) and mechanics vetoed the easy fix ("降 kp 救不了热" - equilibrium torque equals the external load regardless of kp). The fix (moving_gate: hard-gate the three shaping terms on |cmd| > eps) was designed - then DEFERRED by the user because the real robot at the time stood fine: "真机不表现该问题, 为真机不存在的病改奖励表不划算", with written trigger conditions (C-ladder stand row persistently failing, or real robot starting to step/drift at cmd=0) and the known hazard tag (this is exactly the reward-change class that triggers the B-arm signature). When the real robot later DID step and cook, the booked trigger fired and moving_gate moved from debt to to-do with its benefit re-priced: "停止空烧 hip_roll,稳态发热降 ~71%".

    Change

    moving_gate designed with the gate_by_cmd convention; deferral, triggers, and expected heat recovery all pre-registered instead of patching the reward for a then-sim-only symptom.

    Outcome

    The cmd=0 stepping went from mystery to closed arithmetic; the thermal measurement (hip_roll +20 degC) quantitatively confirmed the 90%-of-heat prediction; the reward change waited for real-world justification instead of spending a risky revision early.

    Mechanism

    Clock-driven shaping terms define a perpetual-motion bounty unless gated by command; the resulting idle gait is not a training bug but the table's optimum. Its cost surfaces on hardware as stall-torque heating set by statics (mass x lateral offset), which no gain change can remove - only removing the motion (gating) or widening the stance can.

    Applies when

    • the policy steps in place or creeps at zero command
    • specific joints run hot in idle behaviors
    • deciding when a known reward flaw justifies a risky mid-lineage fix
    “塑形合计 −3.32/步 … 净: 站定亏 42 倍 … 两颗 hip_roll 4.29 / 4.09 N·m(各占限扭 25%),占全机稳态 I²R 的 90% … 吊挂实测 hip_roll 跟踪误差 0.21 rad → 0.0008 rad … 降 kp 救不了热 … 真机不表现该问题, 为真机不存在的病改奖励表不划算”
    train/README.md § C1 FAIL 节 (cmd=0 的 sim/真机分歧记账) / C2 真机 A/B 三b 发热定性
  • Fix a too-deep nominal pose before adding any penalties - the default stance defines the basin training starts innominal-posture-before-penalties
    Mechanism understoodwalkreward-shapingreward-shapingplant-calibration

    Before tuning penalties on a degenerate gait, audit the nominal pose and height targets against morphology and published ratios; if the default stance encodes the degenerate behavior, fix it first - and recompute dependent quantities (init height) by FK, not by hand.

    Symptom

    Policy lived in a crouched shuffle; nominal knee angle was 0.5 rad (28.6 deg) - deeper than published configs (Unitree G1 0.3 rad / 17.2 deg, Booster T1 0.4 rad) - so the policy's starting point and its action-space center both sat inside the crouch basin.

    Context

    Initially ranked "secondary" in the local diagnosis, this was promoted to co-first priority by the cross-check against published reward tables, which states that with nominal knee flexion above ~0.4 rad, fixing the posture must precede adding any penalties ("改这个之前别加任何 惩罚都是白费"). Companion base-height items: walk profile had weakened base_height_l2 to -5.0 (base class -10, field standard -10 to -20, "the second most common cause of death"), and the height target must be the STANDING height (0.384), not the crouch height.

    Change

    Nominal knee 0.5 -> 0.3 rad with init_base_height recomputed by MuJoCo FK (0.3739 -> 0.3802); base_height_l2 restored to -10 with standing height target; both bundled as first-priority alongside the clearance term.

    Outcome

    Part of the v5/v6 package that lifted swing height to 34 mm and tracking to 87%; the crouch basin stopped being the default answer.

    Mechanism

    The nominal pose is the fixed point every regularizer pulls toward and the point where action=0 lands; if that point is itself the degenerate posture, every penalty fights the geometry. Correcting the attractor is prior to shaping the gradient field around it.

    Applies when

    • policy converges to a crouched or collapsed posture
    • nominal joint angles were chosen for stability rather than gait
    • base-height reward targets or weights were locally weakened
    “研究明确说"nominal 膝屈超过 ~0.4 rad 必须先改,改这个之前别加任何惩罚"。我们是 0.50,超标。… base_height_l2 在 walk profile 里被减到 −5.0(基类是 −10)。研究说这是"第二常见死因"且应 −10 ~ −20。改回 −10。目标高度用站立高 0.384 是对的(研究要求 target 必须是*站立*高度而非蹲姿)。”
    train/WALK_DIAGNOSIS.md § 修正 ①(升级优先级) / 修正 ④
  • An edge-triggered landing penalty missed the tail and fired after the harm - penalize overspeed continuously inside the contact windowpenalize-tail-before-touchdown
    Mechanism understoodwalkreward-shapingreward-shaping

    Penalties aimed at impact/violation events must (a) price the excess over a threshold, not the mean, and (b) be active on the approach (state-gated window), not triggered by the event - check your control rate can even see the event you are penalizing.

    Symptom

    The v7 landing penalty (vz^2 on the contact-force rising edge, weight -10) did not bite: landing-velocity 95th percentile stayed at 2.61 m/s against a 0.3 target.

    Context

    Two structural faults were identified: (1) it penalized the MEAN over sparse events - many soft landings dilute the occasional violent slam, while the damage (GRF peaks, motor peak load) lives in the tail; (2) it fired AFTER touchdown - at 50 Hz evaluation the rising edge is aliased by physics decimation, so the read vz is often the already-decelerated post-impact value: underestimated, and with no shaping gradient before contact. Replacement: continuous penalty while the sole is inside a height gate (h < 0.03 m): relu(-vz - 0.30) - only the excess over an allowed approach speed is penalized (tail only), and gradient exists for several frames BEFORE touchdown. The sole-height computation again subtracts the 0.0585 m link offset ("WALK_DIAGNOSIS 坑#1, 别再踩"); the edge-triggered version was kept as a diagnostic only.

    Change

    feet_landing_vel reformulated: edge-event vz^2 -> in-window relu(-vz - v_ok) with v_ok 0.30 (conservative vs the sqrt(L)-scaled human value ~0.19, to be tightened after passing), h_gate 0.03, weight unchanged -10.

    Outcome

    The failure analysis of the first form was written before the second was trained; the v_ok escalation path (0.30 -> 0.45 if the robot becomes afraid to land) was pre-registered in the risk table.

    Mechanism

    Sparse-event mean penalties optimize the average case while the constraint is a quantile; and any penalty evaluated only at/after a discrete event gives the optimizer no gradient along the approach trajectory that determines the event. A state-gated continuous excess penalty fixes both: it prices only violations and shapes the approach.

    Applies when

    • impact/landing penalties fail to move tail percentiles
    • a penalty is triggered by contact edges at a coarse control rate
    • designing constraint-style penalties for rare violent events
    “罚的是均值路径:上升沿是稀疏事件 … 大量软着陆稀释偶发猛砸;而伤害在尾部 … 罚在触地后:50 Hz 评一次,上升沿被物理 decimation 混叠,读到的 vz 常是撞完已减速的值——既低估,又没有触地前的塑形梯度。”
    train/WALK_V8_SPEC.md § 2. 改动 B — 落地惩罚改罚尾部、罚在触地前
  • Decompose the offending quantity by channel first - then penalize the failure event, not the jointspenalize-the-slip-not-the-joint
    Mechanism understoodwalkreward-shapingreward-shapingattribution

    Before penalizing motion to fix a side effect, measure which channels actually carry the offending quantity; prefer penalties conditioned on the failure event that are exactly zero for healthy behavior - and do not medicate behaviors that measurement shows are not sick.

    Symptom

    Heading drift with support-foot yaw slip (v5: 212-284 deg accumulated over 15 s); the previous v6 draft had attacked it by penalizing lateral joints (a roll 4.0 / yaw 2.0 "home" group) - which collapsed training into the standing basin.

    Context

    Before choosing the penalty target, the yaw angular momentum was decomposed by joint group with MuJoCo subtree_angmom weighted by real walking joint velocities: pitch-class joints (hip_pitch + knee) carry 95.3%, hip_roll 3.3%, hip_yaw 1.4%. The failed "home" group had been taxing 2.7/step to manage a 4.7% channel. The replacement, feet_yaw_slip (-0.2, |support-foot yaw rate| while in contact), targets the failure event itself and - decisively - costs a non-slipping gait exactly zero, which "横向回家组做不到". The same rung's do-not-do table applied the complementary principle to foot spacing: measured 196-214 mm, stable, no crossing - "没病不吃药" (no disease, no medicine).

    Change

    Removed joint-usage penalties for the drift problem; added the event-conditional slip penalty (-0.2, realized tax 0.141/step = 12% of tracking) alongside the existing linear-slip term.

    Outcome

    Turn-gain left/right difference improved 70% -> 19% and heading 185 -> 60.3 deg by v6 without a standing-basin collapse; the 2.7/step lateral tax never returned.

    Mechanism

    Penalizing joints taxes every use of a channel including healthy use, and if the channel carries little of the offending quantity the tax buys nothing while pushing the optimum toward immobility. An event-conditional penalty (slip while in contact) prices only the failure, leaving the healthy gait's cost surface untouched - and the channel decomposition tells you in advance whether a joint-side fix can even work.

    Applies when

    • choosing a penalty target for drift/slip/impact problems
    • a proposed penalty taxes joints or motions rather than failure events
    • a previous joint-penalty attempt collapsed the gait
    “pitch 类 (hip_pitch + knee) 占偏航角动量 95.3% … hip_yaw 1.4% … 压 hip_yaw 是管 1.4% 的通道收 2.7/步 的税 —— 上一轮正是这样把策略推进了站立盆地。滑移项不惩罚走路: 不打滑的步态代价为零, 这是横向"回家"组做不到的。”
    train/WALK_V6_MINIMAL.md § ① / ② 新增 feet_yaw_slip
  • Every gate on a penalty is an exit - to stop paying a stance tax the policy parked 2 deg outside a 30 deg uprightness gate (lunging), and, from scratch, just under a height gate (crouching); a positive band and an always-on guard fixed bothpenalty-gate-is-an-escape-hatch
    Replicatedrecoveryreward-shapingreward-shapinggate-battery

    Never gate a penalty on a state the policy can leave by getting worse; use always-on guards for what must never happen and positive, gated bands for what you want, and count every gate on a penalty as one more escape route to check in the logs.

    Symptom

    V2.9 added stance_width_task = relu(0.34 m - foot spacing) x standing gates, weight -10. Three checkpoints scored 0% on acceptance: standing height reached, feet on the ground, angular rate low, but the torso leaned 32.5/32.3/31.9 deg in a fore-aft lunge. The training dashboard read "tax paid off, base_height at full value".

    Context

    The penalty was gated by uprightness (tilt < 30 deg) and standing height; its tax had no time gate (500 steps x -1.65) while the standing income sat behind a 3 s zero gate (~325 steps), so leaning just past 30 deg lost a little gated income and saved the whole tax. base_height has no upright gate, so the lunge still collected it. The first metric also measured full horizontal spacing, so a staggered lunge counted as "wide".

    Change

    Two laws written down: a penalty may only carry gates the policy cannot escape by getting worse (make it an always-on guard) or it becomes a positive band ("not earned" is not "escaped"); and width is measured laterally in the base yaw frame. From scratch (V3.1 P1) with the lateral metric but the same gates, the policy parked just under the height gate instead (base_height 1.176/1.5, h ~ 0.30 m against a 0.3264 gate), the beta curriculum never advanced in 1,700 iterations, and the run was stopped early. P1b flipped the penalty into a positive band +2.0 x clamp(lateral / 0.34) x standing gates; P1c added yaw_guard = -5 x relu(|hip_yaw| - 30 deg), always on, no gate, no exemption.

    Outcome

    P1b: lateral stance 0.364 m, all four categories 100%, MuJoCo mu 1.0/0.4 both 100%. P1c: all six acceptance criteria passed for the first time on the line (hip-yaw saturation 1.2%), the guard's tax converging to -0.016 (almost never touched).

    Mechanism

    A gated tax that is not paid is saved, so the policy moves to the cheapest state just outside the gate; a gated income that is not earned is simply lost, so a positive band has no exit. HoST's style penalties are ungated or binary - the same law seen from the other side.

    Applies when

    • adding a penalty multiplied by an uprightness, height, phase or contact gate
    • a policy settles just beyond a gate threshold
    • training metrics look paid-up while acceptance collapses
    “**罚项的门 = 策略的逃生门**。带直立门的负项可以靠"变得更差"(倾出门外) 全时免税;HoST 的 style 罚全部无门控/二值恰是同一律的反面实证。修律: 负项只许挂"变差逃不掉"的门(上限护栏 always-on),或改正向 band (收不到 ≠ 逃掉)。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §46 定案(血统内第四败 + 两条新律)
  • Size each joint's action authority to its measured working range - lock a channel to zero only when its job is provably elsewhereper-joint-action-scale-lockdown
    Observed oncewalkobservation-designobservation-honestyreward-shapingcontract-freeze

    Set per-joint action scales from measured target ranges and momentum decompositions: full authority for working channels, working-range authority for balance channels, zero for channels whose contribution is measured negligible - and prefer this structural quieting over perpetual reward penalties, keeping the contract dimensions intact.

    Symptom

    "Walks crooked" - roll and yaw channels wandered; hip_yaw peak-to-peak reached 22.3 deg in v11 while contributing essentially nothing to locomotion; a uniform action scale of 0.5 gave every joint the same authority regardless of its actual job.

    Context

    The v12 design replaced the scalar action scale with per-joint scales justified by measurements: pitch-class 0.5 (the gait's entire working channel - untouched); hip_yaw 0.0 - lossless because it carries only 1.4% of yaw momentum (v6 decomposition) and straight-line targets use merely +/-0.006-0.03 ("乱动纯属浪费"); roll 0.2 - NOT zero, because lateral balance and weight transfer are roll's unique job (locking it would degenerate into foot-edge rocking, "比现在更歪"), and 0.2 covers the measured working range +/-0.06-0.19 while "0.5 的另一半全是 '歪歪扭扭'的来源". Costs were accepted consciously: turning demoted to an observation row with a fallback (yaw 0 -> 0.1 in v13). Contract preserved: the 12-dim action interface unchanged, yaw values simply neutralized. The structural lockdown also RETIRED the reward-side hip_yaw_quiet penalty - "A 的 yaw=0 结构性取代,不再付奖励塑形成本".

    Change

    action_scale_joint pitch 0.5 / roll 0.2 / yaw 0.0 wired through robot.yaml -> policy_io (verified: all-ones action gives hip_yaw target exactly 0) -> Isaac action term, guarded by the contract checker ("它就是抓这种双侧不一致的").

    Outcome

    Designed and verified on the shared side before the lineage freeze; stands as the pattern for authority sizing: structure replaces reward shaping wherever a channel should simply not act.

    Mechanism

    Action scale is a per-channel authority budget; uniform budgets give noise channels the same voice as working channels, and reward-side quieting then pays a permanent shaping tax for what a zero scale provides for free. But zeroing is only lossless when decomposition proves the channel's contribution negligible AND no unique function (balance) lives there.

    Conflicts

    Wired and verified on the config/deploy side but never trained - the 2026-08-05 reset suspended v12 before the Isaac-side run.

    Applies when

    • some joints wander without contributing to the task
    • a quieting penalty (deviation/L1) taxes every step forever
    • deciding action-space authority for a new task or robot
    “yaw=0 是无损的:实测它只贡献 1.4% 偏航动量、直行目标只 ±0.006~0.03,乱动纯属浪费。… roll 不能为 0:横向平衡/重心换脚是它的独有职责,锁死会退化成脚缘摇摆(比现在更歪)。0.2 的依据:各代实测 roll 目标只用 ±0.06~0.19,0.5 的另一半全是"歪歪扭扭"的来源。”
    train/WALK_V12_SPEC.md § 3. A —— 逐关节动作幅度(用户"只动 pitch"的安全版)
  • The get-up kept getting faster because standing earlier paid more every step - lowering torque authority barely slowed it, and only zeroing the standing income for the first 3 s moved the pace into the design bandper-step-income-drives-speed-time-gate
    Mechanism understoodrecoveryreward-shapingreward-shapingcurriculum

    When a skill is too fast, find the term that pays for finishing early and gate that income by time; keep the "get into position" term ungated so the policy does not learn to wait, use a ramp instead of a cliff, and confirm with a paired same-level experiment that the drift is motivational before changing it.

    Symptom

    The user judged the get-up too fast (Isaac medians about 0.7-1.6 s) and suspected path dependence: the policy seemed to get faster the longer it trained.

    Context

    Lowering the beta authority 0.40 -> 0.30 cut impact but moved supine only 1.70 -> 2.00 s: coordination-limited, not torque-limited. A paired experiment inside one beta level (checkpoint 15,600 vs 18,499, +2,900 iterations, same ruler) measured the drift: get-up medians -7 to -10%. The spec concluded the motive, not the path, was the cause - per-step standing income pays for every early step, and any lineage (even one from scratch) races toward the fastest solution inside its constraints.

    Change

    V2.5: the standing income (base_height, stand_pose, still, feet_on_ground) multiplied by w(t) = clamp(t/3 s, 0, 1); upright deliberately NOT gated, so righting and sitting up early still pay and the policy is not taught to lie flat and wait; a ramp, not a step. V2.5b: zero before t0 = 3 s, then a 1 s ramp.

    Outcome

    V2.5: Isaac 100%, get-up +17-43% slower, MuJoCo 100/100/98/100% (the best cross-simulator reading yet), still short of the 3.5-4.5 s design band - a linear ramp only discounts early income. V2.5b: MuJoCo supine 2.04 -> 4.10 s and prone 3.18 -> 4.04 s, inside the band; the Isaac pace barely moved (a lineage habit on a gradient-free plateau). Later the zero gate proved harmful when trained from scratch (curriculum-history-is-part-of-the-product) and in the V3.1 lineage (time-gate-vs-wide-stance-retire-the-fix).

    Mechanism

    Constraints on authority or velocity change how the fastest solution looks; the time structure of the task income decides how fast the fastest solution is.

    Applies when

    • a policy is faster or more aggressive than wanted and constraints do not slow it
    • progress-style rewards pay every step spent at the goal
    • performance drifts faster with more training at fixed settings
    “**V2.5 机制(唯一)**:站立收入(base_height/stand_pose/still/feet_on_ground) 乘时间斜坡 w(t)=clamp(t/T_gate,0,1),T_gate=3.0 s;**upright 刻意不门控** (翻正/坐直早期照常拿钱,防"躺平等门开" … **用户假设量化 证实:逐步计酬动机在 β 包络内持续压缩时间,约束挡不住动机 —— V2.5 动机层 修法为正解。**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §39 V2.5 预注册 / §39 补 配对实验
  • The walk phase machine structurally cannot express flight - rebuild the representation for duty < 0.5, teach flight with a mask tax, never a cliff bountyphase-machine-structural-limits
    Mechanism understoodrunreward-shapingreward-shapingcurriculumgate-battery

    When a new gait changes the contact pattern's structure, audit whether the phase/mask representation can express it and rebuild the representation if not; teach the new contact pattern with graded mask-mismatch pressure and count it in acceptance with artifact-proof definitions (minimum segment length), never with cliff bounties.

    Symptom

    Running requires both feet airborne, but the walk-era phase machine switches legs by the sign of sin(phase) - with duty < 0.5 the two swing windows must OVERLAP during flight, which a sign-switching representation cannot express at all.

    Context

    The run phase machine was re-architected rather than patched: per-leg phases (left = phi, right = phi+0.5 mod 1) with leg_phase < duty defining stance, aligned to the walk sin convention at duty=0.5 so the machines agree where their domains overlap. Flight is taught by the SAME mechanism that once cured foot-dragging, direction reversed: in the two planned flight windows the contact mask is (0,0) and feet_contact_number_duty charges -0.3 per foot still on the ground - a mild ~0.16/step tax, deliberately NOT a cliff: "悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文". The reference shape (half-sine bump over swing progress) is zero at window boundaries by construction, eliminating the clearing-window that the C4 probe measured to cost 13-19% on non-sinusoidal references. The shape self-check ("零代价选项是什么") was run on three behaviors: standing pays ref everywhere (known cmd=0 stepping risk, booked), walking pays only the flight-window tax, proper running collects full marks.

    Change

    New rewards.py run section (leg_phase_duty / stance_mask_duty / ref_run / clearance_run / contact_duty) with the walk versions untouched byte-for-byte; flight acceptance metric defined with a segment-length floor (>=40 ms to count) so numeric contact flicker cannot fake flight.

    Outcome

    Flight became expressible and taught by a calibrated mild pressure; the walk lineage's phase code stayed frozen as its own contract.

    Mechanism

    A phase representation defines which contact patterns exist in the reward's vocabulary; duty cycling below 0.5 introduces states (double-flight) outside a half-period sign convention's language, so no weight tuning can teach them. And rare desirable events taught by cliff-shaped bounties invite hacks (jumping in place); the graded mask tax prices the planned pattern without creating a jackpot.

    Applies when

    • extending a walking stack to running/jumping (duty < 0.5)
    • a desired contact pattern never appears despite reward increases
    • defining flight/contact acceptance metrics
    “duty<0.5 时摆动窗 (1−duty)T > T/2,两腿摆动窗在腾空段重叠 —— sin 符号切腿的机制结构上表达不了"双脚同时在空中"。… feet_contact_number_duty 对"还踩着地"持续 −0.3/脚 —— 与 walk 治拖地同一机制,方向相反。不腾空的税 ~0.16/步 … 梯度温和不构成悬崖(悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文)。”
    train/RUN_V0_SPEC.md § 5. 相位机设计
  • The standing-pose reward had been pulling toward the narrow stance the whole line was fighting - a zero-training kinematic audit of the target vector found it, after first auditing the wrong nominalpose-target-geometric-audit
    Mechanism understoodrecoveryattributionreward-shapingattributionplant-calibration

    Before training on a pose target, audit it with forward kinematics - is it geometrically consistent (feet flat, intended stance, intended width) and is it the frame you think it is (action nominal vs standing default)? A posture term's target may itself be the attractor you are fighting.

    Symptom

    Several rungs aimed at widening the stance failed; the stance stayed narrow as if something kept pulling it back.

    Context

    The audit (MuJoCo forward kinematics, no training): every 5 deg of hip roll widens the stance ~5.5 cm (0.271 m at 5 deg, 0.383 m at 15 deg); at 47 deg of hip yaw a wide stance cannot be flat-footed (residual foot tilt ~0.7 x hip roll), which explained the stalled rungs. The first report also said the stand_pose nominal (hip roll 25, knee 60) has a 63 deg residual foot tilt - but that was the action frame's nominal (the limit-midpoint squat), which stand_pose never used, despite a docstring warning not to mix them. stand_pose's real target was DEFAULT_JOINT_POS: the contract's all-zero pose, legs parallel, ~0.22 m apart.

    Change

    The disease statement was corrected in writing: the narrow stance was not an accidental by-product of proxy traps but the target stand_pose had been actively rewarding. A stored "narrow the stance" knife was marked toxic. V2.8 moved the target to a flat 15-deg stance (sigma 3 -> 1.5, flat_feet margin 5 -> 20 deg).

    Outcome

    V2.8 still failed in-lineage (stance unchanged, feet nearly overlapping, mu 0.4 transfer 2%) - see stance-decided-by-get-up-path - and the from-scratch V3.1 removed both roll joints from stand_pose and put width into a task-space term, which is what finally produced a 0.355 m flat stance. The 63 deg finding was kept as a warning: an action nominal used as a standing target would be a ready-made pit.

    Mechanism

    A posture term with a sharp kernel around the wrong target is an active attractor; every other term fighting it pays twice.

    Applies when

    • a posture keeps returning despite penalties against it
    • a reward uses a default or nominal pose as its target
    • the contract has more than one "nominal" (action frame vs standing pose)
    “上文"stand_pose 的 nominal (hip25/knee60) 残倾 63°"**审计错了对象**:那是 **动作参考系 nominal**(限位中点蹲),stand_pose 从未指向它(函数 docstring 原文即警告"两者别混",还是混了 —— 记档)。 … **修正后的病根陈述:窄站距不是代理陷阱的意外副产物,而是 stand_pose 一直在主动奖励的目标本身**”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 勘误(实现时抓到):审计混了两个 nominal —— 真病根比误诊的更直白
  • Prone get-up sat at 0/159 until two gated hinge terms moved the seated feet - first sideways (561 -> 360 mm), then fore-aft (-168 -> +56 mm) - and success went to 158/159 with nothing else changedprone-dead-end-is-foot-placement
    Mechanism understoodrecoveryreward-shapingreward-shapingattributioncurriculum

    When a stuck state and a successful state differ geometrically, penalize the discriminating quantity with a gated hinge that is exactly zero in the state the policy actually reaches (measure it - not the nominal), then re-probe: flattening one axis can move the discriminant to another.

    Symptom

    Prone falls always righted and then sat with the feet splayed wide or tucked behind the hips, from where the policy never stood (0% for four generations).

    Context

    Four lines of evidence pointed at foot position: the configuration probe (ankles 215 mm apart stood 52.3%, 561 mm apart 0.0%); FK showing the action contract's nominal (a = 0) is itself a 465 mm straddle, so the action_rate and still terms were pulling toward the splits; biomechanics (feet tucked under the body cut peak hip-extension torque 148.8 -> 32.7 N*m, -78%); and HoST's foot-displacement term, which this reward table lacked. The earlier "not a reward hole" reading was corrected to "a gradient hole, not a level hole": at the dead point the heaviest term (upright) was saturated with zero gradient, still paid for not moving, and the one live gradient (base_height) pointed at the thigh-horizontal torque barrier. A prone ROM scan had already ruled out pushing up from prone.

    Change

    R0.4: feet_spread_excess = clamp(ankle distance - 0.215, 0, inf) x upright gate, weight -2.0, plus a height-decay factor added after measuring that the policy's real standing stance was 406 mm, not the 215 mm nominal (the plain version would have taxed every successful stand 0.38/s). R0.5: the same shape on the fore-aft axis, feet_fore_seated = |fore-aft offset - 0.05| x upright gate x height decay, target +50 mm (the measured natural offset of standing postures). One variable per rung.

    Outcome

    R0.4: seated ankle distance 561 -> 360 mm, supine/side exactly unchanged, prone 0 -> 1.9%, mid 45.9 -> 62.2%; a probe then showed the discriminant had moved to the fore-aft axis (standing starts +42 to +51 mm, the prone seat -168 mm). R0.5: supine 99.4, prone 99.4, side 100, mid 100%, re-falls 0%; both geometry terms collapsed to ~0 near iteration 13,100 as base_height rose, and the prone fore-aft offset went -168 -> +56 mm - the term's own target, closing the causal chain. The cost, unmeasured at the time: action jitter rose 33% (sum |da|^2 6.82 -> 9.06).

    Mechanism

    An upright-gated hinge is inert while the robot rolls and exactly zero in the achieved stance, so it adds gradient only inside the stuck basin; a seated robot with its feet behind or outside its COM must make a kinematically unfavourable transition to stand, and moving the feet under the body removes it.

    Applies when

    • a get-up or transition skill fails from one start category only
    • successful and failed episodes differ in a measurable geometric quantity
    • a shaping term might tax the posture successful episodes already use
    “`recovery_r0_5`,唯一变量 = 追加 `feet_fore_seated`(与 R0.4 同形状,只换测量轴)。 … 对照 R0.4 的 prone(3.1%,360 mm,**−168 mm**):前后偏移从 −168 走到 +56, 正是这一项的目标量,**判别量被消掉后成功率随之到顶** —— 因果链完整。”
    git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §20 R0.5(前后向脚位置项):成功率门全过 —— prone 0/159 → 158/159
  • Audit rewards by realized contribution (weight x achieved value) - a weight of 2.0 was really paying 0.04realized-contribution-audit
    Mechanism understoodwalkreward-shapingreward-shapingattribution

    Evaluate a reward table by each term's realized per-step contribution under the current policy, never by its weight column; if a term's realized value is ~0, escalating its weight is a no-op - change the term's structure instead.

    Symptom

    Foot dragging persisted through repeated weight escalation: feet_air_time had been raised 0.25 -> 1.0 -> 2.0 across versions with no behavioral change, and the training-side comment even recorded the fact ("几乎没有 单支撑相, 是拖着脚蹭") without the fix changing form.

    Context

    Computing realized per-term contributions in the trained state exposed the economy: feet_air_time contributed weight 2.0 x achieved 0.019 = 0.038 per step, against tracking's +1.20 - lifting the leg earned 3% of what tracking earned, so dragging was the rational optimum no matter the weight escalation. The same table acquitted the energy penalties (total -0.30 negative vs +1.74 positive) that a naive read of weights (-5.0 orientation!) would have blamed.

    Change

    Fix redirected from "raise the weight again" to "add a term whose realized contribution changes the optimum": a clearance penalty sized so its realized magnitude (~0.018/foot when dragging) is comparable to feet_air_time's, enough to flip the optimum without drowning tracking.

    Outcome

    With the term economy corrected (plus posture/range fixes), swing height reached 34 mm and tracking 87% by v6; weight escalation of the old term was abandoned.

    Mechanism

    A reward weight is only a multiplier on whatever the policy currently achieves on that term; when the achieved value is near zero (behavior absent), escalating the weight multiplies near-zero. Optimizer behavior is governed by realized per-step magnitudes, so audits must be conducted in that currency.

    Applies when

    • a behavior persists despite repeated weight increases
    • auditing whether penalties are "too strong" or rewards "too weak"
    • sizing a new reward term against existing ones
    “把 feet_air_time 权重从 0.25 → 1.0 → 2.0 一路加,但没有加高度项。量级算下来:feet_air_time 权重 2.0 × 实得 0.019 = 0.038,而跟踪奖励是 1.2。抬腿的边际收益只有跟踪的 3%,拖地当然是最优解。… 正项 +1.74,负项 −0.30。能量惩罚不是瓶颈,抬腿没收益才是。”
    train/WALK_DIAGNOSIS.md § 决定性证据(#1) / 各项奖励的实际量级

Next page