Training Coach
Doctrine
A report may cite any of these as doctrine-N.
doctrine-1Contract freeze and fingerprint disciplineThe policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.
Case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).
Coach application. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.
doctrine-2Attribution by resolved training params - never eval-override knobsCapability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.
Case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).
Coach application. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.
doctrine-3PASS gates become constraints; FAIL gates become objectivesOnce a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.
Case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.
Coach application. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".
doctrine-4One variable per ladder rung - counted against what the checkpoint sawA rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.
Case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.
Coach application. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.
doctrine-5Pre-register risks, readings, and stop criteria before the ladderBefore a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.
Case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).
Coach application. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.
doctrine-6Plant parameters are measured, never inventedEvery plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.
Case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).
Coach application. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.
doctrine-7Sim2sim gate before sim2real - under deployment conditionsEvery checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).
Case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).
Coach application. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.
doctrine-8Observation honesty - the actor's inputs are a hardware contractThe actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.
Case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).
Coach application. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.
doctrine-9Reward economics are audited in realized currencyReward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.
Case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).
Coach application. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.
doctrine-10The zero-cost option must be the desired behaviorFor every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.
Case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).
Coach application. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.
doctrine-11Measurement discipline: independent referees, signs, distributionsA disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.
Case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).
Coach application. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.
doctrine-12The deployment pipeline is plantIrreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.
Case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).
Coach application. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.
doctrine-13DR budget is finite; its distribution is the measured supportRobustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.
Case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).
Coach application. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.
doctrine-14Gates measure what hardware feels: posture, margins, stripped assistsAcceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.
Case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).
Coach application. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.
doctrine-15Fork and root selection: recoverability, maturity, frozen rewardsChoose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.
Case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).
Coach application. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.
doctrine-16Curricula: verified engagement, lineage counters, disease-phase gatingAutomatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.
Case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).
Coach application. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.
doctrine-17Probe before training: feasibility first, hypotheses in tablesAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.
Case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).
Coach application. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).
doctrine-18External advice is recomputed locally; values transfer as ratiosEvery external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.
Case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).
Coach application. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.
doctrine-19Hardware sessions are scripted experiments, not tuning sessionsReal-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.
Case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).
Coach application. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.
doctrine-20Close questions in writing; restart when the debt is structuralAudited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.
Case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).
Coach application. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.
doctrine-21Name the quantity in the space it lives inA goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.
Case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).
Coach application. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.
doctrine-22Continuation needs a live gradient; a release is chosen by a scanContinue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.
Case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).
Coach application. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.
Experience cards
169 cards matching “unpriced-foot-attitude-is-a-free-variable”.
Deployment power derating damages non-forward axes far more than forward - sweep it in sim before deploying
power-scale-hurts-nonforward-axesTreat deployment power/torque scaling as a plant parameter: evaluate the policy in sim at the exact deployment scale, expect non-dominant axes to degrade first under derating, and either deploy at the training power or train with power randomization.
Symptom
Policies deployed at power-scale 0.8 (a safety derating of commanded torque) looked fine walking forward but were weak at backward and turning, inviting the wrong diagnosis "the skill was not trained well".
Context
Measured repeatedly: on s1e, going 1.0 -> 0.8 cost forward 18% but backward 58%; on C4-ff800, turn tracking was +25%/+40% at pw0.8 vs +75%/+58% at pw1.0, backward 51-52% vs 97-103%, while forward stayed 96-98% at both. Sim evaluation numbers in the plan were all pw1.0, but the robot was being run at 0.8.
Change
Pre-deploy protocol added: sweep the exported policy across power in sim (for PW in 0.8 0.9 1.0: eval_c_matrix --power $PW --seeds 20) and deploy at the first level where both turn directions reach >=50%. For C4 the recommendation was raise the robot to pw1.0 - the sweep showed it nearly free (saturation 47%->33%, left foot-clipping danger zone 25%->6%, cost only tilt 6.7->8.3 deg).
Outcome
Turning "weakness" resolved without any retraining; the sim sweep correctly predicted the real-robot signature at both power levels.
Mechanism
Forward walking is the reward-dominant, torque-cheapest skill with the most margin; backward/turn/sidewalk live closer to the torque envelope, so a uniform torque derating consumes their margin first. Training ran at power 1.0 (the trainer does no power scaling), so deploying at 0.8 is a systematic underactuation the policy never experienced.
Applies when
- deploying with any torque/power derating or safety scale
- secondary skills (backward, turn, lateral) underperform on hardware while forward walking looks fine
- choosing the deployment power level for a new policy
“power 衰减对非前进轴的伤害远大于前进轴(s1e:前进 1.0→0.8 掉 18%,后退掉 58%)。转向是非前进轴,0.8 下很可能明显跟不动。”
train/C_LADDER_RUN.md § 3c. A-2 上机前先定部署力度档 / 3p. 二 Three hardware accounts locked the run design point - and the knee's real speed ceiling is tau_limit/kd, not the firmware limit
feasibility-accounts-lock-design-pointBefore opening a dynamic-gait training line, compute the full account set - tau_limit/kd effective speed ceilings, joint ROM under the intended reference geometry, and thermal RMS at the duty cycle - and let the accounts lock the design point; move only to pre-registered in-table alternates, re-running the accounts first.
Symptom
The run line was believed to require a firmware raise of the RS06 speed limit (10 rad/s) as a hard precondition, and the feasibility script's motor-envelope scan had marked 80/100 mm foot-lift cells "physically feasible".
Context
Three added accounts re-decided everything. (1) Damping tax: in MIT mode tau = kp*(q_des-q) - kd*qd, so sustained rotation is capped at tau_limit/kd = 12/1.5 = 8 rad/s - below the firmware's 10; at peak speeds 6.7-7.9 rad/s the damping term alone eats 10.1-11.9 N*m (84-99% of the torque limit). "提固件 limit_spd 越不过这道税 —— 它是 kd 与限扭的比,不是固件旋钮." (2) Joint ROM: the feasibility script had checked motor envelopes but NOT joint range - the ankle-pitch ROM caps 1:2:1 leg-shortening lift at 62 mm (soft) / 77 mm (hard), so the 80/100 mm "feasible" cells were voided; also firmware-independent. (3) Ankle thermal: duty 0.40 puts ankle RMS at 87% of continuous rating (0.35 -> 93%); long-period big-stride cells hit both ankle torque peak and heat. Verdict: firmware raise DEQUEUED (50 mm design point needs knee 6.7-7.3 < the 8 rad/s effective ceiling < firmware 10); vel_limit stays 10 so sim == robot. The three accounts uniquely lock the design point - 50 mm lift / T 0.60 s / duty 0.40 - "三笔账 唯一锁定,不是调参空间", with pre-registered alternates allowed only inside the table and only after re-running the accounts.
Change
Design point frozen from accounts; hardware precondition reversed by arithmetic rather than by test; reference amplitude (0.84 rad = FK inverse of 50 mm) derived, per-joint action scales sized to the required travel (knee 0.9, hip_pitch 0.6, ankle deliberately NOT amplified - hard limit is adjacent).
Outcome
A firmware work item left the critical path; an infeasible region of the design space was closed before any training; the remaining risk (knee tracking lag from the damping tax) was pre-registered with its own criterion and in-table fallback (duty 0.35) - "这不是'奖励没调好', 是 plant 账".
Mechanism
PD actuators in MIT mode pay kd*velocity out of the same torque budget that tracks position, so the effective speed ceiling is a ratio of configuration constants, invisible to firmware settings; and feasibility is the intersection of ALL constraint families (torque envelope, joint ROM, thermal RMS) - a scan that omits one family certifies impossible cells.
Applies when
- planning running/jumping or any high-rate gait on PD actuators
- a firmware or hardware upgrade is assumed as a training precondition
- a feasibility scan covers motor limits but not ROM or heat
“膝的有效速度顶 = τ_limit/kd = 12/1.5 = 8 rad/s,不是固件的 10。… 提固件 limit_spd 越不过这道税 —— 它是 kd 与限扭的比,不是固件旋钮。… 可行性脚本只查了电机包络没查关节 ROM —— 其 80/100mm 的"物理可行"格作废。… 判决:RS06 提固件对 run v0 不是前置,出队”
train/RUN_V0_SPEC.md § 1. 硬件账判决 / 2. 步态设计点 Narrowing the speed range to stop high-speed falls entrenched crouch-shuffling - judge gait quality at the speed that demands a gait
low-speed-commands-reward-draggingSet command ranges to include speeds that physically demand the target behavior, and evaluate behavior-quality gates at those speeds; when a restriction is added to suppress a failure, check what new optimum it creates at the remaining commands.
Symptom
After the command range was narrowed to (0.15, 0.35) m/s (to treat walk_v1's 134% overspeed and 8.3 s fall at 0.5), the policy settled into crouched foot-dragging; tracking rose monotonically with speed (63% at cmd 0.2, 76% at 0.3, 87% at 0.45), showing low speeds were where the degenerate gait was optimal.
Context
The narrowing advice was the author's own and is retracted in the file: it treated the symptom (falls at speed) while reinforcing the root cause (at 0.15-0.35 m/s, shuffling in a crouch is globally optimal - the Froude number is so low that even humans would not lift their feet). A zero-cost experiment confirmed the flip side: at cmd 0.5 the same policy met BOTH tracking (81%) and clearance (23.0/23.2 mm) standards.
Change
Speed range widened back toward (0.15, 0.5) - upper bound deliberately slightly above the mechanically feasible ~0.44 m/s so the policy finds the boundary itself; acceptance re-pointed: gait-quality criteria (tracking, clearance) judged at 0.45-0.5 m/s, low speed kept only as a survival check.
Outcome
v4 -> v6 progression under the widened range delivered 87% tracking with 34 mm clearance; the "low command = drag" account was confirmed by the monotone tracking-vs-speed curve.
Mechanism
Command distribution is part of the reward: physics prices gaits per speed, and at very low speed the energetic optimum is no swing phase at all. Restricting training to that regime makes the degenerate gait the correct answer to the posed problem - and grading a gait at a speed that does not require stepping measures nothing.
Applies when
- a gait degenerates after a command-range restriction
- quality metrics improve monotonically toward the range boundary
- writing acceptance criteria for gait quality vs survival
“现在看那个建议可能起了反作用:0.15~0.35 m/s 下蹲着蹭就是全局最优,抬腿反而亏。收窄治的是"高速摔倒"的症状,却强化了拖地的病根。… 验收标准里的 cmd 0.2 本身就是拖地速度(Froude 数极低,人在那个速度下也不抬脚)。accept_v2 应把速度跟踪与 clearance 的判定点改到 0.45~0.5 m/s”
train/WALK_DIAGNOSIS.md § ② 放宽速度区间 / ① 零成本实验 Push DR helped one lineage and hurt another at the same dose - robustness budget is conserved and gets borrowed, not created
push-dr-conditional-budget-conservationBefore opening a disturbance-DR rung, measure whether the untrained policy already meets the spec; if training it anyway, expect the benefit to be conditional on the lineage's existing DR load, grade the intensity, and audit retained margins - budget spent elsewhere will be borrowed back.
Symptom
The push rung's outcome flipped with the lineage: direct +/-0.6 m/s push failed outright on first attempt (base walking collapsed - kd1.2 scan 0/3 from iter 3300, sim2sim self-falls with pushes OFF - no PASS point existed); staged +/-0.3 then gave the narrow-kd single-working-point lineage real gains (push survival 1/5 -> 4/5) while the SAME dose made the dual-working-point balanced-band lineage WORSE (20-seed survival 18 -> 12/20 plus across-the-board push regression).
Context
The four-ladder verdict ("四梯定案", s2e/s2f at both intensities) named the pattern: "push DR 收益条件性" - the benefit is conditional on how much robustness budget the lineage has already spent. The law candidate: "DR 总预算守恒, 平衡带鲁棒性从抗扰余量借" - total DR budget is conserved; a lineage already covering a wide plant band pays for push tolerance out of its disturbance margin. Both S2 ladders therefore closed at the friction rung, with the decisive numerator: untrained push tolerance already met the 4-6 N*s requirement, so the rung was not needed at all ("⑥ push 不训(收益条件性,免训 ±0.6 已达 标)"). The same accounting later justified the C-before-S2 ordering ("push/μ 两轮已实证 DR 预算有限且会被重分配") and trimmed the second S2 pass to three rungs.
Change
Push removed from the standing ladder; graded intensity retained as the method IF a lineage ever needs push training; "does the untrained policy already meet the disturbance spec" instituted as the first check before opening any disturbance rung.
Outcome
Two rungs (push, ground mu) deleted from the second S2 pass on measured grounds; the ladder's real yield was re-stated honestly as precision, not robustness (speed gate 0 -> 20/20, zero-command drift 0.98 -> 0.06 m, but push 159 -> 125/160).
Mechanism
A fixed-capacity policy allocates representation and margin across the training distribution; adding a disturbance axis to a lineage that already spans a wide plant family forces reallocation - the new tolerance is bought with existing margins. Lineages with narrow plant coverage have free budget, so the identical DR dose lands as gain. Benefit is a property of (dose x lineage state), never of the dose alone.
Applies when
- proposing push/perturbation training on a hardened lineage
- the same DR rung helped one lineage and hurt another
- accounting where a ladder's robustness gains actually came from
“push DR 收益条件性 —— s2e⑥a (单工作点血统 kd 窄带) ±0.3 得抗推 1/5→4/5; s2f⑥ (双工作点平衡带血统) 同档反而 20-seed 存活 18→12/20 且抗推全面倒退。规律候选: DR 总预算守恒, 平衡带鲁棒性从抗扰余量借。两阶梯均以 ⑤ 摩擦级收官 … 抗推 4~6 N·s 免训已达标。”
train/OMNI_V0_SPEC.md § 4. ⑥ push 四梯定案 (2026-08-07) Add a termination that makes the degenerate strategy fatal - no height cut-off meant crouch-shuffling could live forever
termination-closes-degenerate-basinFor each known degenerate strategy, check whether the termination set makes it fatal; if the robot can live indefinitely inside the degenerate posture, add a termination just past the intended operating envelope rather than escalating penalties.
Symptom
Crouched foot-dragging survived indefinitely because the termination set contained only bad_orientation (40 deg) and base contact - there was no height termination at all, so a deep squat was a viable long-term strategy.
Context
The hypothesis audit found the missing termination (hypothesis 7); the cross-check against published configs found the field practice: Booster terminates at 0.45 m (38% of body height) and the research warning is that the termination height must not be so low that crouching survives it. The proposed value: 0.32 m, just below the walk crouch base height 0.3739 - a deep squat terminates immediately, "断掉蹲着蹭的活路" (cutting off the crouch-shuffle's livelihood).
Change
Add height termination at 0.32 m as a second-priority item of the walk fix package, alongside restoring base_height_l2 to -10.
Outcome
Entered the v5/v6 fix package under which the crouch-shuffle optimum disappeared (34 mm clearance, 87% tracking by v6).
Mechanism
Termination conditions define which strategies exist at all: a reward penalty prices a behavior, but a termination deletes its future returns entirely. Degenerate basins that are merely penalized can remain optimal under enough tracking pressure; a termination placed between the degenerate posture and the intended one makes the basin unreachable as a steady state.
Applies when
- a degenerate but stable behavior persists across reward tunings
- auditing termination conditions for a locomotion task
- a policy exploits the gap between penalized and terminated states
“加终止高度:研究第 6 条"终止高度不能低到让蹲着也能活"。我们完全没有高度终止。建议 0.32 m(略低于 walk 蹲姿基座高 0.3739,深蹲即终止)。… 加终止高度 0.32 m(深蹲即终止,断掉蹲着蹭的活路)”
train/WALK_DIAGNOSIS.md § 修正 ④ / 最终改动清单 第二优先 Train with self-collisions ON (filtering nested-link ghost pairs) - the reward wall prevents, the physics makes cheating impossible
self-collision-physics-plus-reward-wallNever train a contact-risk behavior with self-collisions disabled; enable them with an audited filter list for nested/overlapping pairs (zero contacts across a pose sweep), record the fps cost, and keep a calibrated distance penalty as the preventive layer on top.
Symptom
walk_v8 logged 107 frames of leg-on-leg contact while still earning 0.751 tracking score - because training-side self-collisions were OFF, leg clipping was literally imperceptible to the policy ("碰腿在训练里 根本感知不到").
Context
Enabling self-collisions naively is its own trap: an Isaac audit had shown PhysX auto-filters adjacent bodies (base-hip clean for free) but nested links generate ghost forces - calf and ankle_roll overlap 65 mm at the zero pose, producing 12x body-weight phantom forces. The v10 recipe: enable self-collisions, explicitly filter only the two nested pairs (l/r calf-ankle_roll), then run a zero-contact audit at three poses (nominal stand, walk crouch, swing-extreme) requiring contact count = 0, adding any residual pair to the filter and re-auditing; a 500-iter sanity run for NaN and an fps-cost record (measured -8.8%). Redundancy with the reward-side foot-distance wall was argued, not assumed: "N2 离得远(奖励侧预防),SC 碰了疼(物理侧兜底)" - the reward keeps distance at range, the physics makes contact hurt - so the v8-style "clip legs and still score" outcome becomes physically impossible.
Change
enabled_self_collisions=True + 2-pair filter + three-pose zero-contact audit (re-verified at 0.00 N after the later mass update) + fps budget recorded.
Outcome
Leg contact entered the training signal; the audit protocol caught the nested-pair ghost-force hazard before it corrupted training; combined with the calibrated distance wall, later versions held contact = 0 on hardware and in sim.
Mechanism
A hazard absent from the training physics cannot be learned about, no matter the reward; but collision meshes that interpenetrate at rest inject large fictitious forces if enabled blindly. Filtered enabling plus a pose-swept zero-contact audit gives true contact physics with no phantom energy - and layering prevention (reward) with consequence (physics) covers both learning and enforcement.
Applies when
- real robot self-contacts while training scored it healthy
- enabling self-collisions on a model with nested collision meshes
- deciding between reward-side and physics-side fixes for clipping
“PhysX 自动过滤相邻体(base↔hip_pitch 免费干净),幽灵力只在 calf↔ankle_roll(零位嵌套 65mm,12 倍体重)。… 与 N2 互补不冗余:N2 离得远(奖励侧预防),SC 碰了疼(物理侧兜底)—— v8 那种 107 帧互碰拿 0.751 跟踪分的事从此物理上不可能。”
train/WALK_V10_SPEC.md § 4. SC —— 训练侧自碰撞(范围已探明,比想象便宜) The run policy never left the ground and fell in the second simulator from the frontal plane - its DR (gains and latency only) covered the actuator axis, not the frontal-plane contact and inertia disturbances the doubled stride amplified; "is DR on" is the wrong question
thin-dr-judged-by-channel-coverageJudge a DR recipe by whether its randomized terms cover the channel where the skill can lose stability, not by whether DR is enabled; when a new skill lengthens single support or enlarges motion in one plane, add disturbances in the plane it destabilizes before training.
Symptom
run R1 (6,000 iterations, 78 min): no flight phase ever appeared, and every one of 13 checkpoints failed the eight-gate MuJoCo smoke. In Isaac: zero terminations in 6,000 iterations, 4.2 deg tilt. In MuJoCo at delay 2: 1/6 survived, falls within 1.9-6.2 s at 50.8-58.7 deg, the most saturated joints all roll joints.
Context
The run contract doubled sagittal travel (knee action scale 0.9, knee swing peak 1.14 rad) with a 0.60 s period and 0.40 duty - long single support - while roll/yaw scales were deliberately left at 0.5. DR copied the s1e recipe: kp/kd (0.9, 1.1) and latency on; mass, COM, joint friction and push all off; ground friction pinned at (1.0, 1.0). Flight was read two independent ways: Isaac's per-foot contact reward stayed 0.845-0.857, never above 0.87 - the arithmetic ceiling of a gait with zero flight - and 30 of 36 MuJoCo seeds had flight fraction exactly 0 (the nonzero six were all tumbling falls). Foot lift itself worked (46-59 mm against a 50 mm design point): the walk-era "not enough travel" failure did not recur.
Change
Verdict FAIL, with the pre-registered first knob (exploration noise 1.0 -> 1.2) explicitly rejected as aimed at a different axis. The lesson was generalized and applied at the next line's design review: the one-leg spec made push, body mass, base COM and friction DR mandatory for its permanent single support and banned the thin recipe.
Outcome
The run line did not continue past R1 in the sources. The one-leg V0 with the wider DR passed its friction-variant gate (mu 0.4 and 1.2) inside a 40/40 acceptance.
Mechanism
Randomizing gains and latency covers the actuator's axis; a skill whose failure lives in frontal-plane contact and inertia needs randomization on that channel (push, mass, COM, friction), or the trainer's exact plant becomes the only one the policy can stand on - the omni_s1 transfer trap a second time, this time with DR switched on.
Applies when
- a policy is flawless in the trainer and falls immediately in a second simulator
- reusing a DR recipe from a skill with a different support pattern
- failures concentrate on one axis (roll, yaw) the DR does not touch
“**机理**: 矢状面行程翻倍 (膝摆动峰 1.14 rad) + T 0.60 + duty 0.40 的长单支撑, 把额状面扰动放大了一个量级; 而 roll/yaw 通道按 §3 **刻意没有放大** (仍 0.5), DR 又是 s1e 复刻的薄配方 (mass/COM/关节摩擦/push **四关全关**, 地面摩擦钉死 (1.0, 1.0))。 … 说明**薄 DR 的判据不能只看"有没有开 DR"**, 要看**开的那几项 是否覆盖失稳所在的通道** —— kp/kd 与延迟是执行器轴向的, 对额状面接触/惯性 扰动零覆盖。 … 0.87 正是「零腾空的走路步态」的天花板算术”
git:Lucen V2@origin/run-line:train/README.md § run R1 FAIL (2026-08-09, run 21-30-30_run_r1): 腾空零, 但病根在额状面不在探索 Drop the frozen policy into chosen configurations - a squat 2.7 cm lower than the stuck pose stood 52% of the time, the stuck W-sit 0%, and the interpolation between them showed a wall, not a slope
configuration-probe-wall-not-slopeWhen a policy is stuck, probe the frozen policy from a grid of hand-placed start configurations, including interpolations between the stuck state and a nearby state it escapes from; one read-only experiment separates height, torque, sampling and configuration and tells you whether to prevent entry or train the exit.
Symptom
After R0.3 the policy stood from 62% of starts and never from the W-sit it fell into; height, torque, missing samples and reward were all plausible suspects.
Context
A read-only probe placed the R0.3 policy directly into specified configurations. Squats (hip, knee, ankle) = (-.65,-1.3,-.65) stood 100%, (-1.0,-2.0,-1.0) 89.8%, (-1.2,-2.4,-1.2) at 0.176 m 52.3%; the measured W-sit at 0.203 m 0.0%; the account-(3) hand-over state (146 deg tilt) 34.4%; linear interpolations from the W-sit toward the squat at 25/50/75% stood 0.0/0.0/3.1%. The squat family's quasi-static torque is 16% of the limits, and the W-sit was visited ~9 s per episode in training. FK showed the squat family (-a,-2a,-a) keeps the torso vertical, the feet flat and the COM over the feet all the way from 0.146 m to 0.384 m.
Change
Height, torque and sampling were eliminated in one experiment; the next rungs targeted entering the W-sit (foot placement) instead of escaping it, and seeding the dead point itself was ruled out because it was already visited every episode.
Outcome
Pure configuration: the W-sit (hips externally rotated +/-47 deg, knees folded 110 deg, shins flat, feet beside the body) is a different place from the sagittal squat (feet flat under the COM). The policy's standing skill was bound to a narrow sagittal family, and the wall was confirmed by the interpolation. The foot-placement rungs that followed took prone from 0/159 to 158/159.
Mechanism
A learned skill covers the neighbourhood of the states it succeeded from; a start state outside that neighbourhood fails regardless of height or torque, and an interpolation that stays at zero until close to a working state shows the boundary is sharp.
Applies when
- a policy stalls in a specific posture and several causes are plausible
- deciding between reverse-curriculum seeding and entry-prevention shaping
- a feasibility account says a path exists but the policy does not take it
“**决定性对比:比死点矮 2.7 cm 的蹲姿站立 52.3%,死点 0.0%。** 所以不是高度、 不是力矩(蹲姿族准静态力矩膝 1.96/12、踝 1.24/17,只占 16%)、也不是训练采样 (死点每局被访问 ~9 s)。**是纯位形问题** … 插值实验进一步显示这**不是坡是墙** —— 走到 75% 仍只有 3.1%”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §15 死点位形实验(只读探针,同一个 R0.3 策略放进指定位形) The standing-pose reward had been pulling toward the narrow stance the whole line was fighting - a zero-training kinematic audit of the target vector found it, after first auditing the wrong nominal
pose-target-geometric-auditBefore training on a pose target, audit it with forward kinematics - is it geometrically consistent (feet flat, intended stance, intended width) and is it the frame you think it is (action nominal vs standing default)? A posture term's target may itself be the attractor you are fighting.
Symptom
Several rungs aimed at widening the stance failed; the stance stayed narrow as if something kept pulling it back.
Context
The audit (MuJoCo forward kinematics, no training): every 5 deg of hip roll widens the stance ~5.5 cm (0.271 m at 5 deg, 0.383 m at 15 deg); at 47 deg of hip yaw a wide stance cannot be flat-footed (residual foot tilt ~0.7 x hip roll), which explained the stalled rungs. The first report also said the stand_pose nominal (hip roll 25, knee 60) has a 63 deg residual foot tilt - but that was the action frame's nominal (the limit-midpoint squat), which stand_pose never used, despite a docstring warning not to mix them. stand_pose's real target was DEFAULT_JOINT_POS: the contract's all-zero pose, legs parallel, ~0.22 m apart.
Change
The disease statement was corrected in writing: the narrow stance was not an accidental by-product of proxy traps but the target stand_pose had been actively rewarding. A stored "narrow the stance" knife was marked toxic. V2.8 moved the target to a flat 15-deg stance (sigma 3 -> 1.5, flat_feet margin 5 -> 20 deg).
Outcome
V2.8 still failed in-lineage (stance unchanged, feet nearly overlapping, mu 0.4 transfer 2%) - see stance-decided-by-get-up-path - and the from-scratch V3.1 removed both roll joints from stand_pose and put width into a task-space term, which is what finally produced a 0.355 m flat stance. The 63 deg finding was kept as a warning: an action nominal used as a standing target would be a ready-made pit.
Mechanism
A posture term with a sharp kernel around the wrong target is an active attractor; every other term fighting it pays twice.
Applies when
- a posture keeps returning despite penalties against it
- a reward uses a default or nominal pose as its target
- the contract has more than one "nominal" (action frame vs standing pose)
“上文"stand_pose 的 nominal (hip25/knee60) 残倾 63°"**审计错了对象**:那是 **动作参考系 nominal**(限位中点蹲),stand_pose 从未指向它(函数 docstring 原文即警告"两者别混",还是混了 —— 记档)。 … **修正后的病根陈述:窄站距不是代理陷阱的意外副产物,而是 stand_pose 一直在主动奖励的目标本身**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §44 勘误(实现时抓到):审计混了两个 nominal —— 真病根比误诊的更直白 The get-up kept getting faster because standing earlier paid more every step - lowering torque authority barely slowed it, and only zeroing the standing income for the first 3 s moved the pace into the design band
per-step-income-drives-speed-time-gateWhen a skill is too fast, find the term that pays for finishing early and gate that income by time; keep the "get into position" term ungated so the policy does not learn to wait, use a ramp instead of a cliff, and confirm with a paired same-level experiment that the drift is motivational before changing it.
Symptom
The user judged the get-up too fast (Isaac medians about 0.7-1.6 s) and suspected path dependence: the policy seemed to get faster the longer it trained.
Context
Lowering the beta authority 0.40 -> 0.30 cut impact but moved supine only 1.70 -> 2.00 s: coordination-limited, not torque-limited. A paired experiment inside one beta level (checkpoint 15,600 vs 18,499, +2,900 iterations, same ruler) measured the drift: get-up medians -7 to -10%. The spec concluded the motive, not the path, was the cause - per-step standing income pays for every early step, and any lineage (even one from scratch) races toward the fastest solution inside its constraints.
Change
V2.5: the standing income (base_height, stand_pose, still, feet_on_ground) multiplied by w(t) = clamp(t/3 s, 0, 1); upright deliberately NOT gated, so righting and sitting up early still pay and the policy is not taught to lie flat and wait; a ramp, not a step. V2.5b: zero before t0 = 3 s, then a 1 s ramp.
Outcome
V2.5: Isaac 100%, get-up +17-43% slower, MuJoCo 100/100/98/100% (the best cross-simulator reading yet), still short of the 3.5-4.5 s design band - a linear ramp only discounts early income. V2.5b: MuJoCo supine 2.04 -> 4.10 s and prone 3.18 -> 4.04 s, inside the band; the Isaac pace barely moved (a lineage habit on a gradient-free plateau). Later the zero gate proved harmful when trained from scratch (curriculum-history-is-part-of-the-product) and in the V3.1 lineage (time-gate-vs-wide-stance-retire-the-fix).
Mechanism
Constraints on authority or velocity change how the fastest solution looks; the time structure of the task income decides how fast the fastest solution is.
Applies when
- a policy is faster or more aggressive than wanted and constraints do not slow it
- progress-style rewards pay every step spent at the goal
- performance drifts faster with more training at fixed settings
“**V2.5 机制(唯一)**:站立收入(base_height/stand_pose/still/feet_on_ground) 乘时间斜坡 w(t)=clamp(t/T_gate,0,1),T_gate=3.0 s;**upright 刻意不门控** (翻正/坐直早期照常拿钱,防"躺平等门开" … **用户假设量化 证实:逐步计酬动机在 β 包络内持续压缩时间,约束挡不住动机 —— V2.5 动机层 修法为正解。**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §39 V2.5 预注册 / §39 补 配对实验 Prove a new penalty actually fires - two ways a clearance term silently did nothing
inert-reward-term-auditBefore training with a new reward term, log its realized per-step value under the current policy and confirm it is nonzero where intended - check coordinate zero-points against FK and check who occupies the term's gate; and never weaken the term that creates the states your new term needs.
Symptom
A newly designed swing-height clearance penalty could have trained as a no-op twice over, and the companion advice to lower feet_air_time actively backfired when tried.
Context
Instance 1 (zero-point offset): the proposed code used body_pos_w of the foot link, but that is the ankle_roll_link frame origin, which sits 0.0585 m above the ground even with the foot flat on it - so (0.03 - 0.0585) is always negative and the penalty is永远 0; the 0.0585 offset must be subtracted (verified identical in MuJoCo FK and Isaac). Instance 2 (gate occupancy): the clearance penalty fires only in swing phase; a dragging policy keeps both feet in contact, so the penalty is constantly 0 for exactly the policy it was meant to fix - and worse, any slight lift immediately incurs it, a reverse threshold. Lowering feet_air_time to 0.5 on that advice measurably collapsed air time to 0.0002 (below v3). Corrected understanding: "clearance 是把已有的摆动相抬高, 造出摆动相仍要靠 air_time" - air_time creates the swing phase, clearance raises it.
Change
Fixed the height zero-point; kept feet_air_time as the swing-phase creator with clearance layered on top; both errors documented as corrections to the team's own earlier advice.
Outcome
With both fixed, swing height rose from 22-23 mm (v2) to 29 mm (v5) to 34 mm (v6); the inert-term failure class entered the standing checklist.
Mechanism
A penalty's gradient exists only where its gate is occupied and its argument crosses its threshold; frame offsets shift the threshold out of reach, and phase gates can have zero occupancy under exactly the policy being treated. Terms interact as an ecology - one term must create the states in which another can act.
Applies when
- adding any gated or thresholded penalty (clearance, impact, slip)
- a new term produces no behavioral change at any weight
- body-frame positions are used in reward code
“body_pos_w 是 ankle_roll_link 坐标系原点,平放触地时仍高出地面 0.0585 m。… (0.03 − 0.0585) 恒为负 → 惩罚永远是 0 … clearance 惩罚只在摆动相生效,拖地时两脚始终触地 → 惩罚恒 0;而一旦轻微抬脚就立刻扣分,对正在拖地的策略是反向门槛。… 正确认识:clearance 是"把已有的摆动相抬高",造出摆动相仍要靠 air_time。”
train/WALK_DIAGNOSIS.md § walk_v4 独立验收 — 本文档给的两处代码/建议是错的 A 2-degree joint-zero calibration fix moved the whole runnable envelope - re-test old "cannot run" verdicts after recalibration
zero-offset-calibration-shifts-envelopeDate every hardware verdict with the calibration state; after any zero/mount recalibration, re-test previously condemned policy-power combinations and previously "unexplainable" posture offsets before attributing either to training or model.
Symptom
s1d was on record as "only runs at power 0.7" (kicked wildly at 0.8); after a calibration pass, the same policy ran 12 s at 0.8 with no kicking at all.
Context
The calibration had fixed a 2.08 deg zero offset on r_hip_roll - exactly the constant error source on the dominant joint of the kicking oscillation loop ("恰是乱踢振荡环主导关节的常值误差源"). The three-generation post-calibration hardware sweep also closed a second case: the robot's mysterious "backward lean" disappeared after calibration, and the sim-real posture difference collapsed from opposite-sign 5+ deg to same-sign ~2 deg ("后仰案实质了结") - the lean had been a sensing/zero artifact, not a mass-model error. Booked consequence: if the s1d recovery re-verifies, "真机可跑档整体 上移" - every policy's runnable power envelope shifts up, and downstream lineages' hardware expectations get revised.
Change
Joint-zero and mount calibration promoted from setup chore to a variable that dates hardware verdicts: verdicts about which power/scale levels a policy can run are conditioned on the calibration state they were measured under.
Outcome
One policy rehabilitated at a higher power level; one standing sim-real posture discrepancy closed without touching model or training; a pending re-verification booked rather than asserted.
Mechanism
A constant joint-zero error acts as a persistent disturbance injected at the feedback loop's most-loaded joint; near an oscillation threshold, removing a 2-degree bias is the difference between a stable and an unstable loop. Since the error is additive and machine-side, it shifts every policy's stability envelope simultaneously - which is why verdicts must carry their calibration date.
Conflicts
The s1d rehabilitation awaited one confirming re-run at the time of writing ("待复核一跑坐实") - the offset-as-cause reading is the head suspect, not a closed verdict.
Applies when
- a policy oscillates at a power level others tolerate
- sim and real disagree on a constant posture offset
- deciding whether to re-test old hardware verdicts after maintenance/calibration
“发现①:s1d@0.8 能跑了(旧账「只有 0.7 能跑」)——12s 无乱踢。头号嫌疑 = 标定修正:r_hip_roll offset 修 2.08°,恰是乱踢振荡环主导关节的常值误差源。… 发现②:「后仰」标定后消失 … sim-real 姿态差从反号 5°+ 收敛到同号 2°,后仰案实质了结。”
train/README.md § 真机 @0.8 三代横评(2026-08-07 标定后) Friction DR was demoted after a measurement (94% success at mu 0.4 with no friction randomization) and promoted again when the action contract changed and mu 0.4 fell to 76% - DR priorities belong to a plant and contract, not to a task
friction-priority-re-measured-after-plant-changeRe-measure transfer along the friction axis for every new action contract or plant, not once per task; a DR priority settled under one action parameterization does not carry to the next.
Symptom
Getting up is all scraping and pushing against the ground, and training pinned friction at 1.0, so friction looked like the first thing to randomize.
Context
The MuJoCo gate on R0.5 (5 categories x 10 seeds x 4 friction levels) measured 100/100/98/94% at mu 1.0/0.8/0.6/0.4: degradation showed first as time (prone 3.2 -> 5.3 s), not failure, so friction DR was demoted and the DR budget earmarked for mass/COM. After the switch to the beta-anchored action space, V2.2 read 90/94/90/76%: mu 0.4 was now the weak row.
Change
V2.3 (single variable): friction DR static (1.0, 1.0) -> (0.2, 2.0), dynamic (0.15, 1.6), the HiFAR range keeping the base dynamic/static ratio; restitution untouched. Continued from v2_2.
Outcome
Isaac nominal 99.8% (DR did not hurt the nominal plant); MuJoCo 98/98/96/92% - mu 0.4 76 -> 92%, mu 1.0 back to R3.1's 98% with bounded torque.
Mechanism
How much a policy leans on friction depends on how it moves; the spec records that the sensitivity rose after the action contract changed but does not establish why.
Applies when
- changing the action space, gains or authority of an existing skill
- deciding which DR axis to spend the next rung on
- an earlier sweep justified leaving an axis unrandomized
“**μ 砍到 0.4(训练值的 40%)仍有 94%**,退化先体现在**用时**(prone 3.2→5.3 s) 而不是成败。μ≥0.8 完全无损。→ **§17 曾把"摩擦随机化提到 R4 第一项"当作优先 事项,这条实测把它降级了**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §21 MuJoCo 复核门 ② 摩擦依赖 Run acceptance under measured contact parameters - honest condim/torsional-friction flipped a false PASS into a real-matching FAIL
eval-plant-honesty-contact-paramsPin the evaluation plant's contact model to measured values (contact dimension, torsional/rolling friction, mu) before trusting any gate that involves slip, impact, or drift - a gate can only fail a policy for physics its simulator contains.
Symptom
walk_v5 passed the old acceptance battery yet failed on the real robot (footfall force, drift, kicking) - the evaluation plant was flattering the policy.
Context
The battery was re-run under "honest contact parameters" - condim 4 (adding torsional contact) with measured torsional friction 0.035 - and v5 then FAILED exactly the rows corresponding to its real problems: heading 185 deg (limit 30), support-foot yaw slip 284 deg (limit 80), landing force 1.72x (limit 1.5x), suspended tilt 45.9 deg (limit 10). The slip physics depends on torsional friction, which the default contact model (condim 3) does not even simulate - a slip problem is invisible to an evaluator that cannot represent yaw friction at the foot. Term-sizing measurements for the new rewards were likewise taken under the same honest parameters (cmd 0.45, skipping the 5 s start transient).
Change
Acceptance harness pinned to condim 4 / torsion 0.035 (measured); verdicts issued under defaults declared non-citable for these rows.
Outcome
Sim acceptance verdicts began agreeing with hardware ("现在失败, 与真机一致"); the v6 fixes could be developed and validated against an evaluator that could actually see the disease.
Mechanism
An evaluator is a plant model too: contact dimensionality and friction values decide which failure modes exist in the simulation at all. Evaluating under default contact parameters tests the policy in a world where its real failure is physically impossible, producing structurally false PASSes.
Applies when
- sim acceptance passes policies that fail on hardware
- slip/drift/impact gates run under default simulator contact settings
- setting up a cross-simulator evaluation harness
“accept_v2.py 已加三条判据, walk_v5 在诚实的接触参数下(--condim 4 --torsion 0.035)现在失败, 与真机一致:直行 15s 航向累计 <30° | 185° ✗ … 落脚力峰值 <1.5× 体重 | 1.72× ✗”
train/WALK_V6_MINIMAL.md § 5. 验收 The walk phase machine structurally cannot express flight - rebuild the representation for duty < 0.5, teach flight with a mask tax, never a cliff bounty
phase-machine-structural-limitsWhen a new gait changes the contact pattern's structure, audit whether the phase/mask representation can express it and rebuild the representation if not; teach the new contact pattern with graded mask-mismatch pressure and count it in acceptance with artifact-proof definitions (minimum segment length), never with cliff bounties.
Symptom
Running requires both feet airborne, but the walk-era phase machine switches legs by the sign of sin(phase) - with duty < 0.5 the two swing windows must OVERLAP during flight, which a sign-switching representation cannot express at all.
Context
The run phase machine was re-architected rather than patched: per-leg phases (left = phi, right = phi+0.5 mod 1) with leg_phase < duty defining stance, aligned to the walk sin convention at duty=0.5 so the machines agree where their domains overlap. Flight is taught by the SAME mechanism that once cured foot-dragging, direction reversed: in the two planned flight windows the contact mask is (0,0) and feet_contact_number_duty charges -0.3 per foot still on the ground - a mild ~0.16/step tax, deliberately NOT a cliff: "悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文". The reference shape (half-sine bump over swing progress) is zero at window boundaries by construction, eliminating the clearing-window that the C4 probe measured to cost 13-19% on non-sinusoidal references. The shape self-check ("零代价选项是什么") was run on three behaviors: standing pays ref everywhere (known cmd=0 stepping risk, booked), walking pays only the flight-window tax, proper running collects full marks.
Change
New rewards.py run section (leg_phase_duty / stance_mask_duty / ref_run / clearance_run / contact_duty) with the walk versions untouched byte-for-byte; flight acceptance metric defined with a segment-length floor (>=40 ms to count) so numeric contact flicker cannot fake flight.
Outcome
Flight became expressible and taught by a calibrated mild pressure; the walk lineage's phase code stayed frozen as its own contract.
Mechanism
A phase representation defines which contact patterns exist in the reward's vocabulary; duty cycling below 0.5 introduces states (double-flight) outside a half-period sign convention's language, so no weight tuning can teach them. And rare desirable events taught by cliff-shaped bounties invite hacks (jumping in place); the graded mask tax prices the planned pattern without creating a jackpot.
Applies when
- extending a walking stack to running/jumping (duty < 0.5)
- a desired contact pattern never appears despite reward increases
- defining flight/contact acceptance metrics
“duty<0.5 时摆动窗 (1−duty)T > T/2,两腿摆动窗在腾空段重叠 —— sin 符号切腿的机制结构上表达不了"双脚同时在空中"。… feet_contact_number_duty 对"还踩着地"持续 −0.3/脚 —— 与 walk 治拖地同一机制,方向相反。不腾空的税 ~0.16/步 … 梯度温和不构成悬崖(悬崖式腾空奖励诱发跳跃 hack,v4-clearance 家族老课文)。”
train/RUN_V0_SPEC.md § 5. 相位机设计 Oversized lateral COM randomization (+/-5 cm) deliberately forces leg spread
com-randomization-forces-leg-spreadDR ranges can be behavior-shaping tools, not just robustness padding: oversize a randomization axis to force a strategy the reward struggles to express - and expect a compensating behavior to appear as the cost.
Symptom
Feet drift toward the centerline and even collide; policy has no incentive to keep a lateral support base.
Context
COM randomization ranges were chosen asymmetrically by axis: lateral +/-5 cm ("比常规大,故意的" - larger than usual, on purpose), fore-aft +/-2 cm, vertical +/-2 cm. The oversized lateral range is not robustness padding but a behavioral forcing function. Lucen logged it as directly relevant to its own roll-channel / sideways leg-kick symptom.
Change
Set COM randomization to lateral +/-5 cm, fore-aft +/-2 cm, vertical +/-2 cm, with the lateral band intentionally oversized to make narrow stances fail during training.
Outcome
Effective at separating the feet on the reference robot; side effect - the base began swaying left-right, which then required a foot-centerline distance penalty (see reward-chain-foot-height-landing-spacing).
Mechanism
Randomizing COM laterally makes narrow-stance policies fall for some draws, so PPO discovers wide stances as the only strategy robust across the band - DR used as an implicit reward. The sway side effect appears because the policy hedges against unknown COM by active lateral correction.
Applies when
- feet too close / self-collision in a learned gait
- roll-axis instability suspected to come from narrow stance
- choosing COM or mass-offset DR ranges
“两脚太近甚至互撞 → 先试质心横向随机化 ±5 cm,逼迫策略把脚分开;有效但引发新问题——基座开始左右摇摆 … 横向 ±5 cm(比常规大,故意的,用来逼出分腿)/ 前后 ±2 cm / 垂直 ±2 cm”
Experience.md § 质心随机化范围 (lines 75, 84-86) A joint frozen at the action clamp pays zero action_rate forever - penalize pre-clip saturation to make the cheat cost money
saturation-cheating-zero-rate-costWhenever actions are clipped and any smoothness/rate penalty exists, add a pre-clip saturation penalty so living at the clamp costs more than oscillating - and audit for frozen-at-clamp joints (action std ~0, |a| at exactly the clip value) as a standing acceptance row.
Symptom
With action_rate_l2 raised to -0.2, walk_v7's hip_pitch actions froze at exactly +/-1.000 (the clamp), reproduced bit-for-bit on hardware (splits frozen at +/-0.35 rad); the gait-shaping term joint_pos_ref collapsed to 0.026-0.035. The repo had died in the same trap once before (walk_v0: four joints pinned at +/-1.0).
Context
Mechanism: a joint pinned at the clamp has action-rate cost exactly zero and forever zero - under a strong smoothness tax, "push to the clamp and freeze" becomes the dominant optimum. Lowering the weight (-0.2 -> -0.1) only reduces temptation; the frozen state still costs nothing, so the structural fix adds action_saturation = sum(relu( |a_raw| - 0.9)) at weight -1.0, computed on the PRE-clip network output - post-clip, |a|=1.01 and |a|=3 punish identically and the out-of-range gradient dies (v0's old disease: mean |a| 1.71 soaked in saturation). Economics: freezing at |a|=1.0 now pays 0.1/joint/step (two hips = 40% of alive) vs ~0.0004/step for the healthy reference oscillation - the cheat flips from free to ~250x negative. Honest limits were recorded: A1 does not forbid freezing at 0.89 (the anti-freeze pressure must come from the oscillation demand of joint_pos_ref), and the alternative "rate on post-clip target" was rejected as 换汤不换药 - a pinned target also has zero rate.
Change
v8-A: add action_saturation (-1.0, thresh 0.9, pre-clip) AND halve action_rate_l2 (-0.2 -> -0.1, still 3.3x the v5 value); success criterion pre-declared (joint_pos_ref telemetry returns to v6 scale).
Outcome
Booked as the structural repair of the v7 freeze; also fixed a config hygiene trap discovered on the way - action_rate was assigned twice in __post_init__ (v5 comment line then v7 line), merged to one assignment "别再留两处赋值给下次审计埋雷".
Mechanism
Clipping creates a zero-gradient, zero-cost absorbing region in action space; any penalty on action derivatives makes that region strictly optimal once entered. Only a penalty on clamp proximity itself (measured pre-clip so depth of violation is visible) restores a slope out of the absorbing region.
Applies when
- joints sit at exactly the action clip with near-zero variance
- raising a smoothness penalty degrades gait amplitude
- shaped-oscillation terms collapse after a rate-weight increase
“钉死在钳位的关节 action_rate 代价精确为零且永远为零;−0.2 之下"推到钳位冻起来"成了压倒性最优 … 本仓第二次栽在同一坑(walk_v0 死于四关节钉死 ±1.0)。回调权重(−0.2→−0.1)只降低诱惑不消除作弊 … 算在 clip 前的原始网络输出上 … 作弊收支从"白赚"变成"倒贴 ~250 倍"。”
train/WALK_V8_SPEC.md § 1. 改动 A — 治饱和作弊 Under continuous 3-axis uniform sampling, pure straight-line walking is a zero-measure event the policy never trained
zero-measure-commands-need-mode-samplingEnumerate the exact command points users will actually issue (straight, stop, in-place turn) and give each explicit probability mass via mode sampling with off-axes pinned to zero - never assume a continuous sampler covers its measure-zero subsets.
Symptom
"The robot drifts even in sim when told to walk straight" persisted across reward tunings - because with commands drawn as vx in [0.15,0.5] x vy ~ U(+/-0.2) x wz ~ U(+/-0.6), the event vy=0 AND wz=0 has probability zero: pure straight-line walking was never sampled even once.
Context
Restart evidence item #3: "纯直行是零测度点 … 'sim 里直行就漂'是分布的 必然,不是 reward 没调好" - the drift metric was legitimately drowned by commanded turning (v11's own comment self-documented this). The structural fix is discrete mode sampling: a custom ModeVelocityCommand that first draws a mode by share (stand/forward/back/turn/side/mixed), then draws values only on that mode's axes with all others pinned to exact zero - which is also what preserves single-variable discipline in the C ladder (native 3-axis uniform "采不出'离散模式桶' … 把 C1~C4 的单变量纪律直接毁掉"). The mixed mode later got an ellipsoid constraint rather than a cube for the same reason in reverse - corner combinations of a cube are unrepresentative extremes.
Change
Command generation moved from independent per-axis uniforms to mode-bucket sampling with pinned-zero off-axes (plus 20% rel_standing); acceptance likewise evaluates per mode.
Outcome
Straight-line behavior became a trained, testable mode instead of a measure-zero hope; the C ladder could add one mode per rung with provable isolation.
Mechanism
A policy optimizes expected reward under the command distribution; events of probability zero contribute nothing to the objective, so exact-zero-command behaviors (straight walk, stand, in-place turn) are only learned if the sampler gives them mass. Product-of-uniforms distributions concentrate mass on mixtures and give none to the pure behaviors users actually command.
Applies when
- a "simple" command (straight, stop) underperforms mixtures in sim
- designing command distributions for velocity-tracking tasks
- a ladder needs per-mode isolation for attribution
“纯直行是零测度点:最终 command 为 vx∈[0.15,0.5] × vy∈U(±0.2) × wz∈U(±0.6) 连续均匀,vy=0∧wz=0 从未被专门采样 —— "sim 里直行就漂"是分布的必然,不是 reward 没调好 … Isaac 原生 UniformVelocityCommand 是三轴各自 uniform,采不出"离散模式桶"”
train/OMNI_V0_SPEC.md § 0. 为什么从零 (3) / 三件前置 (1) Curriculum-gate a penalty to the phase where its disease occurs - early on it only taxes exploration
gate-penalties-to-the-disease-phaseFor penalties aimed at late-stage pathologies (freezing, saturation, degenerate attractors), ramp the weight in only after exploration noise has decayed; anchor the terminal weight to measured healthy-vs-sick raw values, and shift all related tripwires to after the ramp completes.
Symptom
The action_saturation penalty, applied from iteration 0 in v8a, taxed exploration itself: with init_noise_std 1.2 the sampled actions paid ~-2.45/step before any policy had formed - while the disease it targets (clamp freezing) is a LATE pathology (v9 froze at iteration ~2624).
Context
v10 re-introduced the same penalty behind a curriculum gate: weight 0 until iter 1000, ramping linearly to -1.0 by iter 2000 - present only when the disease can occur, absent while exploration noise dominates. The trust argument was evidence, not hope: in v8a the term, while active, had pulled joint_pos_ref from 0.041 up to 0.155 and climbing - proof it can extract a policy from the frozen pit. Weight magnitudes were anchored to measured raw values (healthy v5 0.310 / v6 0.106 vs frozen v7 1.145 / v9 1.22 per step: at -1.0 healthy pays 6-18% of tracking, frozen pays 60%+, standing ~0). v10c then isolated the gated term as THE anti-freeze mechanism by single variable, upgraded to untouchable status in v11: "S 的门控机制(v10c 单变量铁案:任何情况下 不许撤,只许调终值)" - and v11 dared to relax other penalties only because S stood guard.
Change
action_saturation gated 0 -> -1.0 over iters 1000-2000 (later terminal value tuned -1.0 -> -0.5 with the gate mechanism itself frozen); tripwires adjusted to respect the gate's timing (freeze check moved to iter 2500-3000 to give the ramped term its effect window).
Outcome
Freezing stopped recurring while early training kept full exploration; the mechanism graduated from experiment to invariant within two versions.
Mechanism
A penalty's incidence depends on who occupies its support: early in training that is exploration noise (whose suppression starves learning), late it is the converged pathology. Time-gating aligns the penalty's presence with its target's presence, buying the constraint without the exploration tax - and tripwire timing must then be computed from the gate schedule, not from ungated precedents.
Applies when
- a structural penalty punishes exploration in early training
- a late-onset pathology (freeze/saturation) needs a standing guard
- deciding when a curriculum ramp should engage
“v8a 实锤它的病根是"罚在采样动作上"——init_noise_std 1.2 的早期等于罚探索(~−2.45/步);而冻结是晚期病(v9 速率 2624 才死平)… 门控让它只在病发期在场。… v8a 里它在场时 joint_pos_ref 从 0.041 爬到 0.155 且仍在升——有从低谷爬出的实证。”
train/WALK_V10_SPEC.md § 2. S 保险 —— action_saturation 课程门控 Gate a new reward term by its command so all old modes score pointwise identical
gate-new-reward-terms-by-commandWhen a reward term must be added mid-lineage, gate it on the condition that defines the new task so every pre-existing situation scores exactly as before - and still watch for value-rescale pathologies inside the new mode.
Symptom
Adding a lateral tracking reward (track_lin_vel_y_exp) ungated would have paid 0-2.0 per step even in modes with cmd_vy = 0 (healthy gait sway of vy ~0.1 already earns 1.28), shifting the whole reward table by a large bias and rescaling the value function - no longer "just adding one mode".
Context
C4 was the C ladder's only true reward surgery. Single-variable discipline required that the change be invisible to every existing mode. The chosen construction: gate_by_cmd=True - the term pays only when |cmd_vy| > 0.02, so for all modes with cmd_vy == 0 the term is pointwise zero, i.e. the reward is pointwise identical to before the change. The same trick appeared earlier in C1: replacing the vy L2 tax with a command-error version that is "对 cmd_vy≡0 逐点同值" (pointwise equal when cmd_vy is 0), explicitly classified as not-a-reward-change.
Change
track_lin_vel_y_exp added with gate_by_cmd=True (weight +2.0, std 0.15); the residual acknowledged honestly - inside the side bucket the values DO change, so the rung still watched the known reward-reshuffle pathology signature (s1c B-arm: scatter -> half-recover -> collapse) as a stop criterion.
Outcome
Old modes provably unaffected (pointwise-equal argument); attribution for any change in old-skill metrics stayed clean through the C4 redo series.
Mechanism
PPO's critic normalizes to the reward scale it sees; an ungated additive term shifts returns in every state and re-scales advantages globally, entangling the new skill with all old ones. Command-gating confines the new term's support to the new mode's state distribution, making "pointwise identical elsewhere" a provable property rather than a hope.
Applies when
- adding a tracking/shaping term for a new command or skill to a lineage that must not regress
- reward change proposed while other skills are still being gated
- reviewing whether a config diff counts as a reward change
“只在 |cmd_vy| > 0.02 时付。不门控的话它对 cmd_vy≡0 的老模式也给 0~2.0 分(健康摇摆 vy≈0.1 → 1.28),等于给整张奖励表加一个大偏置、值函数尺度全变 … 门控后老模式逐点得 0 = 与加项前逐点同值,单变量纪律成立。… 但 side 桶内的值确实变了 —— 这仍是奖励表改版,开级盯 s1c B 臂签名”
train/C_LADDER_RUN.md § 3d. gate_by_cmd=True(重要) A walking policy's tilt cutoff is a legal state for a recovery policy - the default 45 deg fall guard had to be raised for recovery tests and is disabled once the switch owns falls, so the abort chain becomes the recovery timeout, the operator's cut, and the firmware torque limits
fall-guard-becomes-a-stateWhen a new skill makes a safety cutoff's trigger a legal state, replace the cutoff with a bound of the skill's own (a timeout ending in a safe stop) instead of just switching it off; keep the operator's cut and the firmware limits as independent layers, and write every flag change into the run sheet.
Symptom
deploy_policy's default protection stops the robot beyond 45 deg of tilt. A recovery policy starts lying at roughly 90-97 deg, so under the default it is refused on the spot - a flag the first hanging checklist forgot.
Context
The layers in the sources: deploy_policy's tilt cutoff (default 45 deg, a line in the safety chain); for standalone recovery tests the cutoff was raised (110 deg in the spec's A/B sheet; 181 deg, effectively off, in some runbook commands); with --recovery-policy the cutoff is disabled because a fall is now a state, not an exception, and RECOVERY lasting over 15 s ends in a safe stop (the runbook calls it the line where the spotter steps in). Independent of the policy: firmware torque limits checked at start (12/17/11 N*m, set_torque --check), the operator cutting enable at any kicking or oscillation, and in the one-leg teleop a space-bar stop that puts the foot down.
Change
The flag was added to the run sheets, and the FSM replaced the removed cutoff with its own bound (the timeout).
Outcome
The spec records the flag omission and its fix; it does not record the FSM's timeout being exercised on hardware.
Mechanism
A safety cutoff encodes one policy's notion of "abnormal"; a new skill whose normal operation lies beyond it either cannot run or runs with the cutoff off, and only a replacement bound keeps the chain closed.
Applies when
- deploying recovery, fall-damage or acrobatic skills behind existing safety checks
- a run sheet disables a protection flag
- listing the abort chain for a hardware session
“--max-tilt-deg(默认 45°,安全链第 13 行写的那个)。recovery 的合法状态覆盖整个倾角域,把它抬到 181 = 实效关闭 … RECOVERY 超时 15s 会自动安全停(看护介入线)”
RL系统/FOLLOW THIS copy 2.md § FSM 吊挂首测 ② 落地测 / #### Recovery Policy (operator runbook, undated) A real-robot verdict is (policy x deployment stack) - when the stack changes materially, old verdicts expire
stale-verdicts-under-old-stackDate every hardware verdict with the deployment-stack version it was measured under; after any material stack change, re-test before trusting old condemnations or old praises - with the interpretation of each possible result written down first.
Symptom
walk_v5 stood condemned as "kicks wildly" and walk_v6 as "cannot walk unassisted" - but those verdicts were issued under an earlier deployment stack (--heading did not exist yet, several fixes had just landed); only v7 had ever run under the current unified stack.
Context
New sim evidence sharpened the doubt: a same-harness four-version sweep showed v6 was the HEALTHIEST archive at the real operating point (cmd 0.15: dominant frequency locked at 2.50, steps 35:35 perfectly symmetric, foot distance 203/190 mm best of four, mean tilt 4.2 deg, saturation 15%). A full re-test under the unified stack was scheduled with per-version questions and a pre-filled interpretation table ("判读表(预填假设,回来对号)"): e.g. v6 walks + frequency ~2.5 -> old verdict was the stack's fault, v6 becomes the comparison champion; v6 walks but at ~1.25 -> period-doubling on hardware = confirmed plant gap, actuator fitting promoted to mainline; v5 no longer kicks -> the kicking was an old-stack artifact.
Change
All four versions re-queued on hardware under one stack (same torque limits, slew profile, heading loop, logging), with the version-specific legacy profile pinned; verdicts held provisional until re-issued.
Outcome
The re-test design separated policy properties from stack artifacts before any policy was permanently written off - and turned each outcome into a specific conclusion via the pre-filled table.
Mechanism
A deployed behavior is produced by the policy plus everything between it and the motors (heading loop, slew limits, torque caps, clock); verdicts implicitly condition on that whole stack. Fixing the stack invalidates the conditioning, so old failures may be stack artifacts and old successes may not survive either.
Applies when
- deployment tooling (limits, filters, loops) changed since a policy was last judged
- deciding which historical policy is the rightful baseline
- a sim sweep contradicts an old hardware verdict
“只有 v7 在完整的今日部署栈下上过真机 … v5"左右乱踢"、v6"未能自主"的判决全部来自更早的栈(--heading 尚不存在, 部分修复刚落地)——判决已过期。且 2026-08-02 四代同机仿真横测翻出了新证据:v6 在真机工况(cmd 0.15)下是四代里最健康的仿真档案”
train/REAL_SWEEP_V5_V8.md § 0. 为什么重测 A time gate that had cured one lineage's rushing made the from-scratch lineage trade away its stance width twice (0.364 -> 0.235 m, 0.355 -> 0.251 m) - its disease was absent there, so the fix was retired and the pre-gate checkpoint shipped
time-gate-vs-wide-stance-retire-the-fixCarry a fix into a new lineage only if its disease is present there; a mechanism that cured one lineage can be net negative in another, and when doubling a term's weight recovers almost nothing, treat the two objectives as structurally in conflict and remove the one whose purpose is gone.
Symptom
V3.1's phase 2 (the 3 s zero gate on standing income, continued from P1b) kept 100% success on every friction level and slowed the get-up, but the lateral stance drifted 0.364 -> 0.235 m and hip yaw crept to 57 deg against its 60 deg limit. With the width band's weight doubled (P2c, after P1c) it drifted again, 0.355 -> 0.251 m, below the pre-registered 0.30 m failure line.
Context
The zero gate had been introduced in V2.5/V2.5b to slow the old lineage's get-up. In V3.1 the rushing was already absent: P1c got up in 0.90-1.06 s with a worst torque ratio of 73.3%, better than the stamped v2_6c, because the full beta curriculum, second-difference smoothing and pull curriculum had cured the violence inside training.
Change
Recorded as a candidate law with two data points - the zero gate and a wide stance are mutually exclusive here - and the zero gate was removed from the V3.1 recipe. P1c (the pre-gate checkpoint) went through the full stamp-level acceptance instead.
Outcome
P1c passed everything: all six criteria, lateral stance 0.355 m, foot tilt P75 2.0 deg, mu {1.0, 0.8, 0.6, 0.4} x 10 seeds all 100%. recovery_v3_1p1c.onnx was stamped and pushed to the robot channel.
Mechanism
The zero gate moves the income toward "stay stable until the end", and under low-friction DR a wide stance has a slip tail, so survival outbids the width band; doubling the band's price bought back only 0.016 m - an auction that does not converge signals structural conflict, not an under-priced term.
Applies when
- porting reward mechanisms from an old lineage into a fresh recipe
- a width, margin or posture metric erodes during a late training phase
- a weight increase produces a negligible change in its target
“**定律候选(二实证):归零门 × 宽站互斥**。 … 加价翻倍只挽回 0.016,竞拍不收敛)。 … **归零门是 v2_5 血统的历史包袱,对 V3.1 配方是净负资产,P2 阶段除名 —— P1c 即终点形态**。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §49 终验(2026-08-14) A frame-history observation under zero DR memorizes the trainer's plant fingerprint - the estimator must see variation to learn estimation
history-obs-needs-plant-variationIf the observation carries history (stacked frames, RNN), keep at least minimal plant variation (gain/latency jitter) on from the first iteration - "nominal first, robust later" is structurally invalid for estimator-bearing contracts.
Symptom
omni_s1 (fresh 215-dim contract with a 5-frame history window, trained with DR fully off): training all green, yet the MuJoCo gate scored 0/20 on all eight doors - falls within 2 s, seven checkpoints, not one transferred.
Context
The history window exists precisely to let the actor implicitly estimate line velocity and actuator dynamics (the actor is denied base_lin_vel by observation honesty). Under a constant plant that implicit estimator has nothing to estimate - it learns the trainer's exact response fingerprint instead, and any other simulator's micro-differences are out-of-distribution: "5 帧窗按设计就是隐式估计器, plant 恒定时它学到 Isaac 精确响应的指纹". The planned "nominal-first-robust-later" staging was declared STRUCTURALLY incompatible with history observations: "估计器要见过变化才学估计, 否则学背诵" (an estimator must see variation to learn estimation, otherwise it learns recitation). Honest confound note kept: this is mixed with "zero DR does not transfer, period" - but both attributions prescribe the same fix, so no control was run.
Change
S1.1: minimum actuator jitter turned on from day one - kp/kd +/-10%, latency 0-1 frame (friction/COM/mass still nominal, no push - those stay for the S2 ladder).
Outcome
Transfer restored: survival 0/20 -> 20/20, speed 19/20, foot distance 20/20 (remaining failures moved to gait quality, a different disease); the staging doctrine was amended - history-carrying contracts never train under a frozen plant.
Mechanism
A recurrent/history channel fits whatever temporal structure minimizes loss; with a deterministic plant the cheapest structure is the plant's own impulse-response signature, yielding features that are simulator-specific rather than physics-general. Plant variation forces the channel to carry state-estimation features that transfer.
Conflicts
Attribution is explicitly confounded with the simpler "zero DR never transfers" reading ("与「零 DR 本身就不迁移」混杂 … 两种归因处方相同, 不做对照") - the source chose not to spend a control run separating them.
Applies when
- adding frame stacking or recurrence to an actor observation
- a nominal-plant policy fails a cross-simulator gate within seconds
- planning DR staging for a new contract
“frame_hist × 零 DR = plant 指纹过拟合——5 帧窗按设计就是隐式估计器, plant 恒定时它学到 Isaac 精确响应的指纹, MuJoCo 的微小差异即 OOD, 2 s 内摔, 七个 checkpoint 无一迁移。「先标称后鲁棒」的分段与历史观测结构性冲突:估计器要见过变化才学估计,否则学背诵。”
train/OMNI_V0_SPEC.md § 3. S1.1 修订记录 ① When training fails repeatedly, inject the target behavior open-loop - stop tuning rewards for an unverified behavior
open-loop-probe-before-reward-tuningAfter two failed training attempts at a skill, stop training: demonstrate the behavior open-loop on the real plant/sim first, and only resume training once you hold a measured, safe, sign-verified target trajectory.
Symptom
Three sidewalk training rounds failed; hypotheses multiplied (exploration failure / wrong reference waveform / insufficient authority) with no way to pick between them by running more training.
Context
Instead of a fourth reward guess, the team wrote probe_side_ref.py: the candidate reference is injected open-loop on top of a frozen policy's output (bypassing PPO entirely), directly measuring "what happens if the robot literally does this waveform" - separating all three hypotheses in one experiment (5 seeds x 8 s per condition, several waveform families and gains). The probe immediately eliminated the authority hypothesis (full-amplitude execution, 5/5 survival) and localized the problem to the waveform/measurement side. The closing principle was written down after the saga: without a verified target behavior, tuning rewards is "在黑暗里试钥匙" (trying keys in the dark).
Change
Standing method: before opening another training rung on a failing skill, build an open-loop (or task-space) generator of the intended behavior, measure whether the physical system can express it and what it looks like - then train toward a verified, quantified target.
Outcome
The probe chain produced the verified waveform (reversed-sign triangle, half gain), quantified safe amplitude (tilt 8.2 deg at band top, foot distance clear of the wall), exposed the metric bug when probe and training disagreed, and ultimately supplied the feed-forward that made C4 pass in +100 iters.
Mechanism
Training couples exploration, reward design, and feasibility into one opaque outcome; open-loop injection cuts the loop and tests feasibility and waveform alone. A behavior demonstrated open-loop converts the remaining failure into a pure credit-assignment/reward question - and its measured trajectory becomes the reference itself.
Applies when
- repeated training failures on one skill with multiple live hypotheses
- uncertainty whether the platform can physically express the behavior
- a reference trajectory's shape/sign/amplitude is guessed, not measured
“三轮 FAIL 之后不再猜,写 train/probe_side_ref.py 把参考开环注入到策略输出之上(绕过 PPO),直接量「照这个波形做会怎样」,一次分开三个假说:甲 探索 / 乙 波形 / 丙 权限。”
train/C_LADDER_RUN.md § 3f. C4 真因定谳(开环探针) / 3k. 建议的下一步 Diagnose a behavior failure by enumerating hypotheses and auditing each against the actual config, cheapest first
hypothesis-table-code-auditBefore changing anything, write the full hypothesis list for the symptom and audit each against the resolved config and measured magnitudes, cheapest check first; train only on the survivors.
Symptom
Real robot leaned forward "wanting to walk" but dragged its feet instead of lifting them - a symptom with many plausible causes and no obvious single fix.
Context
Seven hypotheses were listed and each checked against the actual training config files (velocity_env_cfg.py, isaac_values.py), ordered by check cost: missing foot clearance term (CONFIRMED, primary - feet_air_time existed but no swing-height term at all); energy penalties dominating (REJECTED - energy terms total -0.19 vs tracking +1.2, 16%); command range too narrow (CONFIRMED - (0.15,0.35)); nominal pose too crouched / action scale too small (HALF - knee 0.5 rad = 28.6 deg deep, scale fine); mixed PD across motor types (REJECTED - already grouped); missing base-height reward (REJECTED - present at -5.0); height-drop termination (REJECTED - none exists, which itself became finding #4 of the fix list).
Change
The audit produced a ranked fix list (add clearance penalty; widen speed range; reduce nominal crouch) with each rejected hypothesis documented so it would not be re-litigated.
Outcome
Three confirmed causes fixed over v5/v6: swing height went 22-23 mm -> 34 mm, tracking 81% -> 87%; the rejected hypotheses stayed rejected (no wasted rungs on energy weights or PD grouping).
Mechanism
Multi-cause symptoms invite guess-and-train loops; a written hypothesis table forces each candidate to be confirmed or rejected against actual values (not impressions), and cost-ordering the checks means most hypotheses die for the price of reading a config.
Applies when
- a real or sim behavior failure has multiple plausible causes
- the team is about to "try a fix" without an audit
- post-mortems keep re-proposing already-rejected causes
“真机现象:躯干前倾像要走,脚抬不起来(拖着蹭)。按成本从低到高逐条核查 … | 1 | 缺 foot clearance | ✅ 成立,首要 | 有 feet_air_time,无任何摆动足高度项 | | 2 | 能量惩罚压过跟踪 | ❌ 不成立 | 能量类合计 −0.19,跟踪 +1.2,只占 16% |”
train/WALK_DIAGNOSIS.md § walk 拖地问题 — 七条假设的代码核查结果 Record where every failed episode ends - an end-state confusion matrix showed all failures finishing seated and overturned a "cannot roll over" diagnosis that per-category success rates hide by construction
end-state-confusion-matrixFor any multi-category acceptance, report where each failed episode ends, not only which category it started in; it costs a few lines and no extra simulation, and it separates "cannot reach the goal" from "reaches the wrong basin".
Symptom
Prone scored 0% for three generations; the working diagnosis was "prone lacks the roll-over skill", and R0.3 spent a run adding prone-to-side roll-arc start states. It bought nothing: the 45-deg roll band itself only moved from 24.2% to 26.6% after 3,000 iterations.
Context
Acceptance reported success per starting category. A final-state table (lying prone / on the side / supine / seated / standing for every failed episode) was added to accept_recovery.py at R0.3.
Change
The confusion matrix became a permanent part of the acceptance output, and the prone diagnosis was rewritten from it.
Outcome
The prone, side and supine columns were all zero - every failure ended seated - and prone had righted its torso in 159/159 episodes (tilt under 30 deg in 100%). The missing ability was standing up from one specific seated configuration, not rolling over, which redirected the next rungs to foot placement and to a configuration probe.
Mechanism
Per-category success rates collapse "reached the wrong basin" and "never reached anything" into the same zero; the end state separates them.
Applies when
- a category sits at 0% and the diagnosis rests on its label
- recovery, manipulation or navigation tasks with distinct terminal states
- an intervention aimed at the presumed cause shows no effect
“**① 末态混淆矩阵 —— 固化(已在 `accept_recovery.py`)。** 它给出的 "趴/侧躺/仰躺三列全 0、所有失败都终于坐姿"是本线最改变决策的一个事实, 而**逐类成功率按构造看不见它**。 … 成本十来行、零额外仿真。 … **② prone 病因更正(旧诊断作废)。** 旧:"缺翻身"。新:**prone 159/159 全部 把躯干翻正**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §15 R0.3 判决 + 三件事的判断 After installing the measured plant, re-run all generations paired on old and new plant - identity metrics carry verdicts, physics metrics re-baseline
plant-swap-invariants-vs-shiftsTreat every plant upgrade as an era boundary: re-run the retained policy set paired (same seeds/flags) on both plants, carry forward only verdicts whose metrics proved plant-invariant, re-baseline the rest - and mine the systematic shifts as measurements of the old plant's biases.
Symptom
With the plant finally fully measured (weighed masses 9.792 kg, bench-identified armature, in-situ friction), no historical sim number was comparable to new runs - "历史 sim 数字跨纪元不可比" - and it was unknown which historical verdicts still held.
Context
The era-2c cross-test ran all seven walk generations on the complete plant under one harness (14/14 survived), then re-ran the same 14 configurations on the OLD plant retrieved from git, same flags and seeds, with a self-check (one historical record reproduced digit-for-digit). The split was clean. Policy-identity metrics moved essentially zero across the plant swap - dominant frequency (v7's period-doubling 1.20 -> 1.20), knee amplitude (9.9 -> 10.0), foot distance (+/-2 mm), saturation (100 -> 100, 0 -> 0) - so all seven cross-generation verdicts (freeze signature, saturation-line closure, knee-collapse location, slip-penalty accounting, N2 lineage, v6's balance, v11's triad) were re-confirmed on the honest plant. Plant-physics metrics shifted systematically with ordering preserved: slip down 10-25% (measured friction makes ground-twisting costlier), landing vertical velocity down 15-50% ("旧 plant 高估落地 凶度"第二次独立证实), landing force mixed (mass up 2.2% vs friction braking the swing - two effects fighting). Exactly ONE behavior-level change: v8's low-speed period-doubling vanished (1.30 -> 2.50, lift normalizing) - confirming it had been machine-dependent bifurcation-edge behavior that armature+friction push off the knife edge, while v7's period-doubling stood untouched: saturation-freeze-driven, a policy property, not a numerical accident.
Change
Era re-baselining protocol: after any plant upgrade, one paired same-seed sweep of all retained generations on old and new plant; verdicts keyed to identity metrics carry over, thresholds re-read against the new-plant table, and the differences themselves become plant-physics findings.
Outcome
Seven verdicts survived with evidence rather than assumption; two causes of the period-doubling family were separated with plant-side proof; and the sim's landing-violence overestimate was independently confirmed a second time.
Mechanism
A policy's structural properties (frequencies, amplitudes, frozen joints) are functions of its weights and survive plant changes; contact-mediated quantities are joint properties of policy and plant and shift when the plant becomes honest. Pairing seeds across plants isolates the plant's contribution exactly, so the sweep both validates history and measures what the old plant had been lying about.
Applies when
- installing measured masses/armature/friction into the sim
- historical thresholds are cited across a plant change
- a hardware-only behavior might be bifurcation-edge sensitivity
“策略身份指标逐位不动:主频(v7 1.20→1.20)、膝摆 … 这些是策略属性,plant 换代携带无损,历史定论因此全部成立。… 唯一行为级变化:v8@0.15 的倍周期消失(主频 1.30→2.50…)——印证当时"分岔边缘、机器相关"的判定:armature+摩擦把 v8 推离刀锋;v7 的倍周期纹丝不动(1.20→1.20),它是饱和冻结驱动的深层属性,不是数值巧合。”
train/README.md § 纪元 2c 全代同机横测 (2026-08-04): 完全体 plant 上历史结论全部存活 Training the final recipe from scratch in one run - every mechanism the lineage had accumulated - produced 0% and a seated robot; the order in which the lineage acquired those mechanisms was part of why it worked
curriculum-history-is-part-of-the-productA recipe that ends a lineage is not a recipe for a from-scratch run: consolidate it as ordered curriculum phases matching how the lineage acquired its mechanisms, check that every curriculum criterion is reachable from the starting policy, and read the run's raw term values, not the total reward, before calling it green.
Symptom
V3.0 trained the lineage's whole final recipe from scratch in one 9,000 iteration run - full beta curriculum, prone-conditioned pull assist, friction DR, the 3 s zero gate on standing income, flat_feet - testing the proposition "the product is defined by its configuration, not by its training history". Training looked all green (reward 26.33, episode length 500, 100% time-outs).
Context
Read in raw units against v2_6c at the same weights, the green was a seated equilibrium: base_height 0.377 vs 0.640, stand_pose 0.205 vs 0.516, flat_feet 0.0000 (zero because it sits outside its height gate, not because the feet were flat). Acceptance: 0.0% in Isaac at the deployed authority, 0.0% on every MuJoCo friction level, and still 0.0% at the training-time authority (100% seated at 0.222 m, upright and still).
Change
The full stdout (316k lines) was read: the beta curriculum's criterion (standing share over 0.35) was met zero times, so beta never left the wide setting and the policy had no experience at the deployed authority; the pull curriculum was stuck on the same criterion. The lineage had escaped the seated basin with immediate income (V2.0-V2.2) and only then added the zero gate to cure rushing (V2.5); from scratch, the zero gate removed the early "stand fast, earn more" gradient needed to escape. Verdict "the curriculum history is part of the product", limited to n = 1. v2_6c stayed the product.
Outcome
V3.1 kept the order as explicit phases: P1 from scratch with immediate income (zero gate off) until the curricula advance, P2 adding the zero gate. P1b/P1c escaped the seated basin and passed; P2 was later judged net negative and dropped (time-gate-vs-wide-stance-retire-the-fix).
Mechanism
Mechanisms that refine a competent policy (time gates, tight authority) can delete the gradient a naive policy needs, and a curriculum whose advancement criterion the naive policy never meets freezes at its first level.
Conflicts
The spec limits the falsification to "this recipe + this curriculum criterion" (n = 1, no seed sweep, no criterion tuning). V3.1's phased run succeeding is consistent with the ordering reading but changed other terms too.
Applies when
- consolidating a long lineage of continuation fixes into one clean recipe
- a from-scratch run with all mechanisms enabled plateaus early
- curriculum state is not logged or never advances
“命题:产物由配置定义,而非训练史定义。 … 训练 log(完整 stdout 316k 行)`[beta_anchor]` 仅初始 1 行,**达标 0 次** … V2.0~V2.2 靠**即时计酬**爬出坐姿盆地(§33),站立巩固后 V2.5 才装归零门 治"过快"(§39)。**课程史是产品的一部分。**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §42 V3.0 判决(2026-08-11):从零单 run 全机制 FAIL 于坐姿盆地 A pull-assist curriculum keyed to a global success share was satisfied by the categories that already worked and withdrew before prone learned anything - conditioning the criterion on prone took it from 2.5% to 98.7%
curriculum-criterion-conditioned-on-lagging-categoryMeasure a curriculum's advancement criterion on the population the scaffold is meant to help; a global success share is met by whatever already works, and the help is withdrawn before the lagging case learns.
Symptom
Under the beta-anchored action, supine and side stood reliably while prone still sat (1.3%). A pull-assist curriculum added to help it was withdrawn completely within ~790 iterations and prone moved only to 2.5% (noise).
Context
The pull assist follows HoST: an upward force on the base, active only when the torso is within 30 deg of vertical, scaled by body weight (HoST's 200 N on G1 = 0.583 BW -> 56 N here, steps of 5.6 N, ten levels to zero); the product must pass with no assist. It had already taught sit-to-stand in V1. In V2.1 its advancement criterion was the standing-time share over all envs (threshold raised to 0.55 because the share was already ~0.53).
Change
V2.2: PullAssistForce with gate_category = "prone" - only envs whose first step classifies them as prone count toward the criterion - and the threshold back at 0.35. A feasibility signal was pre-registered: if prone's share stayed near zero under the full 56 N, return to the roll-over path instead of adding force.
Outcome
The prone-conditioned curriculum withdrew level by level only as prone itself passed: prone 98.7%, and the four-category gate passed for the first time on the line (98.6% overall, re-falls 0%, torque gate PASS).
Mechanism
A pooled success share is filled by the categories that already succeed (supine/side ~53%), so the scaffold is removed on their account before the lagging category has used it.
Applies when
- an assist, guide force or easier setting is withdrawn by a success threshold
- one task category lags while the pooled metric looks healthy
- a curriculum ran to completion without changing the lagging category
“**教训:全局站立占比阈会被存量类别(supine/side ~53%)凑够,拉力在 prone 学会前就撤光了 —— metric 设计失误,不是拉力机制失效**(它在 V1 教会过 坐→站)。 … **V2.2(已启动)**:`PullAssistForce` 加 `gate_category="prone"` —— 达标判据 只统计 prone 类 env”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §34 V2.1 判决(2026-08-10) A policy's gain profile is part of its contract - the one-leg policy needs per-joint gains the default profile lacks, and the manifest refused an evaluation under the default once; the recovery contract's beta was never stamped, a known gap not to repeat
gain-profile-belongs-in-the-stampStamp everything that defines the closed loop a policy was trained in - gains included - into its manifest, and make every consumer refuse a mismatch; a profile field that is not in the stamp is a silent misconfiguration waiting for an operator to forget a flag.
Symptom
A policy trained with hip_roll kp 80 and ankle_roll kp 60 behaves differently, or falls, under the default kp 20/12 profile - and the gain profile is a command-line flag an operator can forget.
Context
The one-leg line added a gain_profile field to the contract so the stamped manifest carries it; the spec's deployment note says the manifest guard blocks rl_default and that it had already bitten once in simulation (an evaluation run without the one-leg profile). The same spec states the general rule - any new profile field must be synced into the manifest builder - and names the counter-example: the recovery line's beta was never put into the manifest. The recovery line itself had decided that its anchored authority is computed from the base rl gains and written into the contract so it cannot drift with the gain flag, and that the older kp x 0.9 profile chosen in the V0 era does not match the beta contract and must not be used.
Change
Gain profile as a contract field checked at load; per-contract gain choices written into the run sheets.
Outcome
Evaluations and hardware runs of the one-leg policy run under rl_oneleg or are refused; the recovery beta gap stayed recorded as known.
Mechanism
A policy is trained against a closed loop whose gains are part of the plant; running it under other gains is an out-of-distribution plant, exactly like a wrong observation scale.
Applies when
- a skill introduces per-joint or skill-specific gains
- deployment gains are chosen by a command-line flag
- adding any new field to a policy profile
“增益档 `--profile rl_oneleg` 必须给 —— manifest 防线会拦 `rl_default`(sim 已咬合一次) … (recovery 的 β 未进 manifest 是已知缺口,不再复制)”
git:Lucen V2@origin/oneleg-line:train/ONELEG_V0_SPEC.md § §9b AGX 真机手顺 要点 / §3 契约 Raising a command bucket's share does not strengthen its per-state gradient - it only starves the other modes
bucket-share-is-not-a-gradient-leverWhen a skill is not learning, first prove its per-state signal is nonzero (ignore-floor and probe checks); only rebalance sampling shares to fix genuine sample starvation, and account the regression risk to the diluted modes before doing it.
Symptom
Sidewalk was not learning, and the reflex proposal was to give the side bucket a larger share of sampled commands.
Context
The C4-redo3 rung explicitly kept the 20/40/20/20 bucket (stand/forward/turn/side) with the reasoning written out: PPO computes advantages per state, so bucket proportion does not change the per-state gradient of side states; at 4096 envs x 20% x 24 steps the rollout already contained ~19.7k sidewalk states - sample count was not the bottleneck. And the cost side was already measured: cutting forward from 60% to 40% had made vx+0.30 die at +400 in an earlier run - more cuts would only collapse it sooner.
Change
Bucket proportions held constant across the entire C4 redo series; the actual bottlenecks (metric frame bug, reward variance penalty, exploration form) were pursued instead.
Outcome
Sidewalk was eventually fixed with zero bucket changes (feed-forward delivery, +100 iters); forward/turn skills never suffered starvation-induced regressions during the redo series.
Mechanism
Policy-gradient credit is assigned per visited state; oversampling a mode multiplies its states in the batch but not the informativeness of each, so if the per-state gradient is ~0 (behavior unreachable or reward indifferent), N times zero is still zero - while the displaced modes genuinely lose data and regress.
Applies when
- proposing to oversample a failing task/command mode
- a majority mode regresses after share rebalancing
- budgeting env count vs mode share for a multi-skill policy
“比例不动:PPO 逐状态算优势,桶占比不改变单状态梯度;4096 env × 20% × 24 = 每 rollout 已有 1.97 万个侧走状态,样本数不是瓶颈;而 forward 60%→40% 已实测让 f30 在 +400 处死掉,再加码只会更早塌。”
train/C_LADDER_RUN.md § 3i. 桶 20/40/20/20 不动(比例不动) The first real-robot get-up was "very violent, kicking on the floor, dangerous" - a sim-perfect policy with no reason to be slow, unbounded absolute targets, no domain randomization and a rate limiter that filtered nothing; the task was restated as "safe, slow, transferable"
first-real-get-up-violent-stage-one-policyDo not put a get-up policy on hardware until its action is bounded (hard bound or state-anchored targets), smoothed, randomized and tested at the real pipeline's latency, and say explicitly that the task is "safe, slow and transferable" - a simulation-perfect policy optimizes only "gets up".
Symptom
On 2026-08-09 the user ran a V0-lineage recovery policy on the real robot and stopped it: very violent, kicking on the floor, dangerous. The planned next rung (a heavier torque_headroom) was never started.
Context
The spec had pre-registered that R0/R1 products stay in simulation and that the real-robot precondition was the R3 smoothing rungs plus a bridge-slew check plus a hanging protocol; the robustness (DR) rungs had not run. In simulation the policy passed 100% with a get-up of about a second. Which ONNX, which gain profile and whether a torque/joint log existed were left "to be recorded later" and never were.
Change
The V0 ladder was stopped at its best product (R3.1, sim only) and a re-rooting proposal was put to the user. The spec's four-layer account: style (the reward pays for standing early and nothing pays for slowness - HumanUP's "Stage I" get-up, "fast but unsafe ... infeasible for real-world deployment"); impact (full-range absolute targets with no hard bound, raw |a| up to 4.77, action saturation 100%, a single-step change of 0.306 saturating hip_pitch); transfer (zero DR, friction pinned at 1.0, the learned leg bracing); link (the bridge's RL slew equals vel_limit, 0.2-0.66 rad per step, while the real pipeline has 1-2 steps of time-varying latency and acceptance ran at delay 0).
Outcome
The line was re-rooted twice (training-side rate limit, then the beta-anchored action space) and gained a hang protocol before the next real attempt; on 08-11 a beta-anchored policy produced the line's first real get-up.
Mechanism
A task reward that pays for standing early selects the fastest feasible get-up; with absolute full-range targets every large target jump is a torque impulse bounded only by the clip; zero DR and braced-leg solutions do not transfer; and a limiter set at the velocity limit does nothing at 50 Hz.
Conflicts
The four layers are the spec's reconstruction from simulation probes and the literature; the real run's policy file, gain profile and log were never recorded, so no layer was confirmed against hardware data.
Applies when
- a first hardware trial of a high-effort skill is being scheduled
- sim success is high but the policy saturates actions or torques
- pre-registered hardware preconditions are not all met
“用户真机反馈:**非常猛、地上乱踢、危险**,叫停(R3.3 torque_headroom 加档已选型 weight −0.5→−1.5,未启动)。真机细节(哪个 onnx、什么档、有无 τ/q log)**待补记** … 任务从"能起来"变成 **"安全、慢、可迁移"** … **链路层**:桥层 slew RL 档 = vel_limit(10/20/33 rad/s ≈ 每拍 0.2~0.66 rad), 对 recovery 形同虚设;真机 1~2 拍时变延迟,验收默认 delay 0。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 真机叫停与换根判决(2026-08-09) The deploy-side walk/recovery switch - into recovery at tilt > 65 deg held 0.3 s, back at tilt < 15 deg with angular rate < 1 rad/s and straight knees held 1 s, a 15 s timeout, last action cleared both ways - and the handoff steps the design required are only partly implemented
walk-recovery-fsm-handoffSpecify a deploy-time controller switch as hysteretic, time-filtered predicates the robot can measure (proxy what it cannot, e.g. straight knees for height), a timeout that ends in a safe stop, and a complete handoff (history, clock, last action, command ramp) - then test that the code performs every handoff step, because the design document is not the implementation.
Symptom
With a recovery policy and a locomotion policy as separate networks, the robot needs a switch: when is it "fallen", when is it "up", and what state must be reset so the next policy does not act on the previous one's history.
Context
The 08-09 design: enter recovery when fallen (tilt > 55 deg or height < 0.60 x 0.384 m) for 150 ms, leave for a stand-hold when upright (tilt < 12 deg, height > 0.85 x 0.384 m, feet steady, |omega| < 0.8) for 400 ms - wide entry, strict exit, hysteresis - then a mandatory handoff trio before walking resumes (reset the walking policy's observation history, restart its phase clock at 0, clear its previous action and latency buffer) and a command ramp instead of a jump. The 08-14 implementation in deploy_policy (--recovery-policy): each policy under its own manifest contract (walk: nominal + scale; recovery: beta-anchored), the same rl_default gains, the tilt cutoff disabled; RECOVERY at tilt > 65 deg for 0.3 s; LOCO again at tilt < 15 deg and |omega| < 1 and knees straight (< 0.35 rad) for 1 s - the robot's computer has no height estimate, and straight knees stand in for height so a V3.0-style upright kneel cannot pass as standing; RECOVERY longer than 15 s ends in a safe stop; last_action cleared on both switches; command forced to 0 during RECOVERY; power scaling applies to LOCO only.
Change
An open account was written down with it: the recovery end state (0.355 m stance, hip yaw -/+27 deg) is outside the walking policy's training start distribution, so the first test must use stand / zero command as LOCO, and the long-term fix is to widen the walking policy's initial states rather than bend recovery's stance to suit walking.
Outcome
The spec records only a successful compile (py_compile); the hang test was left to be done on site. The operator runbook carries the three-step procedure (hang with stand as LOCO, mat and push, then omni walk as LOCO) but no outcome.
Mechanism
Two policies trained separately each assume their own history, clock and last action; a switch that carries any of them across feeds the next policy a state it never saw - the design called this "the walking policy seeing a ghost history".
Conflicts
The 08-09 design requires resetting history, clock and previous action plus a command ramp; the 08-14 implementation records last_action clearing and a zero command during recovery; the one-leg spec of 2026-09-14 lists the handoff hygiene as specified but not implemented - reset_history() is called by nothing (a 215-dim policy would carry four frames of pre-fall history), the phase clock is not zeroed, and there is no command ramp back to LOCO. No hardware run of the FSM is recorded in either source.
Applies when
- switching between separately trained policies on hardware
- a policy with history or phase observations is re-enabled mid-run
- the robot lacks a sensor the switching criterion was designed around
“**判据**:进 RECOVERY = 倾角 >65°(`--fall-tilt-deg`)持续 0.3 s;回 LOCO = §5 真机可测子集:倾角 <15° ∧ |ω|<1 ∧ **膝直 <0.35 rad(NX 无高度观测, 高度门用膝直代理 —— 防 V3.0 型"跪坐但直立"误判)** 持续 1 s (`--recover-hold`);RECOVERY 单次 >15 s(`--recovery-max-time`)安全停。 … **切换卫生**:两向切换 last_action 清零;RECOVERY 态 cmd 强制 0”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §50 FSM 双策略调度(2026-08-14,用户令):deploy_policy --recovery-policy Fix the task first, harden the plant second - DR budget spent on a dying task is wasted
task-shaping-before-plant-hardeningFreeze the task/command distribution before spending DR budget on plant robustness; if the task will still change, schedule plant hardening as a final pass and book the interim robustness gap explicitly.
Symptom
Tempting default ordering was to keep the plant-hardened (S2) lineage and teach it new commands; but the S2 plant adaptation had been earned on the straight-walk task, and the new omni tasks (sidewalk, in-place turn) use completely different contact patterns.
Context
The team had direct evidence that DR robustness is a budget that gets reallocated when the data distribution changes ("push/μ 两轮已实证 DR 预算有限且会被重分配") - robustness trained under one task/command distribution does not persist when training continues under another.
Change
Ladder order set to: first C (task shaping - add command modes until the task family is final), then a second S2 pass (plant hardening) on the C product. The plant-robustness gap this creates mid-ladder is accepted and booked explicitly ("此处不欠账" - the debt is assigned to the second S2 pass, not denied).
Outcome
The first S2 pass was not wasted: its laws (kd bandwidth <-> low mu, push need not be trained, ground mu need not be trained, bistability) let the second pass drop from five rungs to three. The C ladder itself ran on the softer plant band without incident.
Mechanism
DR robustness is carried by the policy's visited-state distribution; changing the task changes that distribution, so robustness bought under the old task partially dissolves. Hardening before the task is final means paying for robustness on states that will no longer be visited - "给一个即将不存在的任务花预算" (spending budget on a soon-to-not-exist task).
Applies when
- deciding ordering between skill/command expansion and DR hardening
- a hardened lineage is proposed as the root for a task change
- robustness regressions appear after adding new command modes
“S2 的 plant 适应是为直行步态调的,C4 侧走/C3 原地转是完全不同的接触模式,先硬化再改任务 = 给一个即将不存在的任务花预算(push/μ 两轮已实证 DR 预算有限且会被重分配)。故顺序改为 先 C(任务定型)→ 再 S2(plant 硬化)。”
train/C_LADDER_RUN.md § 0. 决策逻辑 = 短板可不可恢复 (末段) action_rate weight is the sim2real bandwidth knob - re-tune it whenever a rate limiter is removed
action-rate-weight-vs-bandwidthSet action_rate weight relative to real actuator bandwidth, and re-tune it any time another smoothing/limiting element (filter, slew limiter, gain) changes - reward weights are load-bearing parts of the actuator model.
Symptom
With a low action_rate_l2 weight the policy learns fast actions; the unmodeled part of the actuator response is then excited hardest, and sim2real "直接崩" (collapses outright). With too high a weight, actions become so slow the robot cannot maintain balance.
Context
The reference developer called action_rate_l2 the single most important reward for transfer, with side-by-side video evidence that the high-penalty, slower policy is clearly better on hardware. Lucen context: the team had just removed the SOFT_SPD=1.0 velocity limiter, which had been an implicit actuator-bandwidth constraint - leaving action_rate as the only remaining constraint on action speed.
Change
Decision recorded: after removing SOFT_SPD, re-evaluate the action_rate weight rather than keep the old value, since its effective role changed from "additional smoother" to "sole bandwidth constraint".
Outcome
Logged as a priority follow-up ("重新评估 action_rate 权重 - 拆掉 SOFT_SPD 之后这一项的作用变了"); the failure mode it guards against is training high-frequency actions the real actuators cannot track.
Mechanism
Slower actions stay inside the frequency band where the ideal-PD sim actuator and the real actuator agree; fast actions probe the band where unmodeled delay, inductance, and bandwidth limits dominate, so model error is amplified in exact proportion to action speed. Any removed external rate limit transfers that constraint's entire job onto the action_rate penalty.
Conflicts
The low/high tradeoff evidence is the external developer's report (with video); the Lucen-side entry is a pre-registered risk and decision, not yet an on-robot A/B at the time of writing.
Applies when
- removing or adding an action filter, slew limiter, or low-level speed cap
- real robot shows high-frequency chatter or overheating absent in sim
- tuning smoothness rewards before a hardware deployment
“权重低 → 动作快 → 执行器模型不准的部分被放大,sim2real 直接崩 / 权重高 → 动作慢 → 好迁移,但可能慢到无法维持平衡 … 我们刚拆掉 SOFT_SPD=1.0 的限速器,等于把执行器带宽约束整个移除了。action_rate 惩罚现在是唯一还在约束动作速率的东西,需要重新评估权重”
Experience.md § action_rate_l2 是他认为最关键的 reward (lines 61-70) Tightening the bridge's rate limiter under an unchanged policy cut torque peaks 30-50% and made other things worse - the policy cannot see the limiter, keeps commanding and winds up; a deploy-side limiter is a safety net, not a cure
deploy-rate-limiter-windupA rate or torque limiter added at deployment lowers peaks but the policy still commands as if unconstrained (saturation, windup, new contacts); use it as a safety net mirrored in evaluation, and put the constraint where the policy can learn around it.
Symptom
After the violent first real-robot get-up, the cheapest candidate fix was to tighten the bridge's slew (rate) limit for the recovery policy without retraining.
Context
Probe on R3.1 in MuJoCo (5 categories x 3 seeds, mu 1.0), monkeypatching the limiter with no repository change: TIGHT = RS06 4.0 / RS02 3.0 / RS00 2.0 rad/s (about 0.08/0.06/0.04 rad per policy step) against the current vel_limit setting.
Change
The probe decided the role of the limiter rather than a deployment.
Outcome
Success 14/15 -> 12/15; get-up median 2.35 -> 3.53 s (max 9.30); torque demand peak median hip_pitch 164% -> 111%, knee 166% -> 86%; action saturation still 100%; leg-leg contact 558 -> 860 frames. The limiter was kept only as a real-robot safety net (mirrored into sim2sim evaluation); the cure moved into training - where the next lesson was that a limiter anchored on the last command is itself an integrator (slew-anchor-is-an-integrator).
Mechanism
A policy that never trained with the limiter keeps issuing the targets it learned; the limiter clips them, the target window runs ahead (windup), and the robot follows a trajectory the policy never evaluated.
Applies when
- a trained policy is too violent on hardware and a quick deploy-side fix is tempting
- adding slew, torque or velocity limits in a bridge or firmware
- evaluation and deployment use different limiter settings
“判读:**链路侧收紧立等可取地把 τ 峰值砍 30~50%,但成功率掉、饱和率仍 100%、 腿-腿接触反升** —— 策略感知不到限速器,目标窗口继续狂奔。⇒ 收紧 slew 只配当 **真机侧安全网**(必须同步进 sim2sim 口径,基础设施现成),**不配当治法; 治法必须进训练**。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §26 探针:收紧桥层 slew,r3_1 不重训直接测 The best checkpoint to SHIP is not the best checkpoint to CONTINUE FROM - maturity is capital against adaptation shock
root-maturity-vs-product-qualityDecide shipping points and fork roots separately: gates rank products, but a root candidate must prove itself by surviving a continuation under the next rung's shift (dual-arm if in doubt) - and prefer the more-trained point as root when product metrics conflict with maturity.
Symptom
A band re-audit found s1e-300 beat the incumbent root s1e-500 on nearly every quality gate (stepping 19/20 vs 13/20 with historically-best 26.9 mm swing, speed gate 14/20 vs 2/20, heading 26 vs 54 deg/20 s) - suggesting the root had been mis-picked and the younger point should take over.
Context
The dual-arm control settled it the other way: continuing the S2 PD rung from s1e-500 adapted smoothly (3/3 smoke throughout), while the b300 control arm (same config, from s1e-300) fell into a survival valley under the PD shock (+100 iters: 1/3 -> 0/3), never climbed out within budget, and its 800-iter product scored 13/20 survival - eliminated. Verdict: "幼年点自身指标再好也扛不住新 DR 适应冲击, 成熟度是本钱,s1e-500 根被数据背书" - a young point's own metrics, however good, do not survive new-DR adaptation shock; maturity is capital. The audit still yielded value: the band scan (200-1000, per-100) mapped the lineage's arc (200 dragging -> 300 peak -> 400+ decay -> 900+ drift blowout), and 300 remains the better PRODUCT answer for shipping-as-is questions.
Change
Selection doctrine split into two questions with different answers: best-product point (quality gates at the point itself) vs best-root point (survives adaptation shocks; more training age = more capital), each decided by its own evidence - and root claims settled by a dual-arm continuation test, not by point metrics.
Outcome
s1e-500 kept the root role with data behind it; the S2e ladder built on it passed rung after rung, while the b300 line was closed at the cost of one control arm.
Mechanism
Early checkpoints sit near sharp optima with less accumulated robustness structure; their headline metrics reflect the narrow training distribution, not resilience to distribution shifts. A continuation rung is itself a distribution shift, so the root property being selected for is shock tolerance - observable only by actually continuing, never by static gates.
Applies when
- a younger checkpoint outscores the current root on quality gates
- choosing the base for a robustification or command ladder
- a continuation run stalls in an early survival valley
“b300 对照臂 … PD 冲击下存活谷(+100 起 1/3→0/3),预算尽未爬出,800 档 20-seed 存活 13/20 出局——幼年点自身指标再好也扛不住新 DR 适应冲击,成熟度是本钱,s1e-500 根被数据背书”
train/README.md § omni_s2e_pd (b300 对照臂) / s1e 选点重审 A nonzero response with the same sign for + and - commands is bias, not ability
same-sign-response-is-yaw-biasBefore crediting any directional skill, test both command signs: response must flip sign with the command; a same-signed pair is a bias to subtract, not an ability to report.
Symptom
Root-selection probe showed nonzero wz "tracking percentages" on turn commands, tempting the read that candidates could partially turn.
Context
During C-ladder root selection, s1e-500's measured yaw rate was +0.084 rad/s for cmd +0.3 and +0.093 rad/s for cmd -0.3 - same sign both ways. The same check on the C2 baseline gave wz+0.20 -> -0.13 and wz-0.20 -> +0.12 (again same sign), while the alternative root s2e_pd-1400 gave +0.16 / -0.16 - opposite signs, i.e. a genuine 16% command response.
Change
Reading corrected and written into the execution sheet: percentages on directional commands are meaningless unless the +cmd and -cmd responses have opposite signs; all three candidates were re-classified as "cannot turn, cannot sidewalk - C2/C3/C4 learn from zero". Acceptance criteria thereafter required "tracking >=50% AND left/right opposite-signed".
Outcome
Prevented crediting turn/sidewalk ability that did not exist; the antisymmetry clause became a standing part of every turn and sidewalk PASS condition (C2, C4, C4-redo levels all carry "且左右反号").
Mechanism
A constant yaw (or lateral) bias projects onto any command's sign convention and shows up as fake fractional tracking; only sign-antisymmetry under command reversal distinguishes a feedback response to the command from an open-loop offset.
Applies when
- evaluating turn/sidewalk/any signed-command tracking percentages
- a candidate shows partial tracking on an axis it was never trained on
- writing PASS criteria for a new directional skill
“C2/C3 那些非零的 wz 百分比不是转向能力 —— 转向+ 与 转向− 的实测同号(s1e:cmd +0.3 → +0.084,cmd −0.3 → +0.093 rad/s),那是恒定偏航偏置。… 三个候选都不会转、都不会侧走。”
train/C_LADDER_RUN.md § 0. 读数纠正(重要,别引错) A plateau in a training curve was a population mix, not a half-learned skill - 60% standing at 0.372 m and 40% sitting at 0.19 m - and a category at hard zero stayed at zero through 3,000 more iterations
zero-partial-credit-is-not-an-iteration-problemBefore buying iterations for a plateau, split the metric by category and check whether it is bimodal; a category at hard zero with no partial credit is missing a capability or a reachable state, and more iterations under an unchanged config will only polish the categories that already work.
Symptom
After R0.1 the training-side base_height sat near 0.29 m and the curve was still climbing at the iteration cap, which read as "train it longer".
Context
Candidate A (user decision) was a child-run from R0.1's last checkpoint with zero config change - the logged env.yaml files differ only in log_dir - for 3,000 more iterations (R0.2). The per-category acceptance split was already available: prone had scored 0/156 with no partial credit.
Change
Continue training unchanged, then read the result by category rather than by the pooled curve.
Outcome
supine 91.5 -> 96.4%, side 82.2 -> 87.9%, get-up 0.96 -> 0.84 s, pose error 0.91 -> 0.54 - all improvements to categories that already stood. Prone stayed 0/153; mid 55.9 -> 41.2% was within noise (n=34). Height by category was binary - standing groups 0.372/0.373 m, seated groups 0.187/0.194 m, nothing between - so the pooled 0.29 was 0.61 x 0.372 + 0.39 x 0.19 = 0.30 (measured 0.307): six in ten standing, four in ten sitting. The rendered prone episode was still kneel-sitting at t = 8 s.
Mechanism
The pooled mean of a binary outcome only moves when the mix moves; PPO kept polishing the subpopulation that already succeeded while the failing one produced no advantage signal to follow.
Conflicts
R0.2 recorded the missing capability as "prone lacks rolling over"; R0.3's end-state confusion matrix retracted that - prone had righted its torso in 159/159 episodes and was failing to stand from the W-sit. The lesson that iterations could not fix it holds; the named cause was wrong.
Applies when
- a training curve plateaus while acceptance shows one category at zero
- deciding between "train longer" and "change something"
- pooled training metrics are read as the typical episode
“**分类别 h 中位把"平台 = 人口混合"钉死了**:数值是**二值**的 —— 站立组 0.372/0.373,坐姿组 0.187/0.194,**中间没有过渡态**。 … **这也是本仓此后读该指标的通用告诫:全体混合的期望会把 双峰分布平均成一个不存在的中间值,必须分类别看。** … **结论:A 不能过门,原因确定为 prone 缺"翻身"这一技能,不是迭代不够。**”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §13 R0.2(recovery_r0_2,child-run 续训):A 走完了 —— 推不动 prone