Coach playbook
Lessons 171
What happened in a real run, what changed and why it worked. Open one to read the quote it rests on.
Reward design 36
How reward terms and their weights shape what a policy learns.
- action_rate weight is the sim2real bandwidth knob - re-tune it whenever a rate limiter is removed
- A binary reward band on the swing knee had zero gradient everywhere below it, so the one-leg policy parked in an unloaded "fake touchdown" that Isaac's 5 N threshold scored as success and MuJoCo showed as real pressing - a capped constant-gradient ramp, retrained from scratch, passed 40/40
- Calibrate a wall penalty by measuring healthy and sick policies - healthy pays ~0, the disease pays a wall
- Every new penalty ships with a pre-registered withdrawal clause - if healthy gait must pay above the cap, the term stands down
- An exponential kernel on instantaneous velocity punishes gait oscillation - track the cycle average
- A joint-velocity penalty meant to slow the get-up cut joint speed 16% and left the get-up time unchanged - the knob never moved the variable, so the idea it was meant to test stayed untested
- Before training a one-leg stand the spec named the cheapest cheats - hopping on the support foot, a raised foot resting unloaded, a leg tripod - and gave each a countermeasure and a gate; one still appeared and was caught by exactly those gates
- Exponential tracking kernels go flat exactly when the error is largest - pair them with an L2 term for the far field
- Gate a new reward term by its command so all old modes score pointwise identical
- Audit which joints your imitation term constrains - a task that needs deviation is fighting the reference
- A quadratic clearance reward stalled for 3500 iters near target - switch to an indicator on accumulated height
- Prove a new penalty actually fires - two ways a clearance term silently did nothing
- Three times a joint-angle stand-in for a foot-level quantity was gamed or lied - the absolute ankle roll sold stance width to buy flat feet, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart
- Knee swing collapsed because it directly trades against the slip penalty - price the conflict explicitly and clamp what reward cannot hold
- Torque caps cannot soften footfalls - impact is falling-mass momentum, only the reward can treat it
- The restart reward table lists a reason for every term AND a lesson for every exclusion - absent terms are removed, not zero-weighted
- Ungated phase shaping made standing 42x more expensive than stepping - and the stepping was cooking the hip motors
- Fix a too-deep nominal pose before adding any penalties - the default stance defines the basin training starts in
- An edge-triggered landing penalty missed the tail and fired after the harm - penalize overspeed continuously inside the contact window
- Decompose the offending quantity by channel first - then penalize the failure event, not the joints
- Every gate on a penalty is an exit - to stop paying a stance tax the policy parked 2 deg outside a 30 deg uprightness gate (lunging), and, from scratch, just under a height gate (crouching); a positive band and an always-on guard fixed both
- The get-up kept getting faster because standing earlier paid more every step - lowering torque authority barely slowed it, and only zeroing the standing income for the first 3 s moved the pace into the design band
- The walk phase machine structurally cannot express flight - rebuild the representation for duty < 0.5, teach flight with a mask tax, never a cliff bounty
- Prone get-up sat at 0/159 until two gated hinge terms moved the seated feet - first sideways (561 -> 360 mm), then fore-aft (-168 -> +56 mm) - and success went to 158/159 with nothing else changed
- Audit rewards by realized contribution (weight x achieved value) - a weight of 2.0 was really paying 0.04
- FK-verify a borrowed reference's structure, then size its amplitude by the reference's job - it pins phase, the policy adds lift
- Reward fixes come in causal chains - foot height, then landing impact, then foot spacing
- Before adding a command mode, compute what ignoring it costs - the lazy optimum must lose
- A joint frozen at the action clamp pays zero action_rate forever - penalize pre-clip saturation to make the cheat cost money
- A get-up policy righted itself and sat - three terms paid the seated pose 84% of the return, and the only shaping term that could tell sitting from standing was an exp kernel outputting 5e-5
- A soft joint-limit penalty charged the standing pose itself - the geometric-zero knee sat on its hard limit, so stand_v1 bent its knees to dodge 0.419 per step and leaned 4.1 deg forward; excluding the knee gave 0.24 deg
- A stronger action_rate penalty cut the median torque demand under the gate and left the p99 at 4x the limit - only a hinge on the pre-clip (computed) torque, weighted by comparison with a peer term, collapsed the tail
- Add a termination that makes the degenerate strategy fatal - no height cut-off meant crouch-shuffling could live forever
- A time gate that had cured one lineage's rushing made the from-scratch lineage trade away its stance width twice (0.364 -> 0.235 m, 0.355 -> 0.251 m) - its disease was absent there, so the fix was retired and the pre-gate checkpoint shipped
- Borrow reward values from other robots as ratios (to tracking weight, to leg length) - never as absolute numbers
- With no term on foot attitude the robot stood on the edges of its feet (ankle roll -26 deg) from R0.5 to V2.5b; pricing it fixed that and turned stance width into the next free variable - the user saw both on video before any metric flagged them
Actuator modeling 16
Measuring the motors and the body so simulation behaves like the robot.
- Ideal PD is not enough - add a delay buffer and fit armature/friction/delay per joint
- Armature must be N^2 x rotor inertia, never 0 - measure it no-load
- Model CAN polling skew - joint observations are 6-9 ms stale by read order
- The trainer read a stale USD after the URDF mass update - regenerate derived assets and gate on an automated equality instrument
- Three hardware accounts locked the run design point - and the knee's real speed ceiling is tau_limit/kd, not the firmware limit
- Guessed joint friction was 2.5x low and damping 5x high - measure, then DR around nominal
- Prove an armless get-up exists before training it - a connected static domain, 25% torque on the cheapest path, an 8 mm hand-over gap, a static roll-over - and write down what each scan cannot represent
- Removing a hand trim re-exposed the plant offset it had been silently compensating - and a slope scan told bias from sensitivity
- The latency DR range must cover the measured deployment pipeline - 0-20 ms could not even reach the real 1-2 control steps
- The action-delay was implemented as lerp - beyond one step it extrapolated BACKWARD, so a whole lineage trained on a fictitious actuator
- The real robot's right-leg kicking was over-trained-delay times loop gain - irreducible pipeline latency is plant, model it fully from day one
- Train with self-collisions ON (filtering nested-link ghost pairs) - the reward wall prevents, the physics makes cheating impossible
- Isaac splits Coulomb friction into static and dynamic columns - wiring only static means zero loss during motion, silently discarding the identified value
- Before training a one-leg stand, the accounts and a probe showed the default gains could not hold it at all - kp 20 needs 0.39 rad of error to carry the static roll moment, more than the whole adduction range - so per-joint gains came first, and thermal limits set the session length
- Set torque limits per joint from measured gait peaks - a uniform percentage is the wrong shape, and training must use the deployed numbers
- Export every CAD part in the whole-machine frame so URDF rotations are zero and inertia is exact
Sim-to-real 44
What changes on the real robot, and how randomization and inputs prepare for it.
- Bracket a real-robot A/B with a repeated reference run - battery drain is the confound
- The +/-50 mm lateral COM randomization meant to spread the legs coincided with legs pulling IN - rolled back per its own pre-registered contract
- Oversized lateral COM randomization (+/-5 cm) deliberately forces leg spread
- A constant-value plant rung passed every binary gate with record scores - and shipped 60% thinner posture margins that hardware exposed
- Slowing the gait clock at deployment is out-of-distribution and backfires - lower the commanded speed instead, or train the knob
- An outer heading P-loop at deploy cut drift 10x because its output stays inside the trained command band - then training was aligned to it
- A DR tail the robot never has is pure cost - stage deterministic plant levels instead of one wide uniform
- A walking policy's tilt cutoff is a legal state for a recovery policy - the default 45 deg fall guard had to be raised for recovery tests and is disabled once the switch owns falls, so the abort chain becomes the recovery timeout, the operator's cut, and the firmware torque limits
- The first real-robot get-up was "very violent, kicking on the floor, dangerous" - a sim-perfect policy with no reason to be slow, unbounded absolute targets, no domain randomization and a rate limiter that filtered nothing; the task was restated as "safe, slow, transferable"
- Friction DR was demoted after a measurement (94% success at mu 0.4 with no friction randomization) and promoted again when the action contract changed and mu 0.4 fell to 76% - DR priorities belong to a plant and contract, not to a task
- A policy's gain profile is part of its contract - the one-leg policy needs per-joint gains the default profile lacks, and the manifest refused an evaluation under the default once; the recovery contract's beta was never stamped, a known gap not to repeat
- A hardware run without its log is an anecdote - the first real get-up's policy, gain profile and log were never recorded, two CSVs stayed "to be reported", and runbook commands wrote different policies' logs under one copied filename
- The hip_roll (l+r) asymmetry scalar predicted real-robot lateral drift - promote validated sim scalars into the gate
- A frame-history observation under zero DR memorizes the trainer's plant fingerprint - the estimator must see variation to learn estimation
- IMU observation age cut 52-68 ms to ~4 ms by moving AHRS onto the MCU - as a single variable
- Brief the operator on the lineage's measured zero-command and untrained-axis behavior before handing over the joystick
- When a contract default changes, old policies must run under a pinned legacy profile - a silent clock swap is out-of-distribution on hardware
- Mirror augmentation over an asymmetric default injects systematic error - symmetrize the default first and verify the transform bit-exact
- No parameter tuning on the floor - a failing config retries once, then it is out; anomalies go back to sim
- Privileged signals (true velocity, foot force, foot height) go to the critic only
- Size each joint's action authority to its measured working range - lock a channel to zero only when its job is provably elsewhere
- Every power cycle starts with the same read-only pre-flight - read the buses, check the torque limits against 12/17/11, verify the IMU axes, check the ports after any new USB device - and any reassembly re-measures the joint zeros
- The walking lines' safety setting, power-scale 0.8, broke the recovery policy's full-range contract - it cut the ends of the joint travel (4/50 could not get up) and left the torque spikes untouched; a kp x 0.9 gain profile inside the trained kp band did the job
- Deployment power derating damages non-forward axes far more than forward - sweep it in sim before deploying
- Write each config's expected hardware signature before the session - and if reality disagrees, change the books, not the conclusion
- Push DR helped one lineage and hurt another at the same dose - robustness budget is conserved and gets borrowed, not created
- Push-test protocol - positive side first, fragile side spotted, axes aligned in the log, and cross-machine push counts stay qualitative
- Removing a foot-spacing wall passed every simulated gate and made the feet collide on the real robot - nothing priced stance width in the two-foot phase, the policy narrowed to the simulator's self-collision floor, and real calibration offsets closed the last millimetres; the wall came back with a gate
- A field experiment is allowed when it is pre-scripted, single-variable, and self-reversing (RAM-only writes)
- A reward on a quantity the actor cannot observe teaches "produce less of it", never "correct it" - closed-loop correction needs an outer loop
- Order hardware runs by sim risk, gate each stage on the last, and put the fragile cell last with a spotter
- A proposal in the runbook - torque and action limits as versioned safety tiers (classroom / research / expert) written to motor RAM and read back, separate from the reward's effort penalty - recorded as a proposal, its implementation unrecorded
- A sim veto needs real confirmation too - the worst sim cell was scheduled as the most informative hardware run
- Real-robot trials of a new skill were staged by risk - a hanging dry run with the robot posed by hand, then one short try per category on a mat with the hardest last, then the composed behaviour (switch + walking) last - with the user present and a log every time
- A real-robot verdict is (policy x deployment stack) - when the stack changes materially, old verdicts expire
- A stand gate judged by survival passes a robot that wanders a meter - judge posture instead
- A suspended (no-load) test acquits or convicts the actuator before you blame authority
- Three same-shaped judging errors - task metrics (survival, tracking, displacement) cannot stand in for posture metrics
- Teleop fed the sidewalk axis a command beyond its training band - feet clipped; give each axis its own speed setting
- The run policy never left the ground and fell in the second simulator from the frontal plane - its DR (gains and latency only) covered the actuator axis, not the frontal-plane contact and inertia disturbances the doubled stride amplified; "is DR on" is the wrong question
- Two machines, one configuration - every gain, offset and torque limit changes only in robot.yaml, whoever edits pushes at once, both checkouts show the same commit before the robot moves, and a pulled policy file is size-checked and synced before power-off
- The deploy-side walk/recovery switch - into recovery at tilt > 65 deg held 0.3 s, back at tilt < 15 deg with angular rate < 1 rad/s and straight knees held 1 s, a 15 s timeout, last action cleared both ways - and the handoff steps the design required are only partly implemented
- The recovery line's real-robot verdicts live in three places that disagree - the spec's "fairly stable, first real stand-up" for v2_6, an undated runbook note that only v3_1p1c works, and a first run whose details were never recorded
- A 2-degree joint-zero calibration fix moved the whole runnable envelope - re-test old "cannot run" verdicts after recalibration
Evaluation 27
Testing a policy in simulation before it goes near hardware.
- Cutting swing amplitude 40% raised yaw-momentum demand 53% - the falsified fix is recorded so nobody walks that road again
- mj_objectVelocity returns inertial-principal-axis frame - one API assumption poisoned eval and observations for a whole line
- The shipped checkpoint was chosen by scanning checkpoints on the full gate - neighbours 100 iterations apart failed 1 and 38 cells, late checkpoints degraded - never by taking the last one, and training stopped on signals, not on a schedule
- Score left and right separately - averages hide chirality breaking that mirror augmentation does not prevent
- A single-signal contact detector lied in both directions - foot height flagged 40% false flight on a walking gait, contact force alone flagged false flight during low-friction slip - so flight became force < 5 N AND sole height > 5 mm
- Tightening the bridge's rate limiter under an unchanged policy cut torque peaks 30-50% and made other things worse - the policy cannot see the limiter, keeps commanding and winds up; a deploy-side limiter is a safety net, not a cure
- Record where every failed episode ends - an end-state confusion matrix showed all failures finishing seated and overturned a "cannot roll over" diagnosis that per-category success rates hide by construction
- A 10 s acceptance episode left 6-7 s of standing to observe - a narrow stance held for that window and split on hardware; a gate cannot see instability slower than its own horizon
- Run acceptance under measured contact parameters - honest condim/torsional-friction flipped a false PASS into a real-matching FAIL
- One fixed acceptance matrix for every rung - new skill must PASS while every old skill stays within a regression budget
- When the training reset distribution changes, freeze the acceptance distribution separately and pin its seed - the same checkpoint measured twice differed by 3.8 and 10.2 points
- The 1.5 Hz step-frequency gate was retracted - the author had misread his own actuator data, and the limit fought pendulum dynamics
- Measure yaw rate by integrating heading, not by averaging body-frame angular velocity - the two differed 15x
- The median of a bimodal metric lands in the empty gap - check the distribution, and never judge swing on 3 seeds
- A single run's drift direction may be a limit cycle, not a policy bias - check the sign distribution across seeds
- After installing the measured plant, re-run all generations paired on old and new plant - identity metrics carry verdicts, physics metrics re-baseline
- A ratio metric flipped the verdict - spectral share rose while absolute high-frequency energy fell 16%
- When the eval proxy has a known systematic bias, gate on within-proxy differences, not absolutes
- Training-log reward values and fixed-command eval values live on different distributions - comparing them once claimed a 44% improvement that was really 6-10%
- A nonzero response with the same sign for + and - commands is bias, not ability
- Cross-simulator gate (Isaac Lab to MuJoCo) comes before any hardware attempt
- Single-impulse push recovery is a binary chaotic quantity - cross-machine floating-point divergence can flip the outcome
- The smoke watcher panicked at the wrong operating point and its score goes blind when gates saturate - treat it as a survival sentinel, not a judge
- Add a single-point-suspension test to acceptance - the ground is a free stabilizer that hides divergence
- Two simulators disagreed 1.8x on one policy's torque demand - four explanations were eliminated with numbers, the surviving suspect was never tested, and neither reading was allowed to cancel the other
- A torque-tail penalty was paid for by bracing the legs against each other - the second simulator's leg-contact count caught it, and the first explanation ("the trainer can't see self-collision") was retracted from the run's own config
- Every acceptance run records video of the very rollout that produced the numbers - the seated basin, edge-standing feet, tangled legs and the narrow stance were all seen on video before, or instead of, a metric catching them
Training runs 21
Curricula, resuming a run, and which checkpoint to continue from.
- Record-high training reward hid a fully-failing DR subgroup - aggregate metrics average over draws, gates must test per condition
- An auto-curriculum ratchet capped out at iter 248 and never engaged - stage difficulty manually or verify engagement
- Anchoring the action on the measured joint angle (target = q + beta*a), with a beta curriculum down to tau_limit/kp, bounded torque by construction, removed the re-falls and later stood the robot up on hardware
- Raising a command bucket's share does not strengthen its per-state gradient - it only starves the other modes
- Adapting a lineage to one plant increment needs hundreds of iterations, not thousands - long runs only buy specialization
- Continuing a converged policy on a change that carried no new gradient drifted its transfer from 100/98% to 80/28% over 3,000 iterations while every Isaac gate stayed perfect - scan every checkpoint on the second simulator's friction axis
- A curriculum ramp keyed to the process step counter re-fires on every resume - and shipped policies never saw the penalty
- A pull-assist curriculum keyed to a global success share was satisfied by the categories that already worked and withdrew before prone learned anything - conditioning the criterion on prone took it from 2.5% to 98.7%
- Training the final recipe from scratch in one run - every mechanism the lineage had accumulated - produced 0% and a seated robot; the order in which the lineage acquired those mechanisms was part of why it worked
- The fallen-state reset was designed, not sampled from SO(3) - fixed category shares with jitter, a low drop that settles physically, equal left/right shares for mirror augmentation, and a numeric check before training
- PPO's Gaussian noise cannot compose phase-locked oscillations - deliver them as feed-forward and let the policy learn the residual
- Fine-tuning through a reward-table change was falsified (scatter, half-recover, collapse) - continuation training is legal only with the reward frozen
- Choose the fork root by which candidate's shortfalls are recoverable, not by headline score
- Curriculum-gate a penalty to the phase where its disease occurs - early on it only taxes exploration
- Narrowing the speed range to stop high-speed falls entrenched crouch-shuffling - judge gait quality at the speed that demands a gait
- Late training leaned on sampling noise as a stability crutch - deterministic play collapsed while training metrics stayed green
- Every rung gets written stop criteria - hit any one, stop; tightened when priors say results should come fast
- Before resuming a checkpoint, diff the current cfg against what the checkpoint was trained with
- The best checkpoint to SHIP is not the best checkpoint to CONTINUE FROM - maturity is capital against adaptation shock
- A training-side rate limit anchored on the last commanded target is an integrator inside the balance loop - two unrelated lineages converged to the same 34-43% re-fall rate, a soft penalty could not fix it, and the bandwidth arithmetic said safety and standing could not coexist
- Under continuous 3-axis uniform sampling, pure straight-line walking is a zero-measure event the policy never trained
Experiment method 27
Changing one thing at a time, and tracing a result back to its cause.
- The advisor's "runtime five-stage state machine + per-stage reference poses + RL residual" appeared in none of the three papers it cited - reading the originals changed the plan and downgraded two widely repeated industry claims
- Symmetrizing the config made the gait MORE asymmetric - the asymmetry lived in the policy weights
- Drop the frozen policy into chosen configurations - a squat 2.7 cm lower than the stuck pose stood 52% of the time, the stuck W-sit 0%, and the interpolation between them showed a wall, not a slope
- Freeze the deployment contract, stamp every export, and let an automated checker catch wiring bugs
- When hardware underperforms, audit deployment knobs before prescribing retraining
- Training at scale 0.4 is NOT the twin of deploying 0.5 at power 0.8 - the algebra matches, the learned policy does not
- Changing the gait clock silently flipped a hardwired threshold's meaning - write derived constants as expressions
- Sort external training advice into adopt / already-have / modify / would-trap by recomputing it on your own config
- Design the next run to complete the 2x2 - either outcome then convicts or acquits a factor cleanly
- After seven patch-generations, freeze the lineage as a regression baseline, fix the structural debts, and retrain from zero
- Close a question with an audit, then freeze the wording - later symptoms may not reopen it without new hard evidence
- Diagnose a behavior failure by enumerating hypotheses and auditing each against the actual config, cheapest first
- Compare the achieved reward to the computed ignore-floor to tell "never learned" from "learned but unprofitable"
- Never referee a suspect metric with another metric from the same code - they can share the disease
- Low-friction robustness traced to kd DR bandwidth, not friction training - by digging resolved params across 8 lineages, 3840 cells
- When training fails repeatedly, inject the target behavior open-loop - stop tuning rewards for an unverified behavior
- Multiple changes may share one rung only if their symptom spaces are orthogonal - with the ablation order written in advance
- Real robot walked at half the sim clock for two generations - resolved by racing a reward-side and a plant-side evidence line, not by guessing
- The standing-pose reward had been pulling toward the narrow stance the whole line was fighting - a zero-training kinematic audit of the target vector found it, after first auditing the wrong nominal
- Pre-register the ladder's risks and how each future result will be read - before training
- Fall recovery was defined as the whole chain - any fallen pose, a stable stand, a clean hand-back to walking - and built as a second policy behind a deploy-side switch, not folded into the walking PPO
- Verify changes in the run's resolved config (and checkpoint md5), never in the source you edited
- Four in-lineage attempts to widen the standing stance failed - remove a tax, add a joint-space knife, change the target, add a task-space metric penalty - because the stance was the end state of the get-up path; trained from scratch with the right terms it grew right from day one
- Four consecutive fixes were each continued from the previous fix's degraded state until the user stopped the ladder - "change parameters, don't stack errors" - rolled back to the last good checkpoint and audited the target geometry first
- Foot dragging is an attractor, not a low amplitude - and joint damping is the mode switch, adjustable at deploy time
- Fix the task first, harden the plant second - DR budget spent on a dying task is wasted
- A plateau in a training curve was a population mix, not a half-learned skill - 60% standing at 0.372 m and 40% sitting at 0.19 m - and a category at hard zero stayed at zero through 3,000 more iterations
Principles 22
Rules for running training experiments, each learned from a case. A report cites them as Principle 1, Principle 2 and so on.
Contract freeze and fingerprint discipline
The policy I/O contract (observation layout, scales, history semantics, action pipeline) is frozen and fingerprinted; every exported policy is stamped and verified; contract changes ship as new versioned profiles that leave old artifacts bit-identical, and old policies run forever under their era's pinned profile.
The case. The 215-dim omni contract was frozen with a three-machine digest; the one contract-level extension (lateral feed-forward) went in as a new `omni_ff` profile with the old profile provably untouched, and the contract checker caught two real wiring bugs before any training (`contract-freeze-and-checker`). A silently changed gait-clock default would have fed old policies a 25% slower clock - closed by pinned legacy profiles (`legacy-profile-pinning`). A stale derived USD forked plant mass 2.2% until an automated source-vs-derived instrument gated it (`derived-asset-staleness-check`). A gain profile is part of the closed loop a policy was trained in and belongs in its stamp; the recovery line's anchored authority was left out of its manifest and recorded as the gap not to repeat (`gain-profile-belongs-in-the-stamp`), and a second policy behind a deploy-side switch made the handoff state itself a contract (`recovery-two-policies-and-a-state-machine`, `walk-recovery-fsm-handoff`).
How the Coach applies it. On any proposal touching obs/action semantics, defaults, or derived assets: demand the version/profile plan, the fingerprint update, and the checker extension in the same change; flag any old artifact that would run under new defaults.
Attribution by resolved training params - never eval-override knobs
Capability differences between lineages are explained only by digging each lineage's *resolved* training configuration and eliminating columns; evaluation-side override knobs (kd-scale, power-scale, cycle-time) act on the plant for *every* policy and may serve as deployment mitigations but never as explanations.
The case. Low-friction robustness across 8 lineages x 3840 cells was traced to kd DR *bandwidth* - every lineage had ground friction pinned to (1.0,1.0), so "trained friction" could not be the axis; the parameter axis and the plant axis were explicitly separated after the first attribution conflated them (`kd-bandwidth-mu-law-attribution`). "Weak turning" on hardware was a power-scale plant effect, not a training gap (`deploy-knob-attribution-before-retraining`); slowing the deploy clock was out-of-distribution, not a feature (`cycle-time-override-is-ood`). The ground truth for what a run trained under is the logged per-run config, not the source tree (`resolved-config-is-source-of-truth`).
How the Coach applies it. Whenever asked "why is lineage A better", require the resolved-param table first; kill zero-variance columns; refuse explanations phrased in eval-knob terms; when a knob helps, label it deployment mitigation.
PASS gates become constraints; FAIL gates become objectives
Once a skill passes its gate, that gate converts into a standing regression constraint (budget <= 2/20 against the parent baseline) for all later training; gates currently failing are the only legitimate objectives of the next rung.
The case. The C ladder ran one frozen 13-cell x 20-seed matrix at every rung with promotion = "new skill PASS and old skills within regression budget"; C1 was stopped and re-rooted precisely because it trained away the root's backward PASS (`fixed-acceptance-matrix-per-rung`, `preregistered-stop-criteria-per-rung`). The C4 product shipped only at 260/260 cells with zero regression.
How the Coach applies it. Keep the ledger: every PASS adds a constraint row; propose rungs only against FAIL rows; treat any constraint violation as stop-and-attribute, never "the next rung might win it back".
One variable per ladder rung - counted against what the checkpoint saw
A rung changes one variable, where "one" is counted against the checkpoint's actual training state, not against the current config's diff; batching is allowed only when each change owns a disjoint symptom space with a pre-registered ablation order.
The case. Two rungs failed identically because resuming s1e-500 under the evolved config silently added four plant variables the checkpoint had never seen ("单变量纪律不只看「我改了什么」,还要看「checkpoint 见过什么」" - `resume-state-dr-audit`). v8 legally batched four orthogonal fixes with a written ablation order (`orthogonal-batch-with-ablation-order`); v9 spent one run completing a 2x2 factorial so either outcome convicted a factor (`fill-the-missing-factorial-cell`); v10b's three-way ablation wrongfully convicted the clock and had to be retried fairly.
How the Coach applies it. Before any resume: diff cfg against the checkpoint's logged training state. Before any batch: require the symptom-ownership map and ablation order in writing.
Pre-register risks, readings, and stop criteria before the ladder
Before a ladder or risky rung, write down the known risks, the interpretation of every plausible outcome, and hit-any-one stop criteria - frozen before training, tightened when priors say results should come fast.
The case. The C ladder opened with three numbered risks including the exact falsification condition for its own root choice; A/B arms carried "预注册读法(事后不改)" tables; a level expected to fail was run anyway for its pre-registered diagnostic value (`preregister-risks-and-fork-readings`). Stop criteria caught C4-redo rungs at +200 instead of full caps (`preregistered-stop-criteria-per-rung`); hardware sessions pre-registered per-config expected signatures and the disagreement rule "不改结论改账" (`preregistered-real-expectations`, `feasibility-accounts-lock-design-point`).
How the Coach applies it. Refuse to open a rung without the written risk/reading/ stop block; after results, read conclusions off the pre-registered table and flag any post-hoc reinterpretation.
Plant parameters are measured, never invented
Every plant number carries measurement provenance: armature = N^2 x rotor inertia from no-load tests, friction split by rig and by API column, torque limits shaped by per-joint gait peaks, latency traced through the real pipeline, masses weighed - and DR bands are additive around the measured nominal, sized to the measured dispersion.
The case. Guessed friction was 2.5x low and guessed damping 5x high (`friction-measured-not-guessed`); armature had been 0 with a 9:1 gearbox (81x reflected inertia, `armature-n2-rotor-inertia`); a uniform torque derating was "the wrong shape" vs measured peaks (`torque-limit-shape-by-measured-peaks`); the delay implementation itself was a wrong plant for a whole lineage (`latency-lerp-reverse-extrapolation`); the run design point was locked by three accounts including the tau_limit/kd speed ceiling (`feasibility-accounts-lock-design-point`); identified friction had to land in the right simulator API columns to act at all (`sim-api-friction-columns`). The recovery and one-leg lines opened with the same kind of accounts before any reward existed - a connected static path and the torque along it for an armless get-up, and the gains single support needs to be holdable at all (`get-up-feasibility-accounts-before-training`, `single-support-gain-authority-probe`).
How the Coach applies it. For any plant value in a config review, ask "measured how?"; reject absolute ranges with no nominal; check API column mapping and derived-asset regeneration whenever measured values land.
Sim2sim gate before sim2real - under deployment conditions
Every checkpoint passes a second, independently built simulator before hardware, and both the gate and the smoke loop run under the measured deployment conditions (real pipeline delay, honest contact parameters, the deployment gain/power profile).
The case. The standing order "先sim2sim 再sim2real" (`sim2sim-gate-before-sim2real`); acceptance flipped to match hardware only under measured condim/torsional friction (`eval-plant-honesty-contact-params`); gates moved permanently to `--delay 2` after the kicking incident (`pipeline-latency-is-plant-not-dr`); and the harness itself must be audited - a frame-convention bug in the cross-sim evaluator invalidated a whole line of verdicts (`body-frame-velocity-api-audit`). The recovery line's second simulator caught a torque penalty paid for by bracing the legs together (`torque-penalty-bought-by-leg-bracing`), and a 1.8x torque disagreement between the two plants stayed binding because its one surviving explanation was never tested (`torque-disagreement-between-simulators-unresolved`).
How the Coach applies it. Block any hardware request lacking a second-sim PASS at deployment conditions; when sim2sim and training-side metrics disagree, treat the evaluator as a suspect too.
Observation honesty - the actor's inputs are a hardware contract
The actor observes only signals the real robot produces with realistic noise; privileged truths go to the critic; history windows are estimators and must train under plant variation; rewards on quantities the actor cannot observe buy only average suppression, never closed-loop correction.
The case. Ground-truth velocity/forces went critic-only (`observation-honesty-critic-only`); frame_hist under zero DR memorized the trainer's plant fingerprint - 0/20 transfer (`history-obs-needs-plant-variation`); world-frame yaw rewards could not teach pull-back because heading is unobservable to the actor - correction was routed to the deploy outer loop instead of breaking the contract (`reward-observability-limit`, `deploy-heading-loop-and-align-training`).
How the Coach applies it. Audit every actor-obs element for hardware existence; require minimal plant jitter whenever history/recurrence exists; for each reward, ask "can the actor see this error?" and route correction tasks to outer loops.
Reward economics are audited in realized currency
Reward design decisions are made on realized per-step magnitudes under the actual policy and command distribution: price the do-nothing optimum before adding a mode, compare achieved values to the computed ignore-floor, calibrate thresholds between measured healthy and sick distributions, and ship every new penalty with a withdrawal clause.
The case. feet_air_time at weight 2.0 realized 0.038 vs tracking 1.2 - drag was rational (`realized-contribution-audit`); ignoring a vy command cost 28-180x less than ignoring vx until a gated tracking term was added (`reward-cost-of-ignoring-audit`, `gate-new-reward-terms-by-command`); achieved-vs-floor separated "never learned" from "priced out" (`ignore-floor-diagnosis`); the foot-distance wall was placed between measured healthy (0.6% tax) and sick (55%) policies (`calibrate-threshold-between-healthy-and-sick`); the landing penalty carried a pre-registered stand-down condition and actually stood down (`calibration-threshold-with-withdrawal-clause`); two clearance terms were inert until zero-points and gate occupancy were checked (`inert-reward-term-audit`). A get-up policy sat because three gated terms paid the seated pose 84% of the return and the one term that could tell sitting from standing was an exp kernel reading 4.6e-5 at the real error (`seated-basin-dead-exp-kernel`); a torque-tail term was weighted by its measured steady value beside a peer term after the estimate proved 12x off (`tail-torque-needs-hinge-on-computed-demand`).
How the Coach applies it. Never discuss weights in the abstract: demand the realized-contribution table, the ignore-floor number, and the healthy-pay calibration before any reward edit is approved.
The zero-cost option must be the desired behavior
For every penalty, name what the zero-cost option is; penalize failure events (slip, saturation excess, contact in flight windows), never the motion or joints that healthy behavior uses; make degenerate strategies fatal via termination where penalties cannot price them out.
The case. Joint-usage penalties for drift taxed a 1.4%-of-momentum channel 2.7/step and collapsed training; the slip penalty costs a non-slipping gait exactly zero (`penalize-the-slip-not-the-joint`). A frozen-at-clamp joint pays zero action-rate forever - only a pre-clip saturation penalty flips the cheat economics (`saturation-cheating-zero-rate-cost`). Ungated phase shaping made standing 42x more expensive than stepping and cooked the hip motors (`moving-gate-42x-stand-tax`); crouch-shuffling lived until a height termination deleted it (`termination-closes-degenerate-basin`). A gated penalty is an exit: the policy parked just outside an uprightness gate, then just under a height gate, to stop paying a stance tax, and only a positive band plus an always-on guard closed both (`penalty-gate-is-an-escape-hatch`); a soft-limit penalty that charged the standing pose itself bought a 4.1 deg lean (`soft-limit-penalty-charges-nominal-pose`); an unpriced foot attitude was spent on edge-standing (`unpriced-foot-attitude-is-a-free-variable`); and the one-leg line listed its cheapest cheats before training and still met one through a zero-gradient band (`enumerate-cheapest-cheats-before-training`, `binary-band-reward-fake-touchdown`).
How the Coach applies it. Run the "零代价的选项是什么" audit on every proposed term; convert motion taxes into event-conditional penalties; check the termination set against each known degenerate strategy.
Measurement discipline: independent referees, signs, distributions
A disputed measurement is adjudicated only by an independent algorithm from raw state; directional ability requires sign-antisymmetry under command reversal; bimodal metrics are reported as mode shares (never medians, never 3 seeds); ratios are not comparable when totals change; reward values compare only within one command distribution; single chaotic events never cross machines.
The case. The triple reversal - a good metric was "refuted" by a sibling metric that shared the disease (`independent-referee-for-metric-disputes`, `body-frame-velocity-api-audit`); same-signed +/- responses were bias, not turning (`same-sign-response-is-yaw-bias`); the swing median sat in a bimodal gap (`median-hides-bimodal-distribution`); "v6 is jitterier" died on absolute energies (`ratio-metrics-need-absolute-check`); yaw gain measured 15x wrong in an oscillating frame (`heading-integral-not-body-rate`); a 44% improvement evaporated under same-distribution comparison (`same-distribution-reward-comparison`); drift direction was a limit cycle (`multiseed-sign-test-for-drift`); a cross-machine push cliff was chaos (`single-impulse-recovery-is-chaotic`).
How the Coach applies it. Before accepting any surprising number: ask for the independent recomputation, the sign pair, the distribution shape, and the comparison conditions. Retract in writing when a metric falls.
The deployment pipeline is plant
Irreducible pipeline properties - action latency, rate limits, power/torque scaling, teleop command mappings - are part of the nominal plant, modeled from day one and reproduced in every gate; deploy-side scalings are crutches that flag unmodeled plant, and they cannot be algebraically folded into training constants.
The case. Right-leg kicking was over-trained-delay x loop gain; power 0.8 was a gain-reduction crutch that retired when the delay was modeled (`pipeline-latency-is-plant-not-dr`); power derating damages non-forward axes first (`power-scale-hurts-nonforward-axes`); training at 0.4 scale as the "twin" of deploying 0.5 x 0.8 collapsed 0/20 (`deploy-scaling-not-training-equivalent`); one shared teleop speed sent an out-of-band lateral command and the robot clipped its own foot (`teleop-command-band-per-axis`); the latency DR range had not even covered the measured pipeline (`latency-dr-covers-measured-pipeline`). A rate limiter added at deployment only clipped a policy that kept commanding (`deploy-rate-limiter-windup`); moved into training and anchored on the last command it became an integrator in the balance loop (`slew-anchor-is-an-integrator`); anchored on the measured angle it bounded torque and kept the bandwidth (`beta-anchored-action-target`). The walking lines' safe setting, power-scale 0.8, cut the ends of the recovery policy's full-range travel and left its spikes alone; a gain inside the trained band did the job (`power-derating-cuts-full-range-contract`).
How the Coach applies it. Demand the measured pipeline latency/limits in the plant model and in gate conditions; treat every deploy-side derating as a question ("what is this compensating?"); block per-axis command sources that exceed training bands.
DR budget is finite; its distribution is the measured support
Robustness is a conserved budget: disturbance training on an already-hardened lineage borrows from existing margins; DR ranges span the measured deployment support - no fictitious tails (they buy degenerate gaits), no single constants (they allow thin-margin specialization); harden the plant only after the task distribution is final.
The case. The same push dose helped a narrow lineage and damaged a balanced one - budget conservation (`push-dr-conditional-budget-conservation`); wide latency tails bought drag-glide, constant values shipped 60% thinner tilt margins - the answer is a narrow band on the measured support (`dr-tail-plant-continuation`, `constant-value-dr-overfits-margin`); task-first ordering because hardening a soon-to-change task wastes budget (`task-shaping-before-plant-hardening`); COM randomization used deliberately as a behavior-shaping tool, and rolled back on symptom per its own contract (`com-randomization-forces-leg-spread`, `com-dr-rollback-on-symptom`). DR that is switched on can still be thin: the run policy fell in the frontal plane its gain-and-latency randomization never touched (`thin-dr-judged-by-channel-coverage`), and a friction priority settled under one action contract had to be re-measured under the next (`friction-priority-re-measured-after-plant-change`).
How the Coach applies it. Before any DR rung: check the untrained policy against the spec, the lineage's current DR load, and the measured real-world range; after it: audit retained margins, not just the new tolerance.
Gates measure what hardware feels: posture, margins, stripped assists
Acceptance batteries carry posture-class rows (tilt max median, per-joint L/R asymmetry, temperature) beside task rows, graded margin columns beside binary gates, chirality scored per side, at least one condition that removes the environment's free stabilization, and validated predictive scalars promoted into the gate.
The case. Three same-shaped judging errors - survival, displacement, wz-difference - all missed what the operator felt; posture metrics had the predictive power (`task-metrics-vs-posture-metrics`, `stand-gate-posture-not-survival`); binary survival saturated and hid a 60% margin gap (`constant-value-dr-overfits-margin`); v5 passed everything on the ground and failed suspended (`suspension-probe-removes-free-stabilizer`); the hip_roll (l+r) scalar predicted real drift direction and ordering and entered the battery (`hip-roll-sum-predicts-lateral-drift`); averages hide chirality (`chirality-scored-separately`); gait-quality gates are judged at speeds that demand a gait (`low-speed-commands-reward-dragging`). The recovery line added the rest of the kit: where failed episodes end, not only where they started (`end-state-confusion-matrix`); a frozen acceptance distribution with a pinned seed (`frozen-acceptance-distribution-and-pinned-seed`); video of the metric rollout itself (`video-as-acceptance-record`); and the admission that a 10 s episode cannot see a stance that fails after a minute (`episode-length-bounds-what-a-gate-sees`). The one-leg line removed a foot-spacing wall that no gate measured, and the feet met on hardware (`removed-wall-returns-on-hardware`).
How the Coach applies it. Review every battery for posture rows, margin columns, per-side scoring, and an assist-stripped condition; when operator feel and gates disagree, suspect the metric class first.
Fork and root selection: recoverability, maturity, frozen rewards
Choose fork roots by which candidate's deficits the coming training can pay back (precision is recoverable; lost plasticity, symmetry, and margins are not); prefer mature checkpoints as roots even when younger ones score better as products; never fine-tune through a reward change - continuation is legal only with the reward frozen and plant/DR widening one rung at a time.
The case. s1e-500 beat higher-precision candidates because its exclusive strengths were unrecoverable (`fork-root-recoverable-shortfall`); the b300 arm proved maturity is capital against adaptation shock (`root-maturity-vs-product-quality`); the B-arm scatter/half-recover/collapse signature falsified reward-change fine-tuning and drew the legal boundary for S2 continuation (`fine-tune-reward-change-falsified`).
How the Coach applies it. For root debates, build the exclusive-strengths table and ask "which side can be trained back?"; require dual-arm evidence for maturity claims; classify any proposed continuation as reward-frozen or not before approving.
Curricula: verified engagement, lineage counters, disease-phase gating
Automatic curricula must prove they engage (a saturated ratchet is constant DR wearing a curriculum's name); every ramp counts lineage-cumulative progress, not per-process steps; penalties aimed at late-stage pathologies ramp in after exploration noise decays; difficulty rises on measured per-stratum success, never on schedule.
The case. The s1f ratchet capped at iter 248 and never engaged (`auto-curriculum-engagement-check`); the saturation ramp re-fired at +600 after every resume and no shipped product ever saw the penalty (`curriculum-counter-lineage-steps`); the same penalty worked once gated to the disease phase and became an untouchable mechanism (`gate-penalties-to-the-disease-phase`); record-high aggregate reward hid a fully-failing delay stratum (`aggregate-metrics-mask-subgroup-failure`); bucket share is not a gradient lever (`bucket-share-is-not-a-gradient-lever`). An assist curriculum keyed to a pooled success share was withdrawn on the strength of the categories that already worked (`curriculum-criterion-conditioned-on-lagging-category`); a pace set by per-step income moved only when that income was time-gated (`per-step-income-drives-speed-time-gate`), and the same gate had to be retired in a lineage without the disease (`time-gate-vs-wide-stance-retire-the-fix`).
How the Coach applies it. Ask every curriculum three questions: does it engage (show the internal state)? what does it count (process or lineage)? when is it present (against the pathology's phase)? Check where shipped checkpoints sit relative to every ramp.
Probe before training: feasibility first, hypotheses in tables
After two failed training attempts at a skill, stop training: demonstrate the behavior open-loop, enumerate hypotheses in a written table audited against actual configs cheapest-first, race one probe per side of the sim2real boundary for hardware-only pathologies, and use suspended tests to acquit or convict actuators before blaming authority.
The case. "在黑暗里试钥匙" - four sidewalk rungs failed until an open-loop probe separated exploration/waveform/authority in one experiment (`open-loop-probe-before-reward-tuning`); the foot-drag mystery fell to a seven-hypothesis config audit (`hypothesis-table-code-audit`); the period-doubling was resolved by racing a reward-side and a plant-side evidence line - and both paid off, one per sub-case (`period-doubling-evidence-race`); the suspended test acquitted the roll actuator in one measurement (`suspended-test-isolates-actuator-authority`). A read-only configuration probe told a wall from a slope in the recovery line's seated basin (`configuration-probe-wall-not-slope`), and the fix it pointed to - where the feet are - took prone from 0/159 to 158/159 (`prone-dead-end-is-foot-placement`); a knob that did not move its variable was recorded as no test of the idea (`dof-vel-penalty-is-not-a-pacing-knob`).
How the Coach applies it. When a skill resists training, prescribe the probe before any further reward edits; require verified target trajectories before imitation terms; keep a falsified-fixes list so closed roads stay closed (`amplitude-cut-falsified-yaw-fix`).
External advice is recomputed locally; values transfer as ratios
Every external suggestion is classified adopt / already-have / modify / trap by recomputing its claim on the local reward table and probe data; numeric values transfer only as dimensionless ratios (to tracking weight, leg length, sqrt(gL), control rate); citations are verified to exist.
The case. "Start vy very small" would have destroyed sidewalk learning on this reward table - the gradient scales quadratically (`external-advice-audit-against-own-arithmetic`); swing-height targets and weights transferred correctly only through leg-length and tracking-ratio scaling (`transfer-ratios-not-absolutes`); the "6-step delay" was refused for lacking a control rate (`latency-dr-covers-measured-pipeline`); a borrowed reference's structure was FK-verified and its amplitude re-derived from the division of labor (`reference-structure-fk-amplitude-division`); retrieval agents fabricated verbatim arXiv quotes - only source-verifiable material was used; and one dismissed suggestion later proved right for a different mechanism, and was credited (`cycle-average-tracking-for-gait-quantities`). An advisor's staged state machine turned out to exist in none of the three papers it cited, and reading them changed the plan (`advisor-paraphrase-vs-paper`).
How the Coach applies it. Intercept every "paper X does Y" with the local recomputation; convert absolutes to ratios before comparison; verify quotes; revisit dismissed advice when new mechanisms appear.
Hardware sessions are scripted experiments, not tuning sessions
Real-robot time executes a pre-registered matrix: risk-ordered (baseline first, fragile last with a spotter), stage-gated (suspended smoke before ground), A/B sessions bracketed by a repeated reference run, operators briefed on measured zero-command and untrained-axis behavior, chirality-aware disturbance protocols, no field tuning - the only legal field changes are scripted, single-variable, and self-reversing.
The case. The S2 acceptance sheet (`risk-ordered-real-deployment`, `battery-bracketed-real-ab`, `know-zero-command-behavior`, `push-test-chirality-protocol`, `no-field-tuning-protocol`); the RAM-only torque experiment with automatic power-cycle rollback (`reversible-single-variable-field-experiments`); and the sim-veto rule - even sim's condemnations get one safeguarded hardware check when they judge the purpose-built configuration (`sim-veto-needs-real-confirmation`). The recovery line's first real run went ahead with its preconditions unmet and was stopped as dangerous (`first-real-get-up-violent-stage-one-policy`); after it: a staged hang, mat and floor protocol (`staged-hang-mat-floor-for-get-up`), a fixed power-cycle pre-flight and two-machine discipline (`power-cycle-preflight`, `two-machine-config-discipline`), a fall guard replaced rather than switched off (`fall-guard-becomes-a-state`), and logs that are part of the run (`hardware-log-is-the-attribution-input`).
How the Coach applies it. Turn every hardware request into a runbook with order, gates, brackets, briefing, and anomaly plays; refuse improvised parameter changes on the floor.
Close questions in writing; restart when the debt is structural
Audited questions get frozen verdicts with citable wording and an explicit reopening bar; hardware verdicts are dated by deployment-stack and calibration state and expire when those change; and when successive rungs shuffle symptoms without net progress, freeze the lineage as regression baselines, pay the structural debts, and retrain minimal - carrying laws and instruments, not weights.
The case. The chirality and COM questions were closed with frozen wording and "no reopening without new hard evidence" (`frozen-verdicts-semantic-boundaries`); v5/v6's condemnations expired with the deploy stack (`stale-verdicts-under-old-stack`); a 2-degree calibration fix moved the whole runnable envelope (`zero-offset-calibration-shifts-envelope`); plant upgrades are era boundaries with paired re-baselining (`plant-swap-invariants-vs-shifts`); and the 2026-08-05 reset froze v5-v11, fixed the latency FIFO / manifest / sampling / reward-table debts, and restarted - producing the lineage that reached hardware SOTA (`freeze-lineage-fix-structure-restart`, `minimal-reward-table-with-provenance`). The recovery line's real-robot verdicts ended up in three places that disagree, one of them an undated note in a command file (`write-hardware-verdicts-back`).
How the Coach applies it. Maintain the closed-questions ledger and quote it when symptoms recur; stamp verdicts with stack/calibration versions; when a team is three rungs into symptom-shuffling, raise the restart question explicitly with the freeze-fix-restart pattern.
Name the quantity in the space it lives in
A goal, reward term or acceptance criterion about the feet, the base or the contact state is computed from the quantity itself - world poses, forces, per-category outcomes - never through a joint-angle, single-signal or pooled stand-in that assumes everything else sits at nominal; and every detector is validated on a behaviour known not to contain the event before it becomes a gate.
The case. The recovery line was caught three times: |ankle roll| as "flat feet" sold stance width and the real robot slid into the splits, a hip-roll criterion was confounded by 50 deg of yaw, and the joint table said 0.271 m where the feet were 0.159 m apart; task-space terms produced the first flat, wide stance (`joint-space-proxy-for-task-space-quantity`). Flight detection lied in both directions across two lines - foot height flagged 40% false flight on a walking gait, contact force alone flagged slip chatter as hops (`contact-detector-single-signal-lies`). A pooled height average described a robot that did not exist - six in ten standing, four in ten sitting (`zero-partial-credit-is-not-an-iteration-problem`) - and the walking line had learned the same lesson on yaw rate (`heading-integral-not-body-rate`).
How the Coach applies it. For every reward term and gate row, ask what physical quantity it stands for and whether it is measured directly; flag joint-space or single-signal stand-ins for task-space goals, ask for a detector validated on a negative control, and split pooled metrics by category before reading them.
Continuation needs a live gradient; a release is chosen by a scan
Continue a converged policy only on a change that creates a live gradient, on a short budget, with every checkpoint scanned on the transfer axis; choose a release by running the full battery over a band of checkpoints and stop on signals, never by taking the last one; and when edits to the terminal phase cannot move a behaviour, roll back and retrain with the constraint present from the start, keeping the order in which the lineage acquired its mechanisms as explicit curriculum phases.
The case. A continuation with no new gradient drifted MuJoCo transfer from 100/98% to 80/28% while every Isaac gate stayed perfect, and a live-gradient continuation at the same depth kept it (`converged-continuation-is-poison`). One-leg checkpoints 100 iterations apart failed 1 and 38 of 40 cells, and late ones degraded (`checkpoint-choice-is-a-full-gate-scan`). Four in-lineage stance fixes failed because the stance was the end of the get-up path, and from scratch it grew right (`stance-decided-by-get-up-path`); fixes stacked on degraded states were rolled back by the user (`stop-stacking-roll-back-and-audit`); and the lineage's final recipe, trained from scratch in one run, sat at 0% because the order of its curriculum was part of the product (`curriculum-history-is-part-of-the-product`). The omni line's short adaptation budgets and mature roots are the same law seen from the other side (`continuation-budget-not-from-zero`, `root-maturity-vs-product-quality`).
How the Coach applies it. Before approving a continuation, ask for the new gradient, the budget and the transfer axis in the scan; before approving a release, ask for the scan; after three rungs without progress on the target, propose rolling back to the last good checkpoint and a from-scratch phase plan instead of a fourth patch.