Skip to content
MouseMouse

The walking lines' safety setting, power-scale 0.8, broke the recovery policy's full-range contract - it cut the ends of the joint travel (4/50 could not get up) and left the torque spikes untouched; a kp x 0.9 gain profile inside the trained kp band did the job

A lesson the Coach cites as power-derating-cuts-full-range-contract.
Observed onceSim-to-real

A deployment derating knob means something only relative to the action contract: before reusing a line's "safe setting" on a new skill, check what it does to that skill's reachable range and to the term that makes the spikes, prefer a gain change inside the band the policy was randomized over, verify it in simulation, and re-decide when the contract changes.

Symptom

After the violent first real get-up (2026-08-09), the recovery policy needed a gentler setting for its next hardware test, and the walking and omni lines' standard derating - deploying at power-scale 0.8 - was the obvious candidate.

Context

The V0 recovery contract maps actions to absolute targets over the full joint range: a = +/-1 lands exactly on the URDF limits, and standing puts the knee at the clip. Candidates were compared on R3.1 in MuJoCo (5 categories x 10 seeds) on 2026-08-10 before any hardware time was spent.

Change

A new gain profile, rl_kp090 (kp x 0.9, kd unchanged), recorded in robot.yaml as the recovery hardware-test setting, with power-scale 0.8 explicitly banned for recovery.

Outcome

kp x 0.9: 48/50 got up; median torque demand on hip_pitch/knee fell from 120-125% to 100-104% of the deployment limit; leg-leg contact frames 2,152 -> 1,095; the change sits inside the +/-10% kp randomization the policy trained with. power-scale 0.8: 4/50 could not get up, because under the full-range contract it removes the ends of the travel (the deep squat's tucked legs, the straight standing knee), and the torque spikes (kp x error) did not fall at all. When the line moved to the beta-anchored contract, rl_kp090 was declared a V0-era choice that does not fit (beta is calibrated at kp 30) and deployment returned to rl_default; the deploy switch applies power scaling to the walking side only.

Mechanism

A power scale multiplies the action, which under an absolute full-range mapping shrinks the reachable workspace instead of softening the actuator; the spikes come from the proportional term on large errors, which only a gain change reduces - and a gain change inside the trained randomization band stays in distribution.

Conflicts

The undated operator runbook still carries an R3.1 "B comparison" command at power-scale 0.8 beside the rl_default baseline; the sources do not say whether it was written before the ban or was ever run.

Applies when

  • reusing a power, torque or action scale from one skill on another
  • a policy whose actions map to absolute targets over the full joint range
  • choosing a gentler setting for a first or second hardware trial
“kp×0.9 / kd 不动 —— recovery_r3_1 成功 48/50, τ 需求中位 hip_pitch/knee 120~125% -> 100~104% 部署限, 腿-腿接触 2152 -> 1095 帧; ±10% 在训练 kp DR 带内. ⚠️ power-scale 0.8 对 recovery **禁用**: 全 ROM 契约下 0.8 砍的是行程 端点 (深蹲收腿/站直够不到), 实测 4/50 起不来, 且尖峰 (kp·err) 一点不降 —— 它是 walk/omni 的安全档, 不是 recovery 的.”
git:Lucen-recovery@origin/recovery:robot.yaml § gain_profiles 注释: recovery 真机测试安全档 (2026-08-10) / rl_kp090

Principles that cite it