Skip to content
MouseMouse

A stronger action_rate penalty cut the median torque demand under the gate and left the p99 at 4x the limit - only a hinge on the pre-clip (computed) torque, weighted by comparison with a peer term, collapsed the tail

A lesson the Coach cites as tail-torque-needs-hinge-on-computed-demand.
Mechanism understoodReward design

Judge actuator demand against the deployed limit, read the pre-clip demand (applied torque is censored and gives no gradient on the excess), use an L2 rate penalty for the median and a thresholded hinge on computed demand for the tail, and set a new term's weight from its measured steady magnitude next to a peer term rather than from a back-of-envelope estimate.

Symptom

After R0.5 hip_pitch delivered torque sat at its 12 N*m limit in a typical get-up (demand 119-125% of the limit, p99 4.2x) - zero control margin at exactly the moment modelling error matters.

Context

The 12/17/11 N*m limits are deployment limits written into robot.yaml by set_torque (RS06 at 33% of rated), and simulation uses the same effort_limit - so the gate is judged against them, not the 36 N*m rating (an early reading against the rating was retracted). Applied torque is clipped at the limit - censored data - so demand must be read from computed_torque. With the full-range action contract (hip_pitch scale 1.309, kp 30) a single-step action change of 0.306 already saturates hip_pitch, and action_rate penalizes exactly that change.

Change

R3.0: action_rate_l2 -0.01 -> -0.03 (child-run). R3.1: new torque_headroom = sum relu(|tau_computed|/limit - 0.9)^2, normalized so three motor types share a scale. Its weight was first estimated at -0.1, measured in a 12-iteration run at an effective -0.019 (12x smaller - the estimate had mixed a per-episode-peak p99 with a per-step p99, and at 1% of upright it would have been numerically absent), and set to -0.5 so its steady value (-0.095) matched action_rate's (-0.097).

Outcome

R3.0: sum |da|^2 -64%, success 99.6 -> 100%, delivered median = demand median (the clamp no longer fired in a typical episode), gate PASS at worst 79.9% - but p99 unchanged (hip_pitch 419-432% -> 427-436%). R3.1: p99 hip_pitch -> 148-189% (-57 to -65%), knee 422-439% -> 233-234%, saturation duty -60%, success 100%; the worst joint became hip_roll at 67.7%. Its cost appears in torque-penalty-bought-by-leg-bracing.

Mechanism

A squared-rate penalty presses the whole-episode sum and moves the typical step, not rare spikes; the spikes came from the kp term (large targets while a limb is blocked by the ground - velocity alone could not reach them under vel_limit), and a penalty on applied torque cannot see demand above the clip because every excess sample reads as exactly the limit.

Applies when

  • torque demand saturates actuator limits in high-effort skills
  • a smoothness penalty improves medians but not peaks
  • a new reward term's weight is set by estimate alone
“**必须用 `computed_torque` 而不是 `applied_torque`**:后者被 `effort_limit` 削平, 是删失数据,超限样本全被压成"恰好等于限",对超限部分梯度恒为 0。 … 改按同侪定标取 **−0.5**(稳态 ≈ −0.095,与 `action_rate_l2` 的 −0.097 等量)。”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §24 R3.1(torque_headroom 力矩需求越限罚)

Principles that cite it