Skip to content
MouseMouse

Before adding a command mode, compute what ignoring it costs - the lazy optimum must lose

A lesson the Coach cites as reward-cost-of-ignoring-audit.
Mechanism understoodReward design

Price the do-nothing policy for every new command or objective: compute reward-per-step for "comply" vs "ignore" from the actual table, and only train once ignoring is decisively unprofitable.

Symptom

A new command mode can be silently unlearnable if the reward table makes "ignore the command entirely" nearly free compared to the tracking reward available elsewhere.

Context

For each C rung the team computed the per-step cost of completely ignoring the new command versus ignoring forward: ignoring vx=0.25 costs 1.264/step; ignoring wz=0.20 costs 0.984/step (78% of forward - gradient sufficient, so C2 was certified "zero reward surgery"); but ignoring vy=0.10 cost only 0.020/step - 50-60x weaker, because vy entered the table only as an L2 tax, not a tracking term.

Change

Rule instituted: a rung may claim "no reward change needed" only after this arithmetic shows the ignore-cost is the same order as forward's. For C4 the audit failed, so track_lin_vel_y_exp (+2.0, std 0.15, same form as vx) was added - raising the ignore-cost at cmd_vy=0.10 from 0.020 to 0.718/step (36x).

Outcome

C1/C2/C3 proceeded with zero reward edits, keeping single-variable attribution clean; C4's needed surgery was identified before training instead of after a failed run.

Mechanism

PPO converges to whatever costs least; if the reward margin for obeying a new command is a rounding error against existing terms, the "ignore" policy is the optimum and no amount of training fixes it. The audit prices the lazy optimum explicitly before spending compute.

Applies when

  • adding a command axis or task mode to an existing reward table
  • a new skill trains flat while other skills stay healthy
  • certifying a rung as "no reward change"
“cmd wz 0.20 → 0.984(coarse .473 + fine .491 + L2 .020)… 对照:忽略 vx=0.25 = 1.264(本级 78%,同量级);忽略 vy=0.10 = 0.020(弱 50 倍——那才是 C4 必须加 track_lin_vel_y_exp 的原因)。本级不动奖励表。”
train/C_LADDER_RUN.md § 4. 原地转级(C2)② 奖励梯度已验够

Principles that cite it