Skip to content
MouseMouse

Raising a command bucket's share does not strengthen its per-state gradient - it only starves the other modes

A lesson the Coach cites as bucket-share-is-not-a-gradient-lever.
Mechanism understoodTraining runs

When a skill is not learning, first prove its per-state signal is nonzero (ignore-floor and probe checks); only rebalance sampling shares to fix genuine sample starvation, and account the regression risk to the diluted modes before doing it.

Symptom

Sidewalk was not learning, and the reflex proposal was to give the side bucket a larger share of sampled commands.

Context

The C4-redo3 rung explicitly kept the 20/40/20/20 bucket (stand/forward/turn/side) with the reasoning written out: PPO computes advantages per state, so bucket proportion does not change the per-state gradient of side states; at 4096 envs x 20% x 24 steps the rollout already contained ~19.7k sidewalk states - sample count was not the bottleneck. And the cost side was already measured: cutting forward from 60% to 40% had made vx+0.30 die at +400 in an earlier run - more cuts would only collapse it sooner.

Change

Bucket proportions held constant across the entire C4 redo series; the actual bottlenecks (metric frame bug, reward variance penalty, exploration form) were pursued instead.

Outcome

Sidewalk was eventually fixed with zero bucket changes (feed-forward delivery, +100 iters); forward/turn skills never suffered starvation-induced regressions during the redo series.

Mechanism

Policy-gradient credit is assigned per visited state; oversampling a mode multiplies its states in the batch but not the informativeness of each, so if the per-state gradient is ~0 (behavior unreachable or reward indifferent), N times zero is still zero - while the displaced modes genuinely lose data and regress.

Applies when

  • proposing to oversample a failing task/command mode
  • a majority mode regresses after share rebalancing
  • budgeting env count vs mode share for a multi-skill policy
“比例不动:PPO 逐状态算优势,桶占比不改变单状态梯度;4096 env × 20% × 24 = 每 rollout 已有 1.97 万个侧走状态,样本数不是瓶颈;而 forward 60%→40% 已实测让 f30 在 +400 处死掉,再加码只会更早塌。”
train/C_LADDER_RUN.md § 3i. 桶 20/40/20/20 不动(比例不动)

Principles that cite it