Skip to content
MouseMouse

Brief the operator on the lineage's measured zero-command and untrained-axis behavior before handing over the joystick

A lesson the Coach cites as know-zero-command-behavior.
ReplicatedSim-to-real

Before any teleop/demo, measure and write down the policy's zero-command behavior and per-axis competence, label untrained axes explicitly as not-bugs, and set the floor/procedure to accommodate the known drift.

Symptom

A teleop session was about to start on a policy that does not stand still at zero command and has never been trained on lateral commands - behaviors an unbriefed operator would report as bugs or emergencies.

Context

Three measured facts were written into the teleop instructions ("都有 实测依据, 不是猜"): (1) A/D (lateral) keys will get essentially no response - probe-measured sidewalk tracking ~3%, an untrained axis: "这正是 C4 要解决的事, 不是 bug"; (2) no keypress = cmd 0, and this lineage does not stand still at zero command - a three-generation lineage property: paces in place, drifts right ~5 cm/s, net rotation -30 deg/20 s; sim survival is 20/20 (it will not fall) but it walks away slowly, so leave floor margin especially on the right; (3) S (backward) WILL respond - probe-measured 20/20 survival, 67% tracking untrained, which is also why this root was chosen for the C ladder. Plus a keybinding dry-run while suspended before touching down.

Change

Operator briefing became part of the deployment artifact: expected response per key, expected idle behavior with magnitudes and directions, and the distinction between untrained (expected, not a bug) and abnormal.

Outcome

The session proceeded with correct interpretations available in advance; the known zero-command wander was handled by floor margin and start-with-command procedure rather than misdiagnosed on the spot.

Mechanism

A learned policy's off-nominal behaviors (idle drift, untrained axes) are lineage properties, stable and measurable in sim beforehand; operator surprise converts known properties into false incident reports and unsafe reactions. A briefing transfers the measured behavior model to the person holding the controller.

Applies when

  • handing a learned policy to an operator or demo audience
  • the policy idles in a non-stationary way at zero command
  • some command axes are untrained in the current lineage
“A/D 基本不会有反应 —— s1e 从未训过非零 vy, 选根探针实测侧走跟踪率 ~3% … 这正是 C4 要解决的事, 不是 bug。… 不按键 = cmd 0, 而 s1e 在零指令下不站定 —— 血统属性, 三代实录: 原地踏步 + 右漂 ~5 cm/s + 净旋 −30°/20s。”
train/REAL_RUN_S2.md § 附: WSAD 遥控 上机前必须知道的三条

Principles that cite it