Never referee a suspect metric with another metric from the same code - they can share the disease
independent-referee-for-metric-disputes.To adjudicate a disputed measurement, compute the quantity by an independent method from raw state; never accept a sibling column from the same pipeline as the tiebreaker.
Symptom
A triple reversal on one question: sidewalk sign diagnosis (correct) was retracted using a second metric from the same script, then the retraction itself had to be retracted when that second metric turned out to be the buggy one - two opposite-direction errors on the same problem in one day, both written into the execution sheet.
Context
The probe's net-displacement metric suggested the sidewalk reference sign was inverted. Worried about yaw-drift pollution of net displacement, the author checked the same table's body-frame vy_mean column (~0.003 everywhere, 20-50x smaller) and retracted the sign diagnosis. But vy_mean came from mj_objectVelocity, which was silently reporting vertical velocity due to a frame bug - the "referee" was the diseased measurement. Re-measured with a truly independent computation (xmat.T @ qvel, world trajectory), the original diagnosis was confirmed: saw -0.5 gave vy +0.058/-0.130 (76%/106%), consistent with the net-displacement values all along (yaw pollution was real but only 10-21 deg, nowhere near reversal-sized).
Change
Lesson written twice, verbatim, as a hard rule: when questioning a measurement, the referee must be an independent algorithm (different code path, different physical derivation), e.g. rotate qvel by the body matrix directly, or inspect the raw world trajectory.
Outcome
With the independent referee in place the frame bug was confirmed, fixed, and the whole C4 line re-scored - revealing sidewalk had been working (see body-frame-velocity-api-audit).
Mechanism
Metrics sharing a code path (or an upstream API) share failure modes; agreement between them is evidence about the code, not the world. Only a measurement with an independent derivation can break the tie, because its errors are uncorrelated with the suspect's.
Applies when
- two metrics of the same quantity disagree
- about to retract a conclusion based on a second readout
- auditing evaluation code after a surprising result
“我用一个坏指标去质疑一个好指标,并把撤回写进了执行单。教训(写死):质疑一个测量时,不能用同一份代码里的另一个测量当裁判 —— 它们可能同源同病。裁判必须是独立算法(这次的裁判应该一开始就是 xmat.T · qvel[:3],或直接看世界轨迹)。”