Skip to content
MouseMouse

Anchoring the action on the measured joint angle (target = q + beta*a), with a beta curriculum down to tau_limit/kp, bounded torque by construction, removed the re-falls and later stood the robot up on hardware

A lesson the Coach cites as beta-anchored-action-target.
Mechanism understoodTraining runs

For large-motion skills on position-controlled actuators, bound the action relative to the measured joint angle with a per-joint authority of tau_limit/kp, curriculum the authority down from full range, keep the curriculum state out of the observation and pin acceptance at the deployed authority - and make the deployment code refuse to run the anchored contract without a measured q.

Symptom

The V0 full-range absolute action produced violent targets; the V1 command-anchored rate limit made standing oscillate. Both failure modes came from how the action becomes a target.

Context

V2.0 (user approved, from scratch): BetaAnchorJointPositionAction, target = q_measured + beta_j(m)*a, memoryless per step. beta_j(m) = floor + m*(beta0 - floor), beta0 = the contract half-range (m = 1 reproduces V0 authority), floor = min(tau_limit/kp, beta0): hip_pitch 1.309 -> 0.40, knee 1.047 -> 0.40, hip_yaw -> 0.917, the other joints unchanged - the tightening lands exactly on the joints the V0 torque account convicted. m drops 0.1 per step when a standing-share EMA exceeds 0.35. beta is NOT in the observation, so the 45-dim contract is untouched; acceptance is pinned at m = 0 because the Python curriculum state is not saved in the checkpoint. The deployment chain got a new profile (recovery_v2: action_anchor current_q, explicit per-joint beta written into the contract, independent of the gain profile), and policy_io raises if q is missing rather than silently falling back to the absolute contract; the old profile's check reproduced its pre-change deviation bit for bit.

Change

New action term and beta curriculum; later the RS06 floor was lowered 0.40 -> 0.30 -> 0.25 (kp*beta 7.5 N*m) and the stamped deployment profile was synced to 0.25.

Outcome

First acceptance at m = 0 (v2_0b): re-falls 0% in every category, the torque gate passed for the first time on the line (worst 69.9%), knee jitter 0.004; supine 98.8 / side 88.8% with prone and mid still failing (fixed by the conditional pull curriculum). MuJoCo showed demand at or under the limits (hip_pitch 11.7/12 against V0's 26.8). Lowering beta cut impact (hip_pitch demand 9.7 -> 8.5 N*m) but barely slowed the get-up - it had become coordination-limited. Enabling the policy moves the target only +/-beta around the current pose, so there is no homing fling; the 08-11 real get-up and the later v3_1p1c both run on this contract.

Mechanism

kp*beta caps the proportional torque in a single step with no build-up delay and no memory, giving both a hard impact bound and full balance bandwidth.

Applies when

  • a skill needs full joint range but hardware torque limits are low
  • absolute position targets cause impacts or saturation
  • changing the action semantics of a contract that deployed policies share
“**动作项** `BetaAnchorJointPositionAction`:`target = q_实测 + β_j(m)·a`, 逐步无记忆 … **Play/验收钉 m=0(= floor = 部署档)**:python 课程状态不进 checkpoint, Play cfg 显式 `beta_m_start=0` … 判读:**结构赌注兑现** —— 站姿零再摔 + 力矩账首过(kp·β 封顶按构造)”
git:Lucen-recovery@origin/recovery:train/RECOVERY_V0_SPEC.md § §33 V2.0 预注册(2026-08-10,用户点头开工):β 锚定动作空间,从零训

Principles that cite it