arXiv:2603.09574cs.ROcs.LG2026-03

仅用机载传感器实现人形机器人行走控制,无需外部状态估计。

SCDP: Learning Humanoid Locomotion from Partial Observations via Mixed-Observation Distillation

  • 通过混合观测训练,让扩散模型从传感器历史推断完整运动状态。
  • 仿真中速度控制成功率99-100%,AMASS测试集追踪成功率达93%。
  • 适合部署于真实机器人,支持50Hz实时运行,无外部感知依赖。

将人形机器人行走控制从离线数据集蒸馏为可部署策略仍具挑战,因现有方法依赖复杂且不可靠的全身体态信息。本文提出传感器条件扩散策略(SCDP),仅使用机载传感器即可实现行走控制,无需显式状态估计。SCDP通过混合观测训练解耦感知与监督:扩散模型基于传感器历史进行条件生成,同时被监督预测特权未来状态-动作轨迹,强制模型在部分可观测条件下推断运动动力学。我们进一步引入受限去噪、上下文分布对齐和上下文感知注意力掩码,以促进模型内部隐式状态估计,并防止训练-部署不匹配。在速度指令控制与运动参考跟踪任务上验证了SCDP的有效性。仿真结果显示,速度控制成功率高达99-100%,在AMASS测试集上追踪成功率为93%,性能接近使用特权信息的基线方法。最后,我们将训练好的策略部署于真实G1人形机器人,在50 Hz频率下实现鲁棒行走,无需外部传感或状态估计。

原文摘要 · Abstract (English)

Distilling humanoid locomotion control from offline datasets into deployable policies remains a challenge, as existing methods rely on privileged full-body states that require complex and often unreliable state estimation. We present Sensor-Conditioned Diffusion Policies (SCDP) that enables humanoid locomotion using only onboard sensors, eliminating the need for explicit state estimation. SCDP decouples sensing from supervision through mixed-observation training: diffusion model conditions on sensor histories while being supervised to predict privileged future state-action trajectories, enforcing the model to infer the motion dynamics under partial observability. We further develop restricted denoising, context distribution alignment, and context-aware attention masking to encourage implicit state estimation within the model and to prevent train-deploy mismatch. We validate SCDP on velocity-commanded locomotion and motion reference tracking tasks. In simulation, SCDP achieves near-perfect success on velocity control (99-100%) and 93% tracking success in AMASS test set, performing comparable to privileged baselines while using only onboard sensors. Finally, we deploy the trained policy on a real G1 humanoid at 50 Hz, demonstrating robust real robot locomotion without external sensing or state estimation.

人形机器人扩散模型强化学习传感器融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。