arXiv:2609.07096cs.RO2026-09

用预测模型让机器人在感知不全时仍能稳定走路

RoboDreamer: Anticipatory Humanoid Locomotion with Predictive State-Space Models

论文配图:RoboDreamer: Anticipatory Humanoid Locomotion with Predictive State-Space Models
图 1 · 摘自论文原文
  • 教师先学干净数据,学生通过遮蔽近期观测学习历史推断
  • 在真实机器人上实现遮蔽下稳定行走,延迟低于10毫秒
  • 适合做高鲁棒性机器人控制的开发者参考

人类形态机器人行走需要在感知不完整的情况下保持控制策略稳定,并利用时间上下文实现一致运动。我们提出RoboDreamer,一种两阶段师生框架,结合下一观察一致性与随机连续时间遮蔽。教师在干净观测上训练,学生则在遮蔽近期观测条件下进行知识蒸馏,促使策略从历史中推断缺失的当前信息。推理时复用相同遮蔽接口,实现隐式闭环动作优化和可选多步动作分块。采用Mamba作为时间骨干网络,对比实验表明遮蔽与蒸馏贡献了主要性能提升,而Mamba进一步提升了跟踪精度并满足实时性要求(延迟<10毫秒)。在IsaacLab、MuJoCo及真实Unitree G1机器人上验证了其在观测遮蔽下的鲁棒运动追踪能力与成功部署效果。

原文摘要 · Abstract (English)

Humanoid locomotion requires control policies that remain stable under imperfect sensing while exploiting temporal context for consistent motion. We present RoboDreamer, a two-stage teacher--student framework that combines next-observation consistency with randomized continuous temporal masking. A teacher is first trained on clean observations, and a student is then distilled under masked recent observations, encouraging the policy to infer missing current information from history. At inference, the same masking interface is reused for implicit closed-loop action refinement and optional multi-step action chunking. Mamba is used as the temporal backbone, while matched ablations show that masking/distillation provides a substantial part of the gain and Mamba contributes additional tracking improvements with real-time latency. Experiments in IsaacLab, MuJoCo, and on a Unitree G1 demonstrate robust motion tracking under observation masking and successful real-world deployment.

机器人控制状态预测Mamba实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。