用头戴设备和摄像头就能让机器人学人类动作,省时又安全。
ARMimic: Learning Robotic Manipulation from Passive Human Demonstrations in Augmented Reality
- 仅用消费级头显和摄像头,通过增强现实实现无硬件负担的演示采集。
- 演示时间减少50%,长任务成功率比现有方法提升11%。
- 适合在真实复杂环境中快速部署,尤其适合非专业用户使用。
模仿学习是机器人技能获取的强大范式,但传统示范方法如身体引导和遥操作存在设备繁重、流程干扰等问题。近期基于扩展现实(XR)头显的被动观察显示了前景,但现有方法仍需额外硬件、复杂校准或受限录制条件,制约可扩展性。本文提出ARMimic,一种轻量级、低硬件依赖的框架,仅需消费级XR头显与固定工作区摄像头即可实现无需机器人的规模化数据采集。ARMimic融合自指手部追踪、增强现实机器人叠加与实时深度感知,确保动作无碰撞且运动学可行。其核心是统一的模仿学习流水线,将人与虚拟机器人轨迹视为可互换,支持跨实体与环境泛化。我们在两项操作任务上验证该方法,包括挑战性的长时程碗叠任务。实验表明,相比遥操作,示范时间减少50%;相较于基于遥操作数据训练的ACT基线,任务成功率提升11%。结果表明,ARMimic实现了安全、无缝、真实场景下的数据采集,为多样化现实环境中机器人学习提供了巨大潜力。
原文摘要 · Abstract (English)
Imitation learning is a powerful paradigm for robot skill acquisition, yet conventional demonstration methods--such as kinesthetic teaching and teleoperation--are cumbersome, hardware-heavy, and disruptive to workflows. Recently, passive observation using extended reality (XR) headsets has shown promise for egocentric demonstration collection, yet current approaches require additional hardware, complex calibration, or constrained recording conditions that limit scalability and usability. We present ARMimic, a novel framework that overcomes these limitations with a lightweight and hardware-minimal setup for scalable, robot-free data collection using only a consumer XR headset and a stationary workplace camera. ARMimic integrates egocentric hand tracking, augmented reality (AR) robot overlays, and real-time depth sensing to ensure collision-aware, kinematically feasible demonstrations. A unified imitation learning pipeline is at the core of our method, treating both human and virtual robot trajectories as interchangeable, which enables policies that generalize across different embodiments and environments. We validate ARMimic on two manipulation tasks, including challenging long-horizon bowl stacking. In our experiments, ARMimic reduces demonstration time by 50% compared to teleoperation and improves task success by 11% over ACT, a state-of-the-art baseline trained on teleoperated data. Our results demonstrate that ARMimic enables safe, seamless, and in-the-wild data collection, offering great potential for scalable robot learning in diverse real-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。