用智能眼镜采集人类示范,实现机器人零样本视觉操控。
ActiveGlasses: Learning Manipulation with Active Vision from Ego-centric Human Demonstration
- 戴眼镜采集人类第一视角演示,用同一摄像头做训练和部署。
- 在遮挡和精准操作任务中零样本迁移表现优于基线。
- 适合希望低成本获取自然人机交互数据的机器人研究者。
大规模真实世界机器人数据采集是机器人日常应用的前提。现有方法通常依赖专用手持设备弥合具身差距,不仅增加操作负担、限制可扩展性,也难以捕捉人类日常互动中的自然感知-操作协调行为。为此,我们提出ActiveGlasses系统,通过第一视角人类示范学习机器人操控。佩戴在智能眼镜上的双目相机作为唯一感知设备,用于数据采集与策略推理:操作员佩戴它进行徒手示范,部署时将同一相机安装于6自由度感知机械臂上复现人类主动视觉。为实现零样本迁移,从示范中提取物体轨迹,并采用以物体为中心的点云策略联合预测操控动作与头部运动。在涉及遮挡和精确交互的多个挑战性任务中,ActiveGlasses实现了带主动视觉的零样本迁移,相同硬件条件下持续超越强基线,并在两个机器人平台上实现泛化。
原文摘要 · Abstract (English)
Large-scale real-world robot data collection is a prerequisite for bringing robots into everyday deployment. However, existing pipelines often rely on specialized handheld devices to bridge the embodiment gap, which not only increases operator burden and limits scalability, but also makes it difficult to capture the naturally coordinated perception-manipulation behaviors of human daily interaction. This challenge calls for a more natural system that can faithfully capture human manipulation and perception behaviors while enabling zero-shot transfer to robotic platforms. We introduce ActiveGlasses, a system for learning robot manipulation from ego-centric human demonstrations with active vision. A stereo camera mounted on smart glasses serves as the sole perception device for both data collection and policy inference: the operator wears it during bare-hand demonstrations, and the same camera is mounted on a 6-DoF perception arm during deployment to reproduce human active vision. To enable zero-transfer, we extract object trajectories from demonstrations and use an object-centric point-cloud policy to jointly predict manipulation and head movement. Across several challenging tasks involving occlusion and precise interaction, ActiveGlasses achieves zero-shot transfer with active vision, consistently outperforms strong baselines under the same hardware setup, and generalizes across two robot platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。