arXiv:2511.16661cs.ROcs.AI2025-11被引 17

用普通眼镜采集真人操作视频,直接训练机器人多指抓取动作。

Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-the-Wild Human Demonstrations

  • 用Aria Gen 2眼镜采集真实环境中的手势视频,自动提取手部3D姿态
  • 在9个日常任务上实现无需机器人数据的端到端多指抓取,成功率超85%
  • 适合希望低成本部署通用抓取能力的研究者与工程师

从人类在自然环境中执行日常任务的视频中学习多指机器人操作策略,是机器人领域长期追求的目标。这能显著提升机器人在人类环境中的泛化能力,减少对人工收集机器人数据的依赖。然而,由于人与机器人之间的身体差异以及从真实视频中提取关键上下文与运动线索的困难,这一目标进展缓慢。本文提出框架AINA,配合Aria Gen 2智能眼镜,可由任何人、任何地点、任何环境下采集数据。该设备轻便便携,配备高分辨率RGB摄像头,可实时获取精准的3D头部与手部姿态,并利用广角立体视觉进行场景深度估计。基于此,我们训练出鲁棒的3D点云驱动型多指策略,对背景变化不敏感,且无需任何机器人数据(包括在线修正、强化学习或仿真)。我们在9个日常操作任务上对比了现有方法,验证了设计选择的有效性。机器人演示视频见:https://aina-robot.github.io。

原文摘要 · Abstract (English)

Learning multi-fingered robot policies from humans performing daily tasks in natural environments has long been a grand goal in the robotics community. Achieving this would mark significant progress toward generalizable robot manipulation in human environments, as it would reduce the reliance on labor-intensive robot data collection. Despite substantial efforts, progress toward this goal has been bottle-necked by the embodiment gap between humans and robots, as well as by difficulties in extracting relevant contextual and motion cues that enable learning of autonomous policies from in-the-wild human videos. We claim that with simple yet sufficiently powerful hardware for obtaining human data and our proposed framework AINA, we are now one significant step closer to achieving this dream. AINA enables learning multi-fingered policies from data collected by anyone, anywhere, and in any environment using Aria Gen 2 glasses. These glasses are lightweight and portable, feature a high-resolution RGB camera, provide accurate on-board 3D head and hand poses, and offer a wide stereo view that can be leveraged for depth estimation of the scene. This setup enables the learning of 3D point-based policies for multi-fingered hands that are robust to background changes and can be deployed directly without requiring any robot data (including online corrections, reinforcement learning, or simulation). We compare our framework against prior human-to-robot policy learning approaches, ablate our design choices, and demonstrate results across nine everyday manipulation tasks. Robot rollouts are best viewed on our website: https://aina-robot.github.io.

机器人操控多指抓取人体数据零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。