用人类视频教机器人操作,无需真实机器人数据即可直接执行。
BridgeACT: Bridging Human Demonstrations to Robot Actions via Unified Tool-Target Affordances

- 将人类示范转化为无身体依赖的可操作性表征,打通视觉到动作的桥梁。
- 在真实场景中对未见物体、视角和环境均实现成功操作,成功率显著提升。
- 适合希望零样本迁移机器人技能的研究者或开发者使用。
从人类视频学习机器人操作具有规模和多样性优势,但将示范转化为可执行机器人行为仍具挑战。以往方法依赖机器人数据进行下游适配,或仅学习停留在感知层面的可操作性表征,无法直接支持现实执行。本文提出BridgeACT,一种完全基于人类视频的可操作性驱动框架,无需任何机器人演示数据即可直接学习机器人操作。核心思想是将可操作性建模为与身体无关的中间表示,连接人类示范与机器人动作。BridgeACT将操作分解为抓取位置选择与运动路径规划两个互补问题:首先在当前场景中定位任务相关的可操作区域,再从人类示范中预测任务条件下的3D运动可操作性。最终通过抓取模块与轻量级闭环运动控制器将可操作性映射为机器人动作,实现在真实机器人上的直接部署。此外,复杂操作任务被表示为可操作性操作的组合,实现对多样化任务与物物交互的统一处理。在真实世界操作任务上的实验表明,BridgeACT优于现有基线,且能泛化至未见过的物体、场景和视角。
原文摘要 · Abstract (English)
Learning robot manipulation from human videos is appealing due to the scale and diversity of human demonstrations, but transferring such demonstrations to executable robot behavior remains challenging. Prior work either relies on robot data for downstream adaptation or learns affordance representations that remain at the perception level and do not directly support real-world execution. We present BridgeACT, an affordance-driven framework that learns robotic manipulation directly from human videos without requiring any robot demonstration data. Our key idea is to model affordance as an embodiment-agnostic intermediate representation that bridges human demonstrations and robot actions. BridgeACT decomposes manipulation into two complementary problems: where to grasp and how to move. To this end, BridgeACT first grounds task-relevant affordance regions in the current scene, and then predicts task-conditioned 3D motion affordances from human demonstrations. The resulting affordances are mapped to robot actions through a grasping module and a lightweight closed-loop motion controller, enabling direct deployment on real robots. In addition, we represent complex manipulation tasks as compositions of affordance operations, which allows a unified treatment of diverse tasks and object-to-object interactions. Experiments on real-world manipulation tasks show that BridgeACT outperforms prior baselines and generalizes to unseen objects, scenes, and viewpoints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。