arXiv:2606.11628cs.ROcs.AI2026-06被引 1

从网络视频学通用操作意图,让机器人零样本迁移到新场景。

LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition

论文配图:LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition
图 1 · 摘自论文原文
  • 两阶段框架:先从互联网视频学任务意图,再在仿真中训练机器人动作策略。
  • 仅用1小时手机视频即可学会推移和穿线任务,真实世界五项任务零样本迁移成功。
  • 同一意图模型适配灵巧手与夹爪,实现跨机械体的技能复用。

当前主流机器人学习方法依赖机器人示范或结构化人类数据,成本高且绑定特定机械结构。相比之下,非结构化人类视频具备可扩展性,包含多样化的操作演示,但无法直接映射到机器人动作。本文提出LUCID,一种两阶段框架:首先从大规模互联网视频中学习任务意图模型,该模型在闭环中根据当前观测预测短时序意图(场景下一步应发生什么);其次在大规模并行仿真中训练具身化传感器-运动策略,将意图转化为具体动作。意图接口对不同控制器共享,同一模型可应用于灵巧手、平行夹爪等多种机械体。在五个真实操作任务(搅拌、擦拭、分拣)上,仅通过互联网视频监督即实现零样本迁移至新场景与物体实例;推移(push-T)和穿缆(cable routing)任务则分别使用1小时自采集手机视频进行训练。项目页面:https://lucid-robot.github.io/

原文摘要 · Abstract (English)

The most widely-adopted robot learning pipelines today learn skills from robot demonstrations or structured human data, which are expensive to collect and tied to specific embodiments. In contrast, unstructured human videos provide a scalable alternative. They contain diverse manipulation demonstrations across objects, scenes, and strategies, but are not directly connected to robot action. We propose LUCID, a two-stage framework that learns task intent from unstructured human videos drawn from internet-scale datasets and learns robot control in massively-parallel simulation. The intent model predicts short-horizon intent (what should happen next in the scene) from the current observation in closed loop. An embodiment-specific sensorimotor policy converts this intent into robot actions. The intent interface is shared across controllers, so the same intent model can be applied to different embodiments, from our primary dexterous hand to a parallel-jaw gripper. We evaluate LUCID on five real-world manipulation tasks: stirring, wiping, and binning supervised by only internet video, with zero-shot transfer to novel scenes and object instances; and push-T and cable routing supervised by 1 hr each of self-collected smartphone video. Project page: https://lucid-robot.github.io/.

机器人技能视频学习零样本迁移多体兼容

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。