arXiv:2503.07135cs.ROcs.CV2025-03CVPR被引 70

用真人视频学3D动作,让机器人零样本操作新场景。

VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

  • 从真实世界2D视频中重建3D手部轨迹和交互空间。
  • 在13项任务中零样本表现超越现有方法,支持跨机器人部署。
  • 适合想低成本实现通用机器人的研究者与工程师。

未来机器人需具备执行多种家庭任务的灵活性。如何在减少物理训练的前提下跨越具身差距,仍是核心挑战。本文提出VidBot框架,利用互联网上海量的真实人类视频数据,实现零样本机器人操作。该方法仅使用单目RGB视频,结合深度基础模型与运动结构技术,重建时序一致、度量尺度的3D可操作性表征,不依赖具体机器人形态。通过粗到精的分层学习:先在像素空间识别粗粒度动作,再以扩散模型生成精细交互轨迹,受测试时约束引导,实现上下文感知规划。大量实验表明,VidBot在13个零样本操作任务中显著优于基线方法,并可无缝部署于多种真实机器人系统,为利用日常人类视频提升机器人学习可扩展性开辟新路径。

原文摘要 · Abstract (English)

Future robots are envisioned as versatile systems capable of performing a variety of household tasks. The big question remains, how can we bridge the embodiment gap while minimizing physical robot learning, which fundamentally does not scale well. We argue that learning from in-the-wild human videos offers a promising solution for robotic manipulation tasks, as vast amounts of relevant data already exist on the internet. In this work, we present VidBot, a framework enabling zero-shot robotic manipulation using learned 3D affordance from in-the-wild monocular RGB-only human videos. VidBot leverages a pipeline to extract explicit representations from them, namely 3D hand trajectories from videos, combining a depth foundation model with structure-from-motion techniques to reconstruct temporally consistent, metric-scale 3D affordance representations agnostic to embodiments. We introduce a coarse-to-fine affordance learning model that first identifies coarse actions from the pixel space and then generates fine-grained interaction trajectories with a diffusion model, conditioned on coarse actions and guided by test-time constraints for context-aware interaction planning, enabling substantial generalization to novel scenes and embodiments. Extensive experiments demonstrate the efficacy of VidBot, which significantly outperforms counterparts across 13 manipulation tasks in zero-shot settings and can be seamlessly deployed across robot systems in real-world environments. VidBot paves the way for leveraging everyday human videos to make robot learning more scalable.

机器人操作零样本学习3D动作理解视频驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。