arXiv:2511.21690cs.ROcs.CV2025-11被引 12

用3D轨迹空间建模世界,让机器人从不同视角视频中快速学新任务。

TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos

  • 构建3D轨迹空间抽象表示,剥离外观保留几何结构。
  • 仅需5个目标视频即达80%成功率,推理速度提升50-600倍。
  • 无需物体检测器,可跨人形与机器人直接迁移学习。

在新平台和新场景中仅凭少量示范学习新机器人任务仍具挑战性。尽管人类和其他机器人的视频资源丰富,但本体、相机和环境差异阻碍其直接使用。我们提出统一的符号化表示——紧凑的3D“轨迹空间”,实现跨本体、跨环境、跨任务视频的学习。提出TraceGen世界模型,预测轨迹空间中的未来运动,而非像素空间,从而抽象外观并保留操作所需的几何结构。为大规模训练TraceGen,开发TraceForge数据流水线,将异构的人类与机器人视频转换为一致的3D轨迹,生成包含12.3万段视频和180万个观察-轨迹-语言三元组的语料库。在该语料库上预训练后,得到可迁移的3D运动先验,仅需5个目标机器人视频即可在4项任务中达到80%成功率,且推理速度比当前最优视频基世界模型快50至600倍。在更困难情形下,仅提供5段手持手机拍摄的未校准人类示范视频时,仍可在真实机器人上实现67.5%的成功率,体现TraceGen无需依赖物体检测器或高维像素生成即可跨本体适应的能力。

原文摘要 · Abstract (English)

Learning new robot tasks on new platforms and in new scenes from only a handful of demonstrations remains challenging. While videos of other embodiments - humans and different robots - are abundant, differences in embodiment, camera, and environment hinder their direct use. We address the small-data problem by introducing a unifying, symbolic representation - a compact 3D "trace-space" of scene-level trajectories - that enables learning from cross-embodiment, cross-environment, and cross-task videos. We present TraceGen, a world model that predicts future motion in trace-space rather than pixel space, abstracting away appearance while retaining the geometric structure needed for manipulation. To train TraceGen at scale, we develop TraceForge, a data pipeline that transforms heterogeneous human and robot videos into consistent 3D traces, yielding a corpus of 123K videos and 1.8M observation-trace-language triplets. Pretraining on this corpus produces a transferable 3D motion prior that adapts efficiently: with just five target robot videos, TraceGen attains 80% success across four tasks while offering 50-600x faster inference than state-of-the-art video-based world models. In the more challenging case where only five uncalibrated human demonstration videos captured on a handheld phone are available, it still reaches 67.5% success on a real robot, highlighting TraceGen's ability to adapt across embodiments without relying on object detectors or heavy pixel-space generation.

世界模型跨本体轨迹空间小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。