用3D轨迹建模物体交互,实现无需动作标签的可扩展机器人学习。
$μ_0$: A Scalable 3D Interaction-Trace World Model

- 通过预测关键点的3D轨迹,构建紧凑且与具体机器人无关的运动表示。
- 在2D和3D轨迹预测任务上优于基线模型,包括基于视觉语言模型的方法。
- 可冻结复用,适用于不同机器人平台,媲美带动作标签训练的先进模型。
能够捕捉动作如何引发物理变化的世界模型,使机器人学习摆脱对特定动作标签的依赖。像素空间视频模型虽具备广泛的视觉先验,但将大量模型容量用于密集外观重建;直接建模动作则需依赖具体机器人的动作标签,限制了可扩展性。我们提出 $μ_0$,一种基于3D轨迹的可扩展世界模型。不同于预测密集像素或直接建模动作,$μ_0$ 预测如物体、工具、手部及接触区域等显著交互点的平滑3D轨迹,形成紧凑且与身体形态无关的运动接口。为支持从多样化视频源训练,我们设计TraceExtract系统,自动提取3D监督信号:选取关键点,构建全局对齐轨迹,并将运动片段与分层语言描述关联。该监督信号通过预训练视觉-语言主干与模块化轨迹专家结合,以B样条控制点表示每个查询并预测未来轨迹。实验表明,$μ_0$ 在2D与3D轨迹预测任务上均超越基线,包括轨迹预测模型与分词化视觉语言模型方法。由于$μ_0$ 可冻结复用,可与动作专家配合用于下游机器人部署。尽管无动作标签预训练,其生成的轨迹条件策略性能仍可媲美经动作监督预训练的VLA模型(如 $π_0$)。这些结果确立3D轨迹作为跨体态操作中可扩展、可迁移的表征范式。
原文摘要 · Abstract (English)
World models that capture how actions induce physical change enable scalable robot learning without reliance on embodiment-specific action labels. Pixel-space video models provide broad visual priors but expend model capacity on dense appearance reconstruction, while direct action models require embodiment-specific labels that hinder scalability. We present $μ_0$, a scalable world model based on 3D traces. Rather than predicting dense pixels or directly modeling actions, $μ_0$ forecasts smooth 3D trajectories for salient interaction points such as objects, tools, hands, and contact regions, yielding a compact, embodiment-agnostic motion interface. To enable training from diverse video sources, our TraceExtract system automatically extracts 3D supervision by selecting keypoints, constructing globally aligned traces, and associating motion segments with hierarchical language captions. This TraceExtract supervision pretrains $μ_0$ by combining a pretrained vision-language backbone with a modular trace expert, which represents each query via B-spline control points and predicts future traces. Experiments show that $μ_0$ outperforms baselines in both 2D and 3D trace prediction, including trace prediction models and tokenized VLM methods. Because $μ_0$ is frozen and reusable, it can be paired with action experts for downstream robot embodiments. Despite action-free pretraining, the resulting trace-conditioned policies achieve performance competitive with VLA models pretrained with action supervision, such as $π_0$. These results establish 3D traces as a scalable and transferable representation for cross-embodiment manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。