arXiv:2603.01549cs.CVcs.AI2026-03被引 7

让视觉语言动作模型学会预测物体动态,提升真实世界操作能力

Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation

  • 用4D点轨迹预测辅助训练,让模型隐式学习物理互动规律
  • 在LIBERO-Long任务上提升10%,RoboCasa上提升40%
  • 无需修改推理架构,零额外计算开销,适配主流视觉语言动作模型

人类不仅学习自身运动方式,还理解周围世界对行为的响应。相比之下,当前视觉语言动作(VLA)模型虽具备出色语义理解能力,却常忽视物理交互中的时空动态。本文提出Pri4R,一种简单高效的方法,在训练中利用特权4D信息,赋予VLA模型对世界动态的隐式理解。具体而言,Pri4R为VLA添加轻量级点轨迹头,联合预测3D点轨迹。通过将VLA特征注入该头,模型在共享表示空间中学习演化场景几何,从而获得更符合物理规律的上下文以实现精准控制。由于架构简洁,Pri4R可无缝集成至主流VLA设计,仅需微小改动。推理时保持原有VLA结构不变,不增加输入、输出或计算开销。在仿真与真实世界评估中,Pri4R显著提升复杂操作任务表现,在LIBERO-Long上提升10%,在RoboCasa上提升40%。我们进一步验证3D点轨迹预测是学习动作-世界动态的有效监督信号,并通过大量消融实验验证设计合理性。

原文摘要 · Abstract (English)

Humans learn not only how their bodies move, but also how the surrounding world responds to their actions. In contrast, while recent Vision-Language-Action (VLA) models exhibit impressive semantic understanding, they often fail to capture the spatiotemporal dynamics governing physical interaction. In this paper, we introduce Pri4R, a simple yet effective approach that endows VLA models with an implicit understanding of world dynamics by leveraging privileged 4D information during training. Specifically, Pri4R augments VLAs with a lightweight point track head that predicts 3D point tracks. By injecting VLA features into this head to jointly predict future 3D trajectories, the model learns to incorporate evolving scene geometry within its shared representation space, enabling more physically aware context for precise control. Due to its architectural simplicity, Pri4R is compatible with dominant VLA design patterns with minimal changes. During inference, we run the model using the original VLA architecture unchanged; Pri4R adds no extra inputs, outputs, or computational overhead. Across simulation and real-world evaluations, Pri4R significantly improves performance on challenging manipulation tasks, including a +10% gain on LIBERO-Long and a +40% gain on RoboCasa. We further show that 3D point track prediction is an effective supervision target for learning action-world dynamics, and validate our design choices through extensive ablations. Project page: https://jiiiisoo.github.io/Pri4R/

视觉语言动作物理建模点轨迹强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。