arXiv:2511.17199cs.CV2025-11被引 8

让机器人操作更连贯,通过4维感知提升视觉语言动作模型的时空一致性。

VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation

  • 将时间信息嵌入3D位置,生成4维视觉表征,融合跨注意力机制。
  • 扩展动作表示包含时间维度,实现时空协同规划与预测。
  • 适用于需要流畅连续操作的复杂机器人任务,如精细抓取与移动。

视觉-语言-动作(VLA)模型在通用机器人任务中展现出潜力,但在时空一致的操作方面仍面临挑战,需细粒度表征。现有方法通常将3D位置嵌入视觉表征以提升空间精度,但难以实现动作执行的时间一致性。本文提出VLA-4D,一种具备4D感知能力的通用VLA模型,用于实现时空一致的机器人操作。核心设计包括:1)4D感知视觉表征——提取视觉特征,将1D时间嵌入3D位置形成4D嵌入,并通过交叉注意力机制融合为统一视觉表示;2)时空动作表征——在传统空间动作表示基础上加入时间信息,支持时空规划,并将多模态表示对齐至大语言模型以进行时空动作预测。该框架下,视觉与动作表征协同实现空间平滑、时间连贯的操作。此外,我们扩展了VLA数据集,添加了时间动作标注以微调模型。大量实验验证了该方法在多种机器人操作任务中的优越性。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models show potential for general robotic tasks, but remain challenging in spatiotemporally coherent manipulation, which requires fine-grained representations. Typically, existing methods embed 3D positions into visual representations to enhance the spatial precision of actions. However, these methods struggle to achieve temporally coherent control over action execution. In this work, we propose VLA-4D, a general VLA model with 4D awareness for spatiotemporally coherent robotic manipulation. Our model is guided by two key designs: 1) 4D-aware visual representation. We extract visual features, embed 1D time into 3D positions for 4D embeddings, and fuse them into a unified visual representation via a cross-attention mechanism. 2) Spatiotemporal action representation. We extend conventional spatial action representations with temporal information to enable the spatiotemporal planning, and align the multimodal representations into the LLM for spatiotemporal action prediction. Within this unified framework, the designed visual and action representations jointly make robotic manipulation spatially-smooth and temporally-coherent. In addition, we extend the VLA dataset with temporal action annotations for fine-tuning our model. Extensive experiments have been conducted to verify the superiority of our method across different tasks of robotic manipulation.

机器人操作时空一致性多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。