让视觉语言动作模型学会看时间,提升长序列操作能力。
Temporal Forcing: 4D Representation Alignment for Vision-Language-Action Models

- 引入历史路径,将观测历史编码为时序感知的潜在表示。
- 在LIBERO上达98.8%成功率,物理隐藏任务成功率达43.3%。
- 适合需要理解动态环境的机器人长程操作场景。
近期视觉-语言-动作(VLA)方法通过对齐3D场景几何结构提升了操作性能。然而,这些方法在长序列操作和视觉相似状态间的观察混淆问题上表现不佳,根源在于缺乏时间信息:3D几何仅反映当前状态,而非其随时间演变的过程。为此,本文提出时空强制(Temporal Forcing),一种面向VLA模型的4D表示对齐方法。具体而言,首先引入历史路径,使基础VLA模型能将观测历史归纳为时序感知的潜在表示;随后,该表示与预训练4D基础模型提取的几何特征对齐,后者通过时序一致的几何表示捕捉动态3D世界,从而实现对动态环境的深层理解。实验显示,该方法在LIBERO上达到98.8%成功率,优于基线2.2个百分点;在物理隐藏放置任务中,全任务成功率从20.0%提升至43.3%。代码将公开。
原文摘要 · Abstract (English)
Recent vision-language-action (VLA) methods improve manipulation performance by aligning their representations with 3D scene geometry. However, these methods often struggle with long-horizon manipulation and observation aliasing between visually similar states due to a lack of temporal information: the 3D scene geometry captures only the current state, rather than how it has evolved over time. To resolve this, we present Temporal Forcing, a 4D representation alignment method for VLA models. Specifically, we first introduce a history pathway that enables a vanilla VLA model to summarize observation history into temporally aware latent representations. Then, the latent representations are aligned with the geometric features extracted by a pretrained 4D foundation model, which captures the evolving 3D world through temporally consistent geometric representations, enabling a deeper understanding of dynamic environments. Temporal Forcing reaches 98.8% on LIBERO, outperforming its base model by 2.2 points. On a physical hidden-placement task, it raises full-task success from 20.0% to 43.3%. Code will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。