arXiv:2607.09818cs.ROcs.AI2026-07中稿 · the 2026 IEEE/RSJ …

用时空掩码提升视觉语言动作模型的长程规划能力

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging

论文配图:TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging
图 1 · 摘自论文原文
  • 引入2D时空掩码,显式建模动作序列的时空结构
  • 在LIBERO上仅用0.5B参数实现95.7%成功率,超越更大模型
  • 适合需要复杂长程规划的机器人操作任务

视觉-语言-动作(VLA)模型旨在理解自然语言指令与视觉观测,并生成并执行相应动作作为具身智能体。近期基于自回归标记的动作生成范式虽推动了多个代表性VLA模型的发展,但常将动作生成简化为下一个标记预测,缺乏对动作序列时空结构的显式建模,也未有效解耦视觉-语言表示与动作,限制了其在长时序和复杂场景下的表现。本文提出TS-Mask VLA,一种用于机器人操作的视觉-语言-动作框架。该框架基于两项关键设计:(1) 基于离散扩散的动作专家,配备桥接注意力机制,实现多层从视觉语言模型的条件传递,提升动作生成的准确性和稳定性;(2) 针对离散动作标记的时空二维掩码策略,强化模型对跨时间依赖性和跨维度耦合的理解,生成更结构一致的动作序列。我们在仿真基准和真实任务上进行了大量实验。在LIBERO上,TS-Mask VLA仅使用0.5B参数即达到95.7%的平均成功率,显著优于更大模型;在CALVIN上,其平均序列长度达4.19,展现出优异的长时序性能。全面分析与消融实验进一步验证了设计的有效性。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models aim to understand natural-language instructions and visual observations, and to generate and execute corresponding actions as embodied agents. Recently, autoregressive token-based action generation has driven the development of many representative VLA models. However, this paradigm often reduces action generation to next-token prediction, thereby lacking explicit modeling of the spatiotemporal structure of action sequences and the disentanglement between vision-language representations and actions, which can limit performance in long-horizon and complex scenarios. In this paper, we propose TS-Mask VLA, a vision-language-action framework for robot manipulation. TS-Mask VLA is built upon two key designs: (1) a Discrete Diffusion Action Expert equipped with a Bridge Attention conditioning bridge, which enables multi-layer conditioning from the VLM and facilitates more accurate and stable action generation; and (2) a temporal-spatial 2D masking strategy for discrete action tokens that strengthens the model's understanding of cross-time dependencies and inter-dimensional coupling, leading to more structurally consistent action sequences. We conduct extensive experiments on simulation benchmarks and real-world tasks. On LIBERO, TS-Mask VLA achieves a 95.7 percent average success rate with only 0.5B parameters, outperforming significantly larger models. On CALVIN, it attains the best average sequence length of 4.19 and strong long-horizon performance. Comprehensive analyses and ablations further validate the effectiveness of our design.

机器人操作视觉语言动作时空建模扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。