arXiv:2601.05248cs.RO2026-01被引 28

用隐式时空推理提升机器人视觉-语言-动作模型的响应速度与物理理解能力

LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model

  • 构建潜空间时序思维链,融合视觉、3D结构和机器人本体感知信息
  • 在真实场景中实现13%-14%的成功率提升,且推理延迟更低
  • 适合需要快速响应的机器人抓取、移动和灵巧操作任务

视觉-语言-动作(VLA)模型虽具强泛化能力,但显式语言推理常导致显著延迟,且受限于语言表征,难以捕捉难以言喻的物理特性。为此,我们提出LaST₀,通过潜空间时序思维链(CoT)实现高效推理,建模未来视觉动态、3D结构信息与机器人本体状态,并跨时间保持一致性。采用混合变换器双系统架构:低频推理专家负责潜空间推断,高频执行专家基于机器人导向表示生成动作。训练时设定异步操作频率,支持部署中自适应切换。在10项真实世界任务(桌面、移动及灵巧手操作)中,相比现有最先进方法,成功率分别提升13%、14%和14%。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have recently shown strong generalization, with some approaches seeking to explicitly generate linguistic reasoning traces or predict future observations prior to execution. However, explicit reasoning typically incurs non-negligible inference latency, which constrains the temporal resolution required for robotic manipulation. Moreover, such reasoning is confined to the linguistic space, imposing a representational bottleneck that struggles to faithfully capture ineffable physical attributes. To mitigate these limitations, we propose LaST$_0$, a framework that enables efficient reasoning before acting through a Latent Spatio-Temporal Chain-of-Thought (CoT), capturing fine-grained physical and robotic dynamics that are often difficult to verbalize. Specifically, we introduce a token-efficient latent CoT space that models future visual dynamics, 3D structural information, and robot proprioceptive states, and further extends these representations across time to enable temporally consistent implicit reasoning trajectories. Furthermore, LaST$_0$ adopts a dual-system architecture implemented via a Mixture-of-Transformers design, where a reasoning expert conducts low-frequency latent inference and an acting expert generates high-frequency actions conditioned on robotics-oriented latent representations. To facilitate coordination, LaST$_0$ is trained with heterogeneous operation frequencies, enabling adaptive switching during deployment. Across 10 real-world tasks spanning tabletop, mobile, and dexterous hand manipulation, LaST$_0$ improves mean success rates by 13%, 14% and 14% over prior SOTA VLA methods, respectively.

机器人视觉-语言推理链动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。