让视觉语言动作模型在潜空间推理,实现高效实时机器人控制。
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
- 将多模态思维链融入连续潜变量,统一推理与预测过程。
- 推理延迟降低90%,在仿真与真实机器人任务中表现优于现有方法。
- 适合需要低延迟、高精度的实时具身智能系统应用。
视觉-语言-动作(VLA)模型受益于思维链(CoT)推理,但现有方法存在推理开销大、依赖离散推理表示的问题,与连续感知和控制不匹配。本文提出潜空间推理VLA(LaRA-VLA),将多模态思维链内化为连续潜表示,实现具身动作的统一推理与预测。该框架在推理时无需显式生成思维链,从而实现高效、面向动作的控制。为实现潜空间具身推理,我们设计了一种基于课程的学习范式,逐步从显式的文本与视觉思维链监督过渡到潜空间推理,并最终将潜空间推理动态适配至动作生成。我们构建了两个结构化的思维链数据集,并在仿真基准和长时程真实机器人操作任务上评估了LaRA-VLA。实验结果表明,该方法持续优于当前最优的VLA模型,且相比显式思维链方法推理延迟最高降低90%,验证了潜空间推理在实时具身控制中的有效性与高效性。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models benefit from chain-of-thought (CoT) reasoning, but existing approaches incur high inference overhead and rely on discrete reasoning representations that mismatch continuous perception and control. We propose Latent Reasoning VLA (LaRA-VLA), a unified VLA framework that internalizes multi-modal CoT reasoning into continuous latent representations for embodied action. LaRA-VLA performs unified reasoning and prediction in latent space, eliminating explicit CoT generation at inference time and enabling efficient, action-oriented control. To realize latent embodied reasoning, we introduce a curriculum-based training paradigm that progressively transitions from explicit textual and visual CoT supervision to latent reasoning, and finally adapts latent reasoning dynamics to condition action generation. We construct two structured CoT datasets and evaluate LaRA-VLA on both simulation benchmarks and long-horizon real-robot manipulation tasks. Experimental results show that LaRA-VLA consistently outperforms state-of-the-art VLA methods while reducing inference latency by up to 90\% compared to explicit CoT-based approaches, demonstrating latent reasoning as an effective and efficient paradigm for real-time embodied control. Project Page: https://loveju1y.github.io/Latent-Reasoning-VLA/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。