用视觉推理提升视觉语言动作模型的准确率与实时性
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies

- 用视觉证据替代文本思维链,保持空间精度且无解码延迟
- 在BridgeData V2上将步骤延迟从8.377秒降至0.367秒,提速22.8倍
- 适合需要低延迟实时控制的机器人应用,如真实场景部署
近期工作开始为视觉语言动作(VLA)策略引入显式中间推理。但在具身控制中,文本思维链存在缺陷:无关或弱相关文本会干扰动作预测,自回归文本解码又带来过高延迟,难以支持实时闭环执行。本文提出VISUALTHINK-VLA,一种面向高精度、低延迟VLA策略的视觉中间推理框架。其核心思想是通过紧凑的视觉证据接口实现有效视觉思考,既保留空间精度,又避免解码开销。此外,采用定制化选择性路由机制学习视觉证据标记,实现低延迟推理的同时保持高容量专精能力。我们还构建了VisualEvidence-Kit资源,包含由VisualEvidence-Agent生成的754.7k条VLA指令视觉证据集,用于路径监督与反事实可信度测试。在多个基准和真实机器人评估中,VISUALTHINK-VLA在多数任务上取得最高成功率,同时将推理增强基线的多秒级延迟降至亚秒级。例如,在BridgeData V2上,步骤延迟从8.377秒降至0.367秒,提速22.8倍。
原文摘要 · Abstract (English)
Recent work has begun to equip vision-language-action (VLA) policies with explicit intermediate reasoning. In embodied control, however, textual chain-of-thought is a poor fit: irrelevant or weakly textual information can interfere with action prediction, while autoregressive text decoding adds too much latency for real-time closed-loop execution. We present VISUALTHINK-VLA, a visual intermediate-reasoning framework for accurate, low-latency VLA policies. Our bootstrapping philosophy is to guide action with effective visual thinking: VISUALTHINK-VLA bootstraps action prediction through a compact visual-evidence interface that preserves spatial precision while avoiding decoding overhead. Besides, to further improve performance and efficiency, VISUALTHINK-VLA adopts a tailored selective routing mechanism to learn the visual evidence tokens, enabling low-latency inference while preserving high-capacity specialization. We also introduce VisualEvidence-Kit, a supervision-and-audit resource centered on a VisualEvidence-Agent that constructs a 754.7k VLA instructions VisualEvidence-Set for route supervision and counterfactual faithfulness tests. Across multiple benchmarks and real-robot evaluation, VISUALTHINK-VLA achieves the highest success rate on most benchmarks while reducing the multi-second latency of reasoning-augmented baselines to the sub-second regime. For example, on BridgeData V2, it reduces step latency from 8.377,s with ECoT to 0.367,s, achieving a 22.8 times speedup.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。