提出连续推理机制,让视觉语言动作模型共享可验证的内部语言以提升控制能力。
Continuous Reasoning for Vision-Language-Action

- 用共享高斯隐变量表示连续思考,实现跨模型复用与验证
- 在TX-G2上子任务成功率提升40.4%,HSR上提升26.3%
- 强调推理需可验证、可共享,避免模型私有捷径
自然语言是语言和视觉语言模型的强大推理媒介,但与连续控制的粒度不匹配。文本和显式子目标作用于任务级粒度,而视觉语言动作(VLA)策略需在更细的时间尺度上选择动作;单次推理步骤可能跨越多个动作块,却仅与当前动作弱耦合。这引发一个关键问题:什么能充当VLA的推理介质?我们主张,有效的推理介质必须可在模型间共享、可通过下游动作改进验证,并与时间扩展的控制结构对齐。基于此,我们提出连续推理(Continuous Reasoning for Vision-Language-Action)。模型首先预测以结构化连续思想形式的连续推理,再将其作为共享上下文用于分块动作生成。仅更好的动作预测不足以证明良好推理:若该内部介质无法在不同模型实例间共享并独立通过改进后的下游控制验证,那么新增的隐变量可能仅成为帮助特定行为的模型私有捷径。因此,我们将连续推理实例化为共享的高斯隐接口,并采用指数移动平均教师模型进行自验证训练,要求教师成功利用学生推理来预测目标动作。实证表明,连续推理提升了LIBERO-PRO的鲁棒性,在兼容AgiBot G2的TX-G2上子任务成功率提升40.4%,在HSR上提升26.3%。这表明,VLA中的推理更关乎一种可共享、可验证的内部动作语言,而非额外的标记。
原文摘要 · Abstract (English)
Natural language is a powerful reasoning medium for language and vision-language models, but it is mismatched to the granularity of continuous control. Text and explicit subgoals operate at task-level granularity, whereas vision-language-action (VLA) policies must choose actions at a much finer temporal scale; a single reasoning step can therefore span many action chunks while remaining only weakly coupled to the action needed now. This suggests a different question for VLA: what should play the role of language? We argue that a useful VLA reasoning medium must be shareable across model instances, verifiable through downstream action improvement, and aligned with temporally extended control structure. Based on this view, we propose Continuous Reasoning for Vision-Language-Action. Our model first predicts continuous reasoning in the form of a structured set of continuous thoughts, then reuses them as shared context for chunk-structured action generation. Better action prediction alone does not certify good reasoning: if the same internal medium cannot be shared across model instances and independently verified through improved downstream control, the added latent may simply become a model-private shortcut that helps on seen behaviors without supporting generalizable control. We therefore instantiate continuous reasoning as a shared Gaussian latent interface and train it with a self-verification objective in which an exponential-moving-average teacher must successfully consume the student's reasoning when predicting target actions. Empirically, Continuous Reasoning improves LIBERO-PRO robustness and performs strongly on real robots, raising mean subtask success over π0.5 by 40.4% on TX-G2, an AgiBot G2-compatible variant, and 26.3% on HSR. This suggests that reasoning in VLA is less about extra tokens than about a shareable, verifiable internal language for action.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。