arXiv:2506.17417cs.LG2025-06被引 8

视觉语言模型推理时自验证能力有限,效果不如简单投票。

Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?

  • 用多数投票法比自验证策略更有效提升性能。
  • 强化学习微调的视觉语言模型无法稳定实现'顿悟时刻'。
  • 视觉信息难以融入模型的自我校验过程,适合研究推理机制者关注。

推理时技术如解码缩放和自精炼已被证明能显著提升大语言模型的数学推理能力,这主要归因于通过强化学习(RL)激发的涌现式自纠错与自验证行为。本文探讨该方法是否适用于视觉语言模型(VLMs),尤其是声称具备强视觉数学推理能力的RL微调版本。通过广泛评估,我们得出三个关键发现:第一,生成阶段能力比验证与精炼更重要,简单多数投票始终显著优于以验证为核心的策略(如带自验证的最佳N选一);第二,常与RL微调模型关联的‘顿悟时刻’行为在推理时并未带来可靠的性能提升;第三,视觉信息未能有效融入模型的自验证过程。总体而言,当前强化学习训练的VLMs在视觉模态中从自验证中获益甚微,制约了推理时扩展对视觉数学推理的有效性。

原文摘要 · Abstract (English)

Inference time techniques such as decoding time scaling and self refinement have been shown to substantially improve mathematical reasoning in large language models (LLMs), largely attributed to emergent self correction and self verification behaviors often elicited through reinforcement learning (RL). In this work, we ask whether the same recipe transfers to vision language models (VLMs), especially RL finetuned variants that claim strong visual mathematical reasoning. Through extensive evaluation, we reach three main findings that differ markedly from text only models. First, generation time capability matters more than verification and refinement: simple majority voting consistently and substantially outperforms verification centric strategies such as best of N with self verification. Second, behaviors often associated with RL tuned models at inference time, such as the 'Aha moment,' do not yield reliable reasoning performance improvements. Third, visual information is not effectively integrated into the model's self verification process. Overall, our analysis highlights a key limitation: current RL trained VLMs derive limited benefit from self verification in the visual modality, which constrains the effectiveness of inference time scaling for visual mathematical reasoning.

视觉语言模型推理增强自验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。