arXiv:2509.25848cs.CVcs.AI2025-09中稿 · ICLR被引 28

推理让视觉语言模型更会思考,但也可能忘了看图。

More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models

  • 用视觉锚定策略引导推理,避免模型忽视图像
  • 在多个基准上达到新最好成绩,提升视觉理解能力
  • 适合关注多模态模型可靠性与视觉对齐的研究者

推理已成为大型语言模型的核心能力。通过强化学习(通常为组相对策略优化,GRPO),这些模型能够解决数学和代码生成等复杂任务。在此基础上,近期研究尝试将推理能力扩展至视觉语言模型(VLMs),在多种视觉任务中取得良好效果。然而,我们的研究揭示了多模态推理的双重特性:虽然显著提升了逻辑推理能力并帮助解决难题,但可能逐渐损害感知基础,导致对简单视觉问题识别失败。进一步分析表明,这一现象源于视觉遗忘——长时间推理使模型逐渐忽略视觉输入。为此,我们提出视觉锚定策略优化(VAPO),一种简单而有效的方法,能明确引导推理过程朝向视觉一致路径。所提出的VAPO-Thinker-7B模型显著增强了对视觉信息的依赖,在广泛基准测试中达到新最优性能。项目页面:https://xytian1008.github.io/VAPO/

原文摘要 · Abstract (English)

Reasoning has emerged as a pivotal capability in Large Language Models (LLMs). Through Reinforcement Learning (RL), typically Group Relative Policy Optimization (GRPO), these models are able to solve complex tasks such as mathematics and code generation. Building on these advances, recent research has sought to extend reasoning to Vision-Language Models (VLMs), yielding promising results across diverse visual tasks. Despite this progress, our study uncovers the dual nature of multimodal reasoning: while it substantially enhances logical inference and facilitates performance on challenging problems, it may gradually impair perceptual grounding, leading to recognition failures on otherwise basic visual questions. Through further analysis, we attribute this phenomenon to visual forgetting, wherein prolonged reasoning causes the model to increasingly disregard visual input. To address this, we propose Vision-Anchored Policy Optimization (VAPO), a simple yet effective method that explicitly steers the reasoning process toward visually grounded trajectories. Our result model, VAPO-Thinker-7B, significantly strengthens the model's reliance on visual information and achieves new state-of-the-art results on a wide range of established benchmarks. Project page: https://xytian1008.github.io/VAPO/

多模态推理视觉对齐模型可靠

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。