arXiv:2509.12132cs.CVcs.CL2025-09EMNLP被引 22

提升视觉语言模型的视觉反思能力,让推理更依赖视觉信息。

Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models

  • 构建以视觉为中心的推理数据,冷启动学习视觉反思模式。
  • 强化学习中引入视觉注意力奖励机制,提升视觉信息利用率。
  • 在多个基准上表现更好,推理过程更稳定地依赖视觉输入。

近期文本类“慢思考”推理进展推动了将该能力迁移至视觉语言模型(VLMs)以训练视觉推理模型(VRMs)。然而,这种迁移面临关键挑战:有效的“慢思考”需具备基于视觉信息检查推理过程的能力,即“视觉反思”。定量分析显示,当前VRMs的视觉反思能力有限,其对视觉信息的关注随生成响应长度迅速下降。为此,我们提出新型VRM Reflection-V,通过推理数据构建实现冷启动,并结合强化学习中的奖励设计增强视觉反思。首先,利用代理在VLM与推理大模型间交互,构建以视觉为中心的推理数据,支持冷启动学习视觉反思模式;其次,在强化学习中引入基于视觉注意力的奖励模型,鼓励推理过程更多依赖视觉信息。实验表明,Reflection-V在多个视觉推理基准上均有显著提升,且推理过程中对视觉信息的依赖更强、更一致,有效增强了视觉反思能力。

原文摘要 · Abstract (English)

Recent advances in text-only "slow-thinking" reasoning have prompted efforts to transfer this capability to vision-language models (VLMs), for training visual reasoning models (\textbf{VRMs}). owever, such transfer faces critical challenges: Effective "slow thinking" in VRMs requires \textbf{visual reflection}, the ability to check the reasoning process based on visual information. Through quantitative analysis, we observe that current VRMs exhibit limited visual reflection, as their attention to visual information diminishes rapidly with longer generated responses. To address this challenge, we propose a new VRM \textbf{Reflection-V}, which enhances visual reflection based on reasoning data construction for cold-start and reward design for reinforcement learning (RL). Firstly, we construct vision-centered reasoning data by leveraging an agent that interacts between VLMs and reasoning LLMs, enabling cold-start learning of visual reflection patterns. Secondly, a visual attention based reward model is employed during RL to encourage reasoning based on visual information. Therefore, \textbf{Reflection-V} demonstrates significant improvements across multiple visual reasoning benchmarks. Furthermore, \textbf{Reflection-V} maintains a stronger and more consistent reliance on visual information during visual reasoning, indicating effective enhancement in visual reflection capabilities.

视觉推理慢思考强化学习视觉反思

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。