让AI模型真正看懂图像,通过对比视觉变化来提升推理能力
Thinking with Deltas: Incentivizing Reinforcement Learning via Differential Visual Reasoning Policy
- 用原始图、遮蔽图和扰动图三者差异驱动模型推理
- 在无外部标注情况下,医学与通用任务性能超越现有方法
- 防止模型依赖语言线索,强制关注真实视觉信息
基于可验证奖励的强化学习(RLVR)显著提升了大语言模型的推理能力。然而,将RLVR应用于多模态领域时面临关键问题:感知与推理的脱节。现有方法以文本结果奖励为导向,推理过程仅在语言层面进行,无意中诱导模型绕过视觉感知。我们通过盲测实验证实:即使完全移除视觉输入,当前最优策略仍能保持甚至提升性能,表明这些模型退化为‘盲推理者’,仅依赖语言先验生成看似合理的答案。为此,我们提出‘思考差值’框架,核心为差分视觉推理策略(DVRP)。DVRP通过原始、遮蔽和扰动输入构成的视觉三元组提供内在监督,促使模型在遮蔽输入下最大化推理差异(强化视觉敏感性),同时在扰动输入下最小化差异(确保视觉鲁棒性)。通过严格对齐推理变化与视觉信息的‘差值’,DVRP自然增强视觉理解能力,在无需外部标注或辅助工具的前提下,显著优于现有方法,在通用与医学基准上均取得更优表现。
原文摘要 · Abstract (English)
Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced reasoning capabilities in Large Language Models. However, adapting RLVR to multimodal domains suffers from a critical \textit{perception-reasoning decoupling}. Existing paradigms, driven by text-centric outcome rewards, reasoning in language medium, inadvertently encourage models to bypass visual perception. We empirically validate this through blind experiments: state-of-the-art policies maintain or surprisingly improve performance even when visual inputs are entirely removed. This reveals that these models degenerate into \textit{blind reasoners}, exploiting linguistic priors to generate plausible answers instead of attending to visual evidence. In response, we propose \textbf{Thinking with Deltas}, a framework driven by a \textbf{Differential Visual Reasoning Policy (DVRP)}. DVRP introduces intrinsic supervision via visual triplets, comprising original, masked, and perturbed inputs. It optimizes the model to maximize reasoning divergence from masked inputs (enforcing \textit{visual sensitivity}) while minimizing divergence from perturbed inputs (ensuring \textit{visual robustness}). By aligning reasoning variations strictly with the \textit{Delta} of visual information, DVRP inherently bolsters visual understanding capabilities and significantly outperforms state-of-the-art methods on both general and medical benchmarks, without requiring external annotations or auxiliary tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。