arXiv:2601.06801cs.AIcs.LG2026-01被引 6

让AI模型真正看懂图像,通过对比视觉变化来提升推理能力

Thinking with Deltas: Incentivizing Reinforcement Learning via Differential Visual Reasoning Policy

  • 用原始图、遮蔽图和扰动图三者差异驱动模型推理
  • 在无外部标注情况下,医学与通用任务性能超越现有方法
  • 防止模型依赖语言线索,强制关注真实视觉信息

基于可验证奖励的强化学习(RLVR)显著提升了大语言模型的推理能力。然而,将RLVR应用于多模态领域时面临关键问题:感知与推理的脱节。现有方法以文本结果奖励为导向,推理过程仅在语言层面进行,无意中诱导模型绕过视觉感知。我们通过盲测实验证实:即使完全移除视觉输入,当前最优策略仍能保持甚至提升性能,表明这些模型退化为‘盲推理者’,仅依赖语言先验生成看似合理的答案。为此,我们提出‘思考差值’框架,核心为差分视觉推理策略(DVRP)。DVRP通过原始、遮蔽和扰动输入构成的视觉三元组提供内在监督,促使模型在遮蔽输入下最大化推理差异(强化视觉敏感性),同时在扰动输入下最小化差异(确保视觉鲁棒性)。通过严格对齐推理变化与视觉信息的‘差值’,DVRP自然增强视觉理解能力,在无需外部标注或辅助工具的前提下,显著优于现有方法,在通用与医学基准上均取得更优表现。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced reasoning capabilities in Large Language Models. However, adapting RLVR to multimodal domains suffers from a critical \textit{perception-reasoning decoupling}. Existing paradigms, driven by text-centric outcome rewards, reasoning in language medium, inadvertently encourage models to bypass visual perception. We empirically validate this through blind experiments: state-of-the-art policies maintain or surprisingly improve performance even when visual inputs are entirely removed. This reveals that these models degenerate into \textit{blind reasoners}, exploiting linguistic priors to generate plausible answers instead of attending to visual evidence. In response, we propose \textbf{Thinking with Deltas}, a framework driven by a \textbf{Differential Visual Reasoning Policy (DVRP)}. DVRP introduces intrinsic supervision via visual triplets, comprising original, masked, and perturbed inputs. It optimizes the model to maximize reasoning divergence from masked inputs (enforcing \textit{visual sensitivity}) while minimizing divergence from perturbed inputs (ensuring \textit{visual robustness}). By aligning reasoning variations strictly with the \textit{Delta} of visual information, DVRP inherently bolsters visual understanding capabilities and significantly outperforms state-of-the-art methods on both general and medical benchmarks, without requiring external annotations or auxiliary tools.

强化学习多模态视觉推理语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。