arXiv:2511.15613cs.CVcs.CL2025-11中稿 · CVPR被引 4

通过不确定性引导回看,提升视觉语言模型的推理准确性

When to Think and When to Look: Uncertainty-Guided Lookback

  • 基于图像不确定性动态触发回看提示,优化思维链路径
  • 在MMMU数据集上超越标准思维链,尤其在弱项任务中提升显著
  • 无需训练,适用于多种多模态推理场景

测试时思考(即生成显式的中间推理链)已被证明可提升大语言模型性能,并在大视觉语言模型(LVLMs)中取得显著进展。然而,对思考如何影响视觉推理仍缺乏系统分析。本文首次通过大规模、受控实验对比了来自InternVL3.5和Qwen3-VL系列的十种变体,在MMMU-val数据集上采用充足上下文长度与多轮解码策略进行评估。结果表明,更长的推理链并不总带来更好表现;过长的链常导致偏离图像的错误轨迹,反而低于标准指令模式下的性能。深入分析发现,成功轨迹中显著富含明确回看图像的短语。基于此,我们提出无训练的不确定性引导回看策略,结合不确定性信号与自适应回看提示及广度搜索。该方法在固定模型族和令牌预算下,提升了总体MMMU表现,尤其在标准思考表现较弱的任务中收益最大,优于多个强基线,创下新基准。进一步验证表明,该策略在五个额外基准上具泛化能力,包括两个广泛的多模态套件和数学导向的视觉推理数据集。

原文摘要 · Abstract (English)

Test-time thinking (that is, generating explicit intermediate reasoning chains) is known to boost performance in large language models and has recently shown strong gains for large vision language models (LVLMs). However, despite these promising results, there is still no systematic analysis of how thinking actually affects visual reasoning. We provide the first such analysis with a large scale, controlled comparison of thinking for LVLMs, evaluating ten variants from the InternVL3.5 and Qwen3-VL families on MMMU-val under generous token budgets and multi pass decoding. We show that more thinking is not always better; long chains often yield long wrong trajectories that ignore the image and underperform the same models run in standard instruct mode. A deeper analysis reveals that certain short lookback phrases, which explicitly refer back to the image, are strongly enriched in successful trajectories and correlate with better visual grounding. Building on this insight, we propose uncertainty guided lookback, a training free decoding strategy that combines an uncertainty signal with adaptive lookback prompts and breadth search. Our method improves overall MMMU performance, delivers the largest gains in categories where standard thinking is weak, and outperforms several strong decoding baselines, setting a new state of the art under fixed model families and token budgets. We further show that this decoding strategy generalizes, yielding consistent improvements on five additional benchmarks, including two broad multimodal suites and math focused visual reasoning datasets.

视觉推理思维链解码策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。