长推理链提升数学能力却加剧幻觉,研究发现视觉注意力随推理变弱。
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
- 引入新指标RH-AUC,量化推理长度对视觉感知的影响。
- 大模型在推理与视觉真实间平衡更好,训练数据类型比数量更重要。
- 适合关注多模态模型可信度与评估框架的研究者。
测试时计算能力使多模态大语言模型生成更长的推理链,在多模态数学推理等任务上表现优异。然而,这种推理能力的提升常伴随幻觉增加:随着生成内容变长,模型越来越偏离图像内容,更依赖语言先验。注意力分析显示,更长的推理链导致对视觉输入的关注减少,从而引发幻觉。为系统研究此现象,我们提出RH-AUC指标,量化模型感知准确率随推理长度的变化,用于评估模型在推理过程中是否保持视觉接地性。同时发布RH-Bench诊断基准,涵盖多种多模态任务,以评估推理能力与幻觉之间的权衡。分析发现:(i) 更大的模型通常在推理与感知之间实现更好的平衡;(ii) 这种平衡更多受训练数据类型和领域影响,而非总量。这些结果凸显了需同时考量推理质量与感知保真度的评估框架的重要性。
原文摘要 · Abstract (English)
Test-time compute has empowered multimodal large language models to generate extended reasoning chains, yielding strong performance on tasks such as multimodal math reasoning. However, this improved reasoning ability often comes with increased hallucination: as generations become longer, models tend to drift away from image-grounded content and rely more heavily on language priors. Attention analysis shows that longer reasoning chains lead to reduced focus on visual inputs, which contributes to hallucination. To systematically study this phenomenon, we introduce RH-AUC, a metric that quantifies how a model's perception accuracy changes with reasoning length, allowing us to evaluate whether the model preserves visual grounding during reasoning. We also release RH-Bench, a diagnostic benchmark that spans a variety of multimodal tasks, designed to assess the trade-off between reasoning ability and hallucination. Our analysis reveals that (i) larger models typically achieve a better balance between reasoning and perception, and (ii) this balance is influenced more by the types and domains of training data than by its overall volume. These findings underscore the importance of evaluation frameworks that jointly consider both reasoning quality and perceptual fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。