arXiv:2511.21397cs.CVcs.AI2025-11被引 2

研究视觉干扰项如何影响多模态模型推理,发现其会降低准确率但不延长推理过程。

Understanding the Effects of Distractors on Reasoning Vision-Language Models

  • 构建含干扰项的图像问答数据集 Idis,系统控制语义与数量维度
  • 视觉干扰项导致准确率下降,但推理长度不变,与文本干扰不同
  • 属性计数可揭示干扰项对推理过程的影响机制,适合模型可解释性研究

视觉语言模型在测试时扩展中,无关信息(即干扰项)如何影响推理表现?已有研究表明,文本干扰项会加剧反向缩放现象,使模型推理更长却更无效。本文探究该现象在多模态场景中的存在性。我们提出 Idis(Images with distractors)数据集,系统地在语义和数值维度上调节视觉干扰项。分析显示,视觉干扰项对推理型视觉语言模型的影响与文本干扰截然不同:尽管反向缩放仍存在,但视觉干扰项会降低准确率,而不会增加推理长度。进一步发现,从推理轨迹中提取的属性计数能有效揭示干扰项与推理长度及准确率之间的交互关系。作为验证,我们提出一种简单提示策略,可缓解模型受干扰项驱动的错误预测。

原文摘要 · Abstract (English)

How does irrelevant information (i.e., distractors) affect test-time scaling in vision-language models (VLMs)? Prior work on text-only language models has shown that textual distractors can intensify inverse scaling, causing models to reason longer but less effective reasoning traces. In this work, we investigate whether similar phenomena arise in multimodal settings. We introduce Idis (Images with distractors), a visual question-answering dataset that systematically varies distractors along semantic and numerical dimensions. Our analyses reveal that visual distractors affect reasoning VLMs in a fundamentally different way from textual distractors: although inverse scaling still emerges, visual distractors reduce accuracy without increasing reasoning length. We further show that attribute counts extracted from reasoning traces provide key insights into how distractors interact with reasoning length and accuracy. As a sanity check, we propose a simple prompting strategy that mitigates distractor-driven predictions in reasoning vision-language models.

视觉语言模型干扰项推理分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。