arXiv:2608.19208cs.CLcs.CV2026-08

无关文本会系统性扭曲多模态模型的视觉判断,且影响可量化。

When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models

论文配图:When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models
图 1 · 摘自论文原文
  • 通过控制提示结构,研究无关文本对视觉判断的影响机制。
  • 发现无关文本使模型决策边距产生一致的仿射变换,非随机噪声。
  • 提出用仿射参数衡量模型对视觉信息的忠实度和回答倾向性。

多模态大语言模型(MLLMs)常面临辅助文本上下文,其对视觉任务的影响尚未深入探究。本文在二元视觉判断框架下,将无关文本作为可控干预变量,保持提示结构不变而改变辅助输入。结果发现,无关文本在多个基准测试中均持续干扰模型预测。为超越性能指标,我们通过候选答案的对数概率差定义决策边距,分析表明:条件边距与无条件边距之间存在稳定的仿射关系。这一规律揭示无关文本并非无序随机噪声,而是可估计的偏好偏移。进一步地,拟合出的仿射参数可作为视觉承诺保持度与方向性回答偏见的度量。该工作为理解无关上下文影响提供了边距层面的诊断视角,并为未来噪声上下文鲁棒性研究奠定基础。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored. In this paper, we investigate the influence of task-irrelevant context by formulating it as a controlled intervention within a binary visual judgment framework. By maintaining an invariant prompt structure while varying auxiliary inputs, we observe that irrelevant text consistently biases model predictions across diverse benchmarks. To move beyond performance metrics, we characterize this sensitivity through a decision margin defined by the log-probability difference between binary candidates. Our analysis reveals a robust geometric regularity: contextconditioned margins follow a consistent affine transformation of their context-free counterparts. This finding demonstrates that irrelevant context does not manifest as unstructured stochastic noise but as a estimable distortion of model preference. We further interpret the fitted affine parameters as metrics for visual commitment preservation and directional answer bias. These findings provide a margin-level diagnostic view of irrelevant-context effects in MLLMs and offer a basis for future studies on noisy-context robustness

多模态模型偏见决策边距鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。