揭示多模态模型在信息冲突时的决策规律,关键看信心差距和固有偏好。
When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs
- 用不确定性差值分解模型选择模态的内在机制
- 发现信心越低的模态被跟随概率越小,存在平衡点
- 适合研究多模态推理、模型可解释性的学者参考
多模态大模型在不同模态提供矛盾信息时需做出选择,这一过程称为模态跟随。以往研究仅依赖粗粒度的数据集统计,忽略了模型对单模态推理的信心影响。本文提出新框架,将模态跟随分解为两个核心因素:相对推理不确定性(单模态预测间的信心差异)和固有模态偏好(当不确定性均衡时模型的稳定倾向)。通过构建可控数据集系统调节视觉与文本输入的推理难度,使用熵作为细粒度不确定性指标,发现模态跟随概率随其相对不确定性增加而单调下降。在模型对两模态跟随概率相近的平衡点处,可提取其固有偏好作为有效指标。该方法摆脱传统宏观比率的混淆,更准确刻画模态偏差,分离出模型能力与数据偏差的影响。进一步层间探查显示,在平衡点附近的模糊区域,模型在不同层间出现模态振荡,解释了外部观察到的犹豫现象。研究确立相对不确定性与固有偏好为模态跟随的双重支配原则,提供了量化框架与机制理解。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) must resolve conflicts when different modalities provide contradictory information, a process we term modality following. Prior work measured this behavior only with coarse dataset-level statistics, overlooking the influence of model's confidence in unimodal reasoning. In this paper, we introduce a new framework that decomposes modality following into two fundamental factors: relative reasoning uncertainty (the case-specific confidence gap between unimodal predictions) and inherent modality preference( a model's stable bias when uncertainties are balanced). To validate this framework, we construct a controllable dataset that systematically varies the reasoning difficulty of visual and textual inputs. Using entropy as a fine-grained uncertainty metric, we uncover a universal law: the probability of following a modality decreases monotonically as its relative uncertainty increases. At the relative difficulty level where the model tends to follow both modalities with comparable probability what we call the balance point, a practical indicator of the model's inherent preference. Unlike traditional macro-level ratios, this measure offers a more principled and less confounded way to characterize modality bias, disentangling it from unimodal capabilities and dataset artifacts. Further, by probing layer-wise predictions, we reveal the internal mechanism of oscillation: in ambiguous regions near the balance point, models vacillate between modalities across layers, explaining externally observed indecision. Together, these findings establish relative uncertainty and inherent preference as the two governing principles of modality following, offering both a quantitative framework and mechanistic insight into how MLLMs resolve conflicting information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。