提出新基准,揭示多模态大模型在感知与陈述信号间的偏见差异。
Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts

- 构建三模态冲突数据集,分离感知与陈述信号
- 发现视觉偏见普遍存在,听觉更依赖陈述信息
- 早期表征中已可解码偏见,难通过表面方法消除
多模态大语言模型(OLLMs)同时处理视觉、音频和文本,但跨模态冲突下的模态偏见尚未被充分研究。现有基准将同一模态内的两种证据形式混淆:感知信号(如狗的照片或录音)和命题信号(如‘这是只狗’的陈述),导致测量到的模态偏见无法清晰归因于任一来源。为此,我们提出Tri-PvP,一个包含8,000个样本的三模态冲突基准,涵盖视觉与音频的感知或命题形式。评估五个OLLM后,发现多数模型存在稳健的视觉偏见,且在不同证据形式下均成立。关键发现是证据形式偏见存在系统性不对称:模型对视觉更倾向感知信号,对音频则更依赖命题信号。通过层间线性探针与对比解码分析表明,模态偏见在早期表示层即可线性解码,且难以完全缓解,提示需超越表面干预的深层缓解策略。
原文摘要 · Abstract (English)
Omni-modal large language models (OLLMs) jointly process vision, audio, and text, yet their modality bias under cross-modal conflict remains underexplored. Existing benchmarks conflate two distinct forms of evidence within a single modality: perceptual signals (e.g., a photograph or recording of a dog) and propositional signals (e.g., the declarative claim "this is a dog"), such that any measured modality bias is inherently confounded with evidence-form bias, precluding clean attribution to either source. To address this, we introduce Tri-PvP, an 8,000-sample tri-modal conflict benchmark crossing vision, audio, and text, where vision and audio each take perceptual or propositional form. Evaluating five OLLMs, we find robust visual bias across most models and evidence-type conditions. Crucially, we reveal a systematic asymmetry in evidence-form bias: models exhibit a stronger bias toward perceptual signal in vision but propositional in audio. Further analyses via layer-wise linear probing and contrastive decoding reveal that modality bias is already linearly decodable from early representation layers and can only be partially mitigated, calling for mitigation strategies beyond surface-level interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。