发现大模型更偏好视觉信息而非文本,且这种偏好随推理过程逐步形成。
Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models

- 通过冲突基准测试量化多模态模型的模态偏好
- 十款主流模型中多数呈现显著视觉偏好而非文本主导
- 利用中间层信号诊断跨模态幻觉,无需额外训练数据
原生多模态大语言模型(OLLMs)已从分步架构转向统一表征空间,但其内部存在尚未充分探索的模态偏好现象。本文构建了一个基于冲突的新型基准与模态选择率度量,系统量化了十款代表性OLLMs的模态偏好。结果揭示显著范式转变:不同于传统视觉语言模型的文本主导,多数OLLMs表现出明显的视觉偏好。进一步层间探查表明,该偏好并非静态,而是在中后层逐步涌现。基于此机制,我们利用内部信号诊断跨模态幻觉,在三个下游多模态基准上取得媲美监督方法的表现,且无需任务特定数据。本工作为理解并提升OLLM可信性提供了机制解释与实用工具。代码与资源已开源:https://github.com/icip-cas/OmniPreference。
原文摘要 · Abstract (English)
Native Omni-modal Large Language Models (OLLMs) have shifted from pipeline architectures to unified representation spaces. However, this native integration gives rise to a critical yet underexplored phenomenon: modality preference. To bridge this gap, we first systematically quantify modality preference of OLLMs using a newly-curated conflict-based benchmark and the modality selection rate metric. Our evaluation of ten representative OLLMs reveals a notable paradigm shift: unlike the ``text-dominance'' of traditional VLMs, most OLLMs exhibit a pronounced visual preference. To further understand the underlying mechanism, we conduct layer-wise probing and demonstrate that such modality preference is not static but emerges progressively in the mid-to-late layers. Building upon these insights, we leverage these internal signals to diagnose cross-modal hallucinations, achieving competitive performance across three downstream multi-modal benchmarks without task-specific data. Our work provides both a mechanistic understanding and a practical tool for building more trustworthy OLLMs. Our code and related resources are publicly available at: https://github.com/icip-cas/OmniPreference
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。