发现大模型更信文字而非音频,导致判断失误。
When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
- 构建首个音频文本冲突测试集MCR-BENCH
- 模型在音频文字矛盾时90%以上倾向选文字
- 适合关注多模态可靠性的研究者与工程师
大型音频-语言模型(LALMs)具备音频感知能力,可处理音视频与文本的多模态输入。然而,其在面对音频与文本信息冲突时的表现尚未充分研究。本文提出MCR-BENCH,首个专门评估LALMs在不一致音文对中信息优先级的综合性基准。通过在多种音频理解任务上的广泛评测,发现当音文信息冲突时,模型表现出显著的文字偏倚,频繁忽视音频证据,导致音频主导任务性能大幅下降,引发实际应用中的可靠性问题。我们进一步分析了文本偏倚的影响因素,探索了监督微调等缓解策略,并发现模型即使面对矛盾输入仍存在持续高置信度。这些结果凸显了训练中需加强模态平衡和设计更稳健的融合机制。
原文摘要 · Abstract (English)
Large Audio-Language Models (LALMs) are enhanced with audio perception capabilities, enabling them to effectively process and understand multimodal inputs that combine audio and text. However, their performance in handling conflicting information between audio and text modalities remains largely unexamined. This paper introduces MCR-BENCH, the first comprehensive benchmark specifically designed to evaluate how LALMs prioritize information when presented with inconsistent audio-text pairs. Through extensive evaluation across diverse audio understanding tasks, we reveal a concerning phenomenon: when inconsistencies exist between modalities, LALMs display a significant bias toward textual input, frequently disregarding audio evidence. This tendency leads to substantial performance degradation in audio-centric tasks and raises important reliability concerns for real-world applications. We further investigate the influencing factors of text bias, and explore mitigation strategies through supervised finetuning, and analyze model confidence patterns that reveal persistent overconfidence even with contradictory inputs. These findings underscore the need for improved modality balance during training and more sophisticated fusion mechanisms to enhance the robustness when handling conflicting multi-modal inputs. The project is available at https://github.com/WangCheng0116/MCR-BENCH.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。