研究发现语音大模型在音文冲突时更信文本,即使被指令听音频。
When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitration
- 构建跨语言音文冲突数据集,量化模型听音频的倾向性
- 多数模型在指令听音频时仍选择文本,错误率高达23.2%
- 模型决策偏好受输入可读性影响,非仅信息量决定
当音频与文本内容冲突时,具备语音能力的语言模型更倾向于遵循文本,而非音频,即便明确指示应听音频。我们提出ALME(Audio-LLM Modality Evaluation)数据集,包含8种语言下57,602个受控的音文冲突样本,并引入文本主导率(TDR)来衡量模型在被要求听音频时仍跟随矛盾文本的频率。Gemini 2.0 Flash和GPT-4o的TDR分别达16.6%和23.2%,远高于以转录文本替代音频的基线(分别为1.6%和0.9%)。结果表明,文本主导不仅源于信息量,也反映决策阶段对不同模态的可访问性差异。将转录文本标记为故意损坏,使TDR下降80%;强制显式转录则使TDR上升14%。微调消融实验进一步显示,仲裁行为更多取决于大模型推理能力,而非音频输入路径本身。在四种音频-大模型中均观察到相似定性模式,存在显著跨模型与跨语言差异。
原文摘要 · Abstract (English)
When audio and text conflict, speech-enabled language models follow text far more often than they do when arbitrating between two conflicting text sources, even under explicit instructions to trust the audio. We introduce ALME (Audio-LLM Modality Evaluation), a dataset of 57,602 controlled audio-text conflict stimuli across eight languages, together with Text Dominance Ratio (TDR), which measures how often a model follows conflicting text when instructed to follow audio. Gemini 2.0 Flash and GPT-4o show TDR 10--26$\times$ higher than a baseline that replaces audio with its transcript under otherwise identical conditions (Gemini 2.0 Flash: 16.6% vs. 1.6%; GPT-4o: 23.2% vs. 0.9%). These results suggest that text dominance reflects not only information content, but also an asymmetry in arbitration accessibility, i.e., how easily the model can use competing representations at decision time. Framing the transcript as deliberately corrupted reduces TDR by 80%, whereas forcing explicit transcription increases it by 14%. A fine-tuning ablation further suggests that arbitration behavior depends more on LLM reasoning than on the audio input path alone. Across four audio-LLMs, we observe the same qualitative pattern with substantial cross-model and cross-linguistic variation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。