arXiv:2606.18924cs.SD2026-06被引 3

揭示音频大模型文本主导偏见的内部机制并提出无训练干预方案

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs

论文配图:Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs
图 1 · 摘自论文原文
  • 通过追踪层间表征传播,解析多模态冲突下的内部行为机制
  • 发现文本路径主动压制音频信息,而非消除其内容
  • 提出回补技术增强音频表征,有效缓解文本主导偏差

尽管音频大语言模型在多模态理解上表现优异,但存在文本主导偏见——模型盲目偏好文本信息而忽视声学证据,导致幻觉。然而,当音频与文本输入矛盾时,模型内部如何运作仍不明确。本文首次进行机制性分析,追踪各层内部表征传播。研究发现:(i) 文本主导现象在不同模型中系统性存在;(ii) 文本与音频虽通过功能独立路径处理,但在深层最终汇聚至共享语义空间;(iii) 文本路径并非抹除音频信息,而是主动抑制完整音频表征。基于此,我们提出无需训练的回补(back-patching)干预方法,将深层音频激活重新注入早期层,增强音频表征以突破文本压制。实验表明,该方法持续降低文本主导程度,为冲突情境下的机制性多模态对齐提供新路径。

原文摘要 · Abstract (English)

While Audio Large Language Models (Audio LLMs) excel at multimodal understanding, they suffer from text dominance, a bias where models blindly favor text over acoustic evidence, causing hallucinations. However, the internal mechanisms underlying how these models behave when audio and textual inputs contradict each other remain unexplored. In this work, we present the first mechanistic analysis of this phenomenon by tracing the propagation of internal representations across layers. Our investigation reveals three key findings: (i) text dominance is systematically and empirically across models; (ii) while text and audio rely on functionally distinct pathways, they ultimately converge into a shared semantic space in late layers; and (iii) the text pathway does not erase audio information, but rather actively suppresses intact audio representations. Building on these insights, we leverage back-patching, a training-free intervention that routes late-layer audio activations back into earlier layers. This amplifies the audio representations, enabling them to overcome textual suppression. Our evaluation shows that back-patching consistently reduces text dominance, paving the way for mechanistic multimodal alignment under conflict.

多模态可解释性音频生成模型偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。