arXiv:2609.04637cs.CLcs.AI2026-09

揭示音频大模型如何真正依赖声音而非文字作答

Tracing Audio Grounding and Answer Selection in Audio LLMs

论文配图:Tracing Audio Grounding and Answer Selection in Audio LLMs
图 1 · 摘自论文原文
  • 通过对比训练前后模型对无声或无关音频的敏感度,追踪音频影响路径
  • 早期到中期层的表征受声学信息主导,后期层则强化音频对最终答案的影响
  • 训练中特定层级权重变化最大,是音频证据被采纳的关键区域

音频大语言模型(Audio LLMs)在音频理解方面取得进展,但仍可能仅凭文本线索或语言先验推断答案。常见做法是使用答案无法仅从文本推断的数据进行训练,虽能提升性能,但模型内部的变化机制尚不明确。本文探究使音频真正决定答案所需发生的内部变化。研究发现:(1) 将音频替换为静音或无关音频时,训练后模型性能下降远大于预训练模型;(2) 声学信息在早期至中期层最显著影响答案选项的表征,而训练主要增强中晚期层中音频对最终预测的影响;(3) 训练过程中,特定层级的权重变化最大。这些结果提供了训练如何加强音频证据使用的机制性解释。

原文摘要 · Abstract (English)

Audio Large Language Models (Audio LLMs) have advanced in audio understanding, yet they can still predict the answer by reasoning from textual cues or linguistic priors rather than the provided audio. A common remedy is to train models on data whose answers cannot be inferred from text alone. This approach can improve performance, but what changes within the model remains unclear. In this paper, we ask what must happen inside the model for the audio to actually determine the answer. Our findings are threefold. (1) Replacing the audio with silence or unrelated audio causes substantially larger performance degradation in the trained model than in the pretrained model. (2) Acoustic information most strongly shapes the model's representations of the answer choices in early-to-middle layers, while training mainly increases the influence of audio information on the final prediction in middle-to-late layers. (3) The weights learned during training have their largest impact in specific layer bands. Together, these results provide a mechanistic account of how training strengthens the use of acoustic evidence in Audio LLMs.

音频理解大模型机制因果分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。