arXiv:2608.27026cs.SDcs.MM2026-08

不同任务下大音频模型的信息传递路径差异显著,影响其泛化能力。

Direct or Mediated? Task-Dependent Audio Information Routing in Large Audio Language Models

  • 通过分层注意力剔除法分析信息路由路径
  • ASR依赖直接读取音频,AQA依赖中间提示整合
  • 提示词中的音频信息仍可解码,问题出在后续利用

大型音频语言模型(LALMs)在多种音频理解任务中表现优异,但通常仅在单一连贯音频片段上评估,其在非标准输入下的行为尚不明确。本文通过将两个音频片段拼接为单个输入的受控实验,发现任务相关性导致显著性能差异:自动语音识别(ASR)保持稳定,而音频问答(AQA)性能大幅下降。进一步分析显示,两种任务采用不同信息路径:ASR主要通过答案词直接检索音频标记,而AQA更依赖音频信息先整合进提示词再生成答案的间接路径。尽管在拼接输入下AQA性能急剧恶化,我们仍可在中后层解码器中有效解码与任务相关的音频特征。这表明失败并非因音频信息丢失,而是下游生成阶段难以获取或利用提示词中介的信息。研究揭示了大音频模型中的任务依赖信息路由机制,提示信息利用能力可能制约其泛化性能。

原文摘要 · Abstract (English)

Large Audio Language Models (LALMs) have demonstrated strong performance across a wide range of audio understanding tasks. However, they are typically evaluated on single, coherent audio segments, leaving their behavior under less familiar input configurations underexplored. We study this issue through a controlled setting in which two audio segments are concatenated into a single input. Across multiple LALMs, we observe a striking task-dependent robustness gap: automatic speech recognition (ASR) remains comparatively stable, whereas audio question answering (AQA) degrades substantially. To investigate the mechanisms underlying this disparity, we analyze how audio information is routed through LALM decoders using layer-wise attention knockout. The results reveal distinct task-dependent pathways. ASR relies primarily on direct retrieval from audio tokens by answer tokens, whereas AQA depends more strongly on a mediated route in which audio information is first integrated into prompt tokens and subsequently accessed during generation. We further probe prompt-token representations under audio concatenation and find that task-relevant audio attributes remain readily decodable, particularly in middle and later decoder layers, even when AQA performance deteriorates sharply. This dissociation indicates that the failure cannot be explained by complete loss of audio information from the decoder states and is instead consistent with a downstream bottleneck in retrieving or utilizing prompt-mediated information during answer generation. Together, our findings reveal task-dependent audio information routing in LALMs and highlight information utilization as a potential limitation on their generalization.

音频模型信息路由大模型机制AQA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。