揭示大模型如何融合声音与文本信息
Causal Tracing of Audio-Text Fusion in Large Audio Language Models
- 用因果追踪分析模型内部信息流动过程
- 发现不同模型有渐进或突然的融合策略
- 适合研究多模态模型机制的学者参考
尽管大型音频语言模型(LALMs)在各类任务中表现优异,但其如何以及在何处融合声学特征与文本上下文仍不明确。我们采用因果追踪方法,研究LALMs在音频理解过程中的内部信息流。通过对DeSTA、Qwen和Voxtral进行逐层与逐标记分析,评估各隐藏状态的因果影响。逐层分析揭示了不同的融合策略:DeSTA呈现渐进式整合,而Qwen则表现为突发式的晚期融合。逐标记分析显示,最终序列标记作为信息瓶颈,模型在此处决定性地从音频中检索相关信息。此外,在中间标记位置观察到类似注意力的查询机制,触发模型提取与任务相关的声音上下文。这些发现清晰刻画了多模态融合在LALMs中的发生时机与位置。
原文摘要 · Abstract (English)
Despite the strong performance of large audio language models (LALMs) in various tasks, exactly how and where they integrate acoustic features with textual context remains unclear. We adapt causal tracing to investigate the internal information flow of LALMs during audio comprehension. By conducting layer-wise and token-wise analyses across DeSTA, Qwen, and Voxtral, we evaluate the causal effects of individual hidden states. Layer-wise analysis identifies different fusion strategies, from progressive integration in DeSTA to abrupt late-stage fusion in Qwen. Token-wise analysis shows that the final sequence token acts as an informational bottleneck where the network decisively retrieves relevant information from the audio. We also observe an attention-like query mechanism at intermediate token positions that triggers the model to pull task-relevant audio context. These findings provide a clear characterization of when and where multi-modal integration occurs within LALMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。