arXiv:2506.05140cs.CLcs.AI2025-06中稿 · ASRU 2025被引 17

首次揭示大音频语言模型如何感知声音属性,为优化提供关键依据。

AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models

  • 通过词汇投影追踪属性信息在层与位置的演变过程。
  • 早期层解析属性更准,失败时信息随深度下降。
  • 模型依赖直接查询输入而非隐藏状态整合信息,适合改进推理机制。

理解大音频语言模型(LALMs)的内部机制对于解释其行为和提升性能至关重要。本文首次深入分析了LALMs如何内部感知和识别声音属性。通过对三种先进LALMs应用词汇投影,我们追踪了属性信息在各层及词元位置上的演化。结果发现,当识别失败时,属性信息通常随网络深度增加而减少;而在早期层中解决属性问题与更高准确率相关。此外,LALMs主要依赖对音频输入的直接查询来预测属性,而非在提及属性的位置聚合隐藏状态中的必要信息。基于这些发现,我们提出一种增强LALMs的方法。研究结果为声音属性处理提供了新见解,有助于未来模型改进。

原文摘要 · Abstract (English)

Understanding the internal mechanisms of large audio-language models (LALMs) is crucial for interpreting their behavior and improving performance. This work presents the first in-depth analysis of how LALMs internally perceive and recognize auditory attributes. By applying vocabulary projection on three state-of-the-art LALMs, we track how attribute information evolves across layers and token positions. We find that attribute information generally decreases with layer depth when recognition fails, and that resolving attributes at earlier layers correlates with better accuracy. Moreover, LALMs heavily rely on querying auditory inputs for predicting attributes instead of aggregating necessary information in hidden states at attribute-mentioning positions. Based on our findings, we demonstrate a method to enhance LALMs. Our results offer insights into auditory attribute processing, paving the way for future improvements.

音频理解模型可解释性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。