arXiv:2410.07530cs.SDcs.AI2024-10被引 6

用生成模型直接合成可听的音频解释,让声音模型的决策过程听得懂。

Audio Explanation Synthesis with Generative Foundation Models

  • 利用音频基础模型的生成能力,从嵌入空间找关键特征。
  • 在关键词识别和情感识别任务上生成可听解释,效果优于传统方法。
  • 适合想理解声音模型决策逻辑的研究者和开发者。

音频基础模型在各类任务中表现优异,但其决策过程缺乏可解释性,亟需改进。现有方法多通过输入空间元素对最终决策的影响来解释模型。本文提出一种新方法,利用音频基础模型的生成能力,结合特征归因技术,在嵌入空间中识别关键特征,并优先保留这些特征生成可听的音频解释。在关键词识别和语音情感识别等标准数据集上的实验证明,该方法能有效生成有意义的音频解释。

原文摘要 · Abstract (English)

The increasing success of audio foundation models across various tasks has led to a growing need for improved interpretability to understand their intricate decision-making processes better. Existing methods primarily focus on explaining these models by attributing importance to elements within the input space based on their influence on the final decision. In this paper, we introduce a novel audio explanation method that capitalises on the generative capacity of audio foundation models. Our method leverages the intrinsic representational power of the embedding space within these models by integrating established feature attribution techniques to identify significant features in this space. The method then generates listenable audio explanations by prioritising the most important features. Through rigorous benchmarking against standard datasets, including keyword spotting and speech emotion recognition, our model demonstrates its efficacy in producing audio explanations.

音频解释生成模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。