arXiv:2504.14076cs.SDcs.LG2025-04中稿 · International Join…被引 6

将音频嵌入转化为可解释的概念表示,提升模型透明度。

Transformation of audio embeddings into interpretable, concept-based representations

  • 用对比学习模型CLAP构建音视频共享空间,实现语义对齐。
  • 转化后的概念表示在下游任务中表现不劣于原嵌入,且更可解释。
  • 发布三个专用音频概念词表,助力音频理解研究。

音频神经网络在下游任务中已达到顶尖水平,但其黑箱结构使得内部音频表示难以解读。本文利用CLAP(一种将音频与文本映射到共享嵌入空间的对比学习模型),探索了音频嵌入的语义可解释性。提出一种后处理方法,将CLAP嵌入转换为基于概念的稀疏表示,兼具可解释性与性能。定性和定量评估表明,该概念表示在下游任务上表现不劣于甚至优于原始嵌入。此外,微调概念表示可进一步提升性能。最后,作者发布了三个面向音频的专用概念词表,支持音频嵌入的可解释分析。

原文摘要 · Abstract (English)

Advancements in audio neural networks have established state-of-the-art results on downstream audio tasks. However, the black-box structure of these models makes it difficult to interpret the information encoded in their internal audio representations. In this work, we explore the semantic interpretability of audio embeddings extracted from these neural networks by leveraging CLAP, a contrastive learning model that brings audio and text into a shared embedding space. We implement a post-hoc method to transform CLAP embeddings into concept-based, sparse representations with semantic interpretability. Qualitative and quantitative evaluations show that the concept-based representations outperform or match the performance of original audio embeddings on downstream tasks while providing interpretability. Additionally, we demonstrate that fine-tuning the concept-based representations can further improve their performance on downstream tasks. Lastly, we publish three audio-specific vocabularies for concept-based interpretability of audio embeddings.

音频表示可解释性概念解析CLAP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。