让音频分类器的解释直接以时间域音频形式输出,音质更好且不失准确性。
LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
- 用时间域解码器直接生成可听音频解释
- 用户研究显示音质显著提升,忠实度无下降
- 适合需要可听解释的音频模型可解释性研究
神经网络通常为黑箱,决策机制不透明。现有研究提出多种后验解释方法以缓解此问题。本文提出 LMAC-TD,一种基于 L-MAC(可听音频分类器映射)的后验解释方法,通过训练解码器直接在时间域生成解释。该方法引入 SepFormer——一种流行的基于 Transformer 的时域语音分离架构。用户研究表明,LMAC-TD 显著提升了生成解释的音频质量,同时未牺牲解释的忠实度。
原文摘要 · Abstract (English)
Neural networks are typically black-boxes that remain opaque with regards to their decision mechanisms. Several works in the literature have proposed post-hoc explanation methods to alleviate this issue. This paper proposes LMAC-TD, a post-hoc explanation method that trains a decoder to produce explanations directly in the time domain. This methodology builds upon the foundation of L-MAC, Listenable Maps for Audio Classifiers, a method that produces faithful and listenable explanations. We incorporate SepFormer, a popular transformer-based time-domain source separation architecture. We show through a user study that LMAC-TD significantly improves the audio quality of the produced explanations while not sacrificing from faithfulness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。