arXiv:2605.10153cs.SDcs.LG2026-05

为音频分类模型提供无需微调的可解释性框架

APEX: Audio Prototype EXplanations for Classification Tasks

论文配图:APEX: Audio Prototype EXplanations for Classification Tasks
图 1 · 摘自论文原文
  • 基于四种声学维度生成可解释原型,分别捕捉瞬时事件、时间模式等
  • 在不改变原模型输出的前提下,实现对音频特征的精准定位与解释
  • 适合需要理解音频分类决策过程的研究者与工程师

可解释人工智能(XAI)在图像分类中取得显著进展,但音频领域尚缺乏成熟解决方案。现有方法将视觉归因技术直接应用于频谱图,忽略了视听信号的根本差异。尽管原型推理具有潜力,但声学相似性具有多维特性。本文提出 APEX(Audio Prototype EXplanations),一种用于预训练音频分类器的后处理解释框架。APEX 不需对原始主干网络进行微调,严格保持输出不变性。该框架将解释分解为四个视角:基于方块的原型用于定位瞬时事件,基于时间的原型揭示时间模式,基于频率的原型突出频带特征,基于时空联合的原型融合二者信息。由此生成直观、实例驱动的解释,符合声学特性,相比传统梯度方法提供更清晰的语义理解。

原文摘要 · Abstract (English)

Explainable AI (XAI) has achieved remarkable success in image classification, yet the audio domain lacks equally mature solutions. Current methods apply vision-based attribution techniques to spectrograms, overlooking fundamental differences between visual and acoustic signals. While prototype reasoning is promising, acoustic similarity remains multidimensional. We introduce APEX (Audio Prototype EXplanations), a post-hoc framework for interpreting pre-trained audio classifiers. Crucially, APEX requires no fine-tuning of the original backbone and strictly preserves output invariance. APEX disentangles explanations into four perspectives: Square-based prototypes to localize transient events, Time-based for temporal patterns, Frequency-based highlighting spectral bands, and Time-Frequency-based integrating both. This yields intuitive, example-based explanations that respect acoustic properties, providing greater semantic clarity than standard gradient-based methods.

可解释性音频分析原型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。