arXiv:2410.01482stat.MLcs.AI2024-10ICML被引 4

用小波变换统一解释模型决策,比传统方法更准确且适配多模态数据。

One Wave To Explain Them All: A Unifying Perspective On Feature Attribution

  • 在小波域计算特征重要性,同时捕捉位置与尺度信息。
  • 在图像、音频、体积数据上表现优于或媲美现有梯度法。
  • 适合需要结构化解释的多模态模型可解释性研究者。

特征归因方法旨在通过识别影响模型决策的输入特征来提升深度神经网络的透明度。像素级热图已成为高维输入(如图像、音频表示、体数据)归因的标准方法,虽直观便捷,却难以捕捉数据的内在结构。此外,归因计算的领域选择常被忽视。本文证明,小波域可实现信息丰富且有意义的归因,适用于任意输入维度,并提供统一的特征归因框架。所提出的波形归因方法(Wavelet Attribution Method, WAM)利用小波系数的空间与尺度局部化特性,揭示模型决策的‘何处’与‘何物’。实验表明,WAM 在多个模态(包括音频、图像、体积)上定量匹配或超越现有基于梯度的方法。此外,我们探讨了WAM如何连接归因与模型鲁棒性及透明性的更广泛议题。

原文摘要 · Abstract (English)

Feature attribution methods aim to improve the transparency of deep neural networks by identifying the input features that influence a model's decision. Pixel-based heatmaps have become the standard for attributing features to high-dimensional inputs, such as images, audio representations, and volumes. While intuitive and convenient, these pixel-based attributions fail to capture the underlying structure of the data. Moreover, the choice of domain for computing attributions has often been overlooked. This work demonstrates that the wavelet domain allows for informative and meaningful attributions. It handles any input dimension and offers a unified approach to feature attribution. Our method, the Wavelet Attribution Method (WAM), leverages the spatial and scale-localized properties of wavelet coefficients to provide explanations that capture both the where and what of a model's decision-making process. We show that WAM quantitatively matches or outperforms existing gradient-based methods across multiple modalities, including audio, images, and volumes. Additionally, we discuss how WAM bridges attribution with broader aspects of model robustness and transparency. Project page: https://gabrielkasmi.github.io/wam/

模型解释小波变换特征归因

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。