通过频域扰动揭示模型依赖的特征,提升解释性。
Frequency-Aware Model Parameter Explorer: A new attribution method for improving explainability
- 基于FFT设计α加权扰动,分别调节高低频成分。
- 在ImageNet上对四类模型测试,最高提升12.04%。
- 无需人工基线,可自适应探索频域特征重要性。
当前先进归因方法依赖对抗样本生成,但普遍采用全通滤波器,丢弃对深度神经网络特征归因至关重要的高频细节。我们提出一种新方法——频域感知参数探索器(FAMPE),通过选择性扰动高低频成分,探查模型依赖的频谱特征,并将频域分析直接转化为归因信号。该方法采用基于FFT的α加权扰动策略,通过能量驱动的频谱截断分别调控高频与低频成分,关键在于首次将频域探索与模型参数归因直接结合。不同于以往侧重迁移性或不可察觉性的方法,FAMPE专为解释性设计,在固定α=0.1条件下,于ImageNet上对四类架构(含CNN与Vision Transformers)测试,相较AttEXplore在Inception-v3上提升4.25%,在MaxViT-T上提升12.04%;样本级最优选择显示,低频主导图像显著受益于高频扰动,表明自适应频域探索潜力巨大。消融实验确认高频扰动对归因精度贡献更大,而过度低频噪声会破坏全局结构一致性。
原文摘要 · Abstract (English)
State-of-the-art attribution methods rely on adversarial sample generation that applies an all-pass filter across the frequency spectrum, discarding fine-grained high-frequency information that is demonstrably important for accurate feature attribution in deep neural networks. By generating adversarial samples that selectively perturb high- and low-frequency components, we can probe which spectral features a model relies on most -- directly translating frequency-domain exploration into attribution signals. Building on this insight, we propose FAMPE (Frequency-Aware Model Parameter Explorer), a novel attribution method that introduces an FFT-based α-weighted perturbation scheme -- separately modulating high- and low-frequency components via an energy-driven spectral cutoff -- and, crucially, integrates this frequency-aware exploration directly into model parameter exploration for attribution, a connection that has not been established in prior work. Unlike prior frequency-aware adversarial approaches that target transferability or imperceptibility, FAMPE's specific formulation is designed and validated exclusively for explainability, translating spectral structure into fine-grained attribution maps without requiring any manual baseline selection. Evaluated on ImageNet across four architectures spanning CNNs and Vision Transformers, at fixed α= 0.1 FAMPE outperforms AttEXplore by 4.25% on Inception-v3 and 12.04% on MaxViT-T, with per-sample oracle selection further revealing that low-frequency-dominated images systematically benefit from high-frequency perturbations -- underscoring the potential of adaptive spectral exploration. Our ablation studies confirm that high-frequency perturbations are disproportionately responsible for attribution precision, while excessive low-frequency noise degrades global structural coherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。