arXiv:2606.14466cs.SDcs.AI2026-06中稿 · ICML

音频深度伪造检测的解释易被操纵,模型判断不变却可乱画解释热图。

The Perceived Fragility of Explanations in Audio Models: Manipulation of Attribution with Unchanged Predictions

论文配图:The Perceived Fragility of Explanations in Audio Models: Manipulation of Attribution with Unchanged Predictions
图 1 · 摘自论文原文
  • 用听觉感知框架生成听不见的干扰,让解释热图失真但分类结果不变。
  • 在多个主流模型上验证,攻击成功率超90%,且解释与预测脱钩。
  • 适合关注AI可解释性安全的研究者和音频内容审核系统开发者。

本文研究音频深度伪造检测中后处理解释方法的脆弱性。以往解释操纵研究多聚焦图像并使用标准 $L_p$ 范数,而本文提出一种心理声学框架,通过优化听觉不可察觉的扰动,实现模型归因与最终分类的解耦。我们在严格保持预测不变的条件下,评估了当前最先进的多种架构。通过结合领域特定的感知音频质量指标与解释对齐度量,框架证明攻击者可在不改变深伪标签的前提下,系统性地扭曲自动化解释热图。完整代码已开源:https://github.com/cncPomper/Audio-XAI。

原文摘要 · Abstract (English)

This paper investigates the fragility of post-hoc explanation methods in audio deepfake detection. While previous work on explanation manipulation focused on images using standard $L_p$ metrics, we introduce a psychoacoustic framework that optimizes inaudible perturbations to decouple model attributions from final classifications. We evaluate this vulnerability across state-of-the-art architectures under strict prediction-preserving constraints. By evaluating the manipulation cost through domain-specific perceptual audio quality metrics alongside explanation alignment criteria, our framework demonstrates that an adversary can systematically distort automated explanation heatmaps while preserving the predicted deepfake label. Full code available at: https://github.com/cncPomper/Audio-XAI

音频深度伪造可解释AI对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。