arXiv:2602.17711cs.SDeess.AS2026-02被引 2

剖析音频反伪造模型内部策略,发现误判根源。

Interpreting Multi-Branch Anti-Spoofing Architectures: Correlating Internal Strategy with Empirical Performance

  • 通过谱特征与树模型分析各分支贡献度,量化决策策略。
  • 识别出四种运行模式,其中错误专精导致攻击A17、A18性能暴跌。
  • 揭示传统指标忽略的结构依赖,提升模型可靠性评估能力。

多分支深度神经网络如AASIST3在语音反欺骗任务中达到顶尖水平,但其内部决策机制远不如输入级显著性方法透明。现有可解释性研究多聚焦于可视化输入伪影,而对各分支在不同欺骗攻击下的协作或竞争关系缺乏系统刻画。本文构建了针对AASIST3组件级别的解释框架:对十四个分支及全局注意力模块的中间激活使用协方差算子建模,其主特征值构成低维谱签名;这些签名训练CatBoost元分类器,生成TreeSHAP分支归因,并转化为归一化贡献份额与置信度得分(Cb),以量化模型运行策略。基于ASVspoof 2019基准中的13种欺骗攻击分析发现,存在四种运行原型——从有效专精(如A09,EER 0.04%,C=1.56)到无效共识(如A08,EER 3.14%,C=0.33)。关键发现为‘错误专精’模式:模型高置信地依赖错误分支,导致攻击A17和A18性能严重下降(EER分别为14.26%和28.63%)。该结果将内部架构策略与实证可靠性直接关联,揭示标准性能指标所忽视的特定结构依赖。

原文摘要 · Abstract (English)

Multi-branch deep neural networks like AASIST3 achieve state-of-the-art comparable performance in audio anti-spoofing, yet their internal decision dynamics remain opaque compared to traditional input-level saliency methods. While existing interpretability efforts largely focus on visualizing input artifacts, the way individual architectural branches cooperate or compete under different spoofing attacks is not well characterized. This paper develops a framework for interpreting AASIST3 at the component level. Intermediate activations from fourteen branches and global attention modules are modeled with covariance operators whose leading eigenvalues form low-dimensional spectral signatures. These signatures train a CatBoost meta-classifier to generate TreeSHAP-based branch attributions, which we convert into normalized contribution shares and confidence scores (Cb) to quantify the model's operational strategy. By analyzing 13 spoofing attacks from the ASVspoof 2019 benchmark, we identify four operational archetypes-ranging from Effective Specialization (e.g., A09, Equal Error Rate (EER) 0.04%, C=1.56) to Ineffective Consensus (e.g., A08, EER 3.14%, C=0.33). Crucially, our analysis exposes a Flawed Specialization mode where the model places high confidence in an incorrect branch, leading to severe performance degradation for attacks A17 and A18 (EER 14.26% and 28.63%, respectively). These quantitative findings link internal architectural strategy directly to empirical reliability, highlighting specific structural dependencies that standard performance metrics overlook.

反欺骗可解释性深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。