揭秘深度伪造语音检测器真正依赖的音频线索
What Do Deepfake Speech Detectors Actually Hear?
- 用时间对齐的自监督表征+积分梯度定位判断依据
- 三款检测器分别依赖环境声、音素异常、词边界等不同线索
- 通过遮蔽实验验证线索有效性,结果可信
深度伪造语音检测器常仅输出单一评分,却不说明判断依据、证据位置或驱动因素。本文提出一种基于时间对齐自监督表征的音频原生可解释性方法,采用积分梯度技术定位决策证据。将该方法应用于三个基于WavLM的检测器(AASIST、CA-MHFA、SLS)在ASVspoof 5上的表现,并人工标注高贡献区域以赋予语义。尽管性能相近,三者依赖的线索各异:AASIST关注非语音/环境特征,CA-MHFA聚焦局部音素异常,SLS则依赖词边界与频谱完整性。通过因果遮蔽主线索验证发现性能显著下降,进一步支持所揭示的检测语义。
原文摘要 · Abstract (English)
Deepfake speech detectors often output a single score without explaining why an audio sample is flagged, where in the signal the evidence lies, or what cues drive the decision. We propose an audio-native explainability pipeline using Integrated Gradients on time-aligned self-supervised representations to localize decision evidence over time. We apply the proposed method to three WavLM-based detectors (AASIST, CA-MHFA, SLS) on ASVspoof 5 and manually annotate the highest-attribution regions to provide a semantic meaning of the most important cues. Despite similar performance, the detectors rely on different cues: AASIST emphasizes non-speech/environment cues, CA-MHFA focuses on localized phoneme artifacts, and SLS relies on word boundaries and spectral integrity. We move beyond speculative reasoning and validate our findings by causal masking of the primary detector cues. Observed performance degradation further supports the explained detector semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。