arXiv:2607.08586eess.AScs.SD2026-07

通过音素级分析,解释语音伪造检测模型为何判断一段语音为伪造。

Why Do You Say It Like That? A Phoneme-Level Framework for Explainable Speech Deepfake Detection

论文配图:Why Do You Say It Like That? A Phoneme-Level Framework for Explainable Speech Deepfake Detection
图 1 · 摘自论文原文
  • 基于梯度加权激活映射与语音识别,生成对齐音素的显著性图。
  • 在ASVspoof 5数据集上检测性能与现有方法相当。
  • 揭示了攻击方式和说话人相关的可解释语音伪造线索。

随着自监督表示(如wav2vec 2.0和HuBERT)的应用,语音深伪检测的准确率不断提升,但理解模型为何将语音判定为真实或伪造仍是开放挑战。为提升人工智能的可信度与可解释性,本文提出一种音素级分析框架,将模型预测与可测量的语音单位关联。该后处理可解释方法适用于基于卷积神经网络的多种语音深伪检测系统,结合梯度加权类激活映射与语音识别,生成与音素和停顿对齐的显著性图。该流程揭示了统计上显著的、与攻击类型和说话人相关的语音伪造线索,且以人类可理解的方式呈现。在ASVspoof 5数据集上的实验表明,该方法检测性能与同类架构相当,同时提供跨说话人和伪造条件的语言学解释。

原文摘要 · Abstract (English)

As the accuracy of speech deepfake detection improves with the use of self-supervised representations such as wav2vec 2.0 and HuBERT, understanding why the speech is classified as bona fide or deepfake remains an open challenge. In pursuit of more trustworthy and interpretable artificial intelligence, we introduce a phoneme-level analysis framework that connects model predictions to measurable phonetic units. Our post-hoc explainability method is generally applicable to a variety of speech deepfake detection systems based on convolutional neural networks since it leverages Gradient-weighted Class Activation Mapping in conjunction with speech recognition to generate saliency maps aligned with phonemes and pauses. This pipeline reveals statistically significant attack- and speaker-dependent phonetic cues associated with spoofed speech in terms that humans can understand. Experiments using ASVspoof 5 show comparable detection performance to similar architectures while providing linguistic interpretations across speakers and spoofing conditions.

语音伪造可解释性音素分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。