用可解释的线性预测检测语音伪造,速度快且结果可信
Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction

- 基于维纳-霍普夫预测构建可解释的检测框架
- 在多个数据集上表现媲美顶尖模型,计算量更低
- 适合需要透明决策过程的语音安全场景
合成语音生成技术的快速发展使语音伪造检测成为多媒体取证的关键挑战。现有方法虽具备高准确率,但多依赖黑箱架构,可解释性差且计算复杂度高。本文提出一种基于维纳-霍普夫线性预测的可解释性设计音频深度伪造检测框架,结合轻量级2D卷积神经网络进行处理。该设计实现了分类结果与信号声学特性之间的直接透明关联。在基准数据集上的实验表明,该方法在保持显著更低计算复杂度的同时,达到与当前最优方案相当的检测性能。通过Grad-CAM进行可解释性分析显示,分类器重点关注低阶预测系数以及静音和过渡区域,说明维纳-霍普夫预测能有效捕捉合成语音中的混响特征和细微统计异常。最后的鲁棒性实验表明,经微调后,模型在添加噪声、MP3压缩和电话滤波等常见后处理降质条件下仍能恢复检测性能。
原文摘要 · Abstract (English)
The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics. While recent approaches achieve high detection accuracy, they typically rely on black-box architectures that offer limited interpretability and high computational complexity. In this paper, we propose an explainable-by-design audio deepfake detection framework based on Wiener-Hopf linear prediction, processed by a lightweight 2D Convolutional Neural Network (CNN). This design enables a direct and transparent connection between classification outcomes and the acoustic properties of the signal. Experimental results on benchmark datasets demonstrate competitive detection performance while maintaining significantly lower computational complexity compared to state-of-the-art solutions. The interpretability analysis using Grad-CAM reveals that the classifier focuses on low-order predictor coefficients and on silence and transitional regions, suggesting that the Wiener-Hopf predictor captures reverberation characteristics and subtle statistical inconsistencies in synthetic speech. Finally, robustness experiments show that fine-tuning effectively recovers detection performance under common post-processing degradations, including additive noise, MP3 compression, and telephone filtering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。