arXiv:2604.13400eess.AS2026-04中稿 · Oral Presentation …被引 2

用传统机器学习在双采样率下建立可解释的假音频检测基线

Classical Machine Learning Baselines for Deepfake Audio Detection on the Fake-or-Real Dataset

  • 提取音高、语音质量与频谱特征,构建可解释的检测模型
  • 最优RBF-SVM在双采样率下达93%准确率,误报率仅7%
  • 揭示音高变化与频谱丰富度是关键判别特征,适合对比研究

深度学习使合成语音高度逼真,引发欺诈、冒充和虚假信息担忧。尽管神经网络检测器进展迅速,但仍需透明的基准方法来揭示哪些声学特征能可靠区分真实与合成语音。本文基于Fake-or-Real(FoR)数据集,提出一种可解释的经典机器学习基线方法。从2秒音频片段中,分别在44.1 kHz(高保真)和16 kHz(电话质量)采样率下提取韵律、语音质量与频谱特征。通过ANOVA与相关热图分析,识别出在真实与合成语音间显著不同的特征。随后训练逻辑回归、LDA、QDA、高斯朴素贝叶斯、SVM与GMM等多类分类器,并以准确率、ROC-AUC、EER及DET曲线评估性能。成对McNemar检验确认模型间差异显著。最优模型为RBF-SVM,双采样率下测试准确率达约93%,误报率(EER)约7%;线性模型准确率约为75%。特征分析表明,音高变异性与频谱丰富度(谱中心、带宽)是关键判别线索。该结果为未来深度伪造音频检测器提供了强而可解释的基准。

原文摘要 · Abstract (English)

Deep learning has enabled highly realistic synthetic speech, raising concerns about fraud, impersonation, and disinformation. Despite rapid progress in neural detectors, transparent baselines are needed to reveal which acoustic cues reliably separate real from synthetic speech. This paper presents an interpretable classical machine learning baseline for deepfake audio detection using the Fake-or-Real (FoR) dataset. We extract prosodic, voice-quality, and spectral features from two-second clips at 44.1 kHz (high-fidelity) and 16 kHz (telephone-quality) sampling rates. Statistical analysis (ANOVA, correlation heatmaps) identifies features that differ significantly between real and fake speech. We then train multiple classifiers -- Logistic Regression, LDA, QDA, Gaussian Naive Bayes, SVMs, and GMMs -- and evaluate performance using accuracy, ROC-AUC, EER, and DET curves. Pairwise McNemar's tests confirm statistically significant differences between models. The best model, an RBF SVM, achieves ~93% test accuracy and ~7% EER on both sampling rates, while linear models reach ~75% accuracy. Feature analysis reveals that pitch variability and spectral richness (spectral centroid, bandwidth) are key discriminative cues. These results provide a strong, interpretable baseline for future deepfake audio detectors.

音频伪造机器学习可解释性检测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。