提出噪声感知的音频深度伪造检测框架,提升真实场景下识别效果。
Toward Noise-Aware Audio Deepfake Detection: Survey, SNR-Benchmarks, and Practical Recipes
- 构建可控信噪比的测试基准,模拟真实噪声环境。
- 在10-0 dB信噪比下,微调使误报率降低10-15个百分点。
- 适合关注真实场景中伪造音频检测的研究者与开发者。
尽管基于预训练编码器(如WavLM、Wav2Vec2、MMS)的音频深度伪造检测技术迅速发展,但在真实采集条件下(如家庭/办公室/交通背景噪声、混响、消费级通道)的性能仍显著低于实验室纯净条件的结果。本文调研并评估了当前主流音频深度伪造检测模型在噪声环境下的鲁棒性,提出一个可复现的框架:将MS-SNSD噪声与ASVspoof 2021 DF语音混合,实现受控信噪比(SNR)下的评估。信噪比作为噪声强度的度量指标,覆盖从接近纯净(35 dB)到极低信噪比(-5 dB)的范围,用于量化模型性能的渐进退化。研究了多条件训练与固定信噪比测试对预训练编码器(WavLM、Wav2Vec2、MMS)的影响,报告了二分类和四分类(真实性 × 噪声类型)任务下的准确率、ROC-AUC与等错误率(EER)。实验表明,在10–0 dB SNR范围内,微调可使各骨干模型的等错误率降低10–15个百分点。
原文摘要 · Abstract (English)
Deepfake audio detection has progressed rapidly with strong pre-trained encoders (e.g., WavLM, Wav2Vec2, MMS). However, performance in realistic capture conditions - background noise (domestic/office/transport), room reverberation, and consumer channels - often lags clean-lab results. We survey and evaluate robustness for state-of-the-art audio deepfake detection models and present a reproducible framework that mixes MS-SNSD noises with ASVspoof 2021 DF utterances to evaluate under controlled signal-to-noise ratios (SNRs). SNR is a measured proxy for noise severity used widely in speech; it lets us sweep from near-clean (35 dB) to very noisy (-5 dB) to quantify graceful degradation. We study multi-condition training and fixed-SNR testing for pretrained encoders (WavLM, Wav2Vec2, MMS), reporting accuracy, ROC-AUC, and EER on binary and four-class (authenticity x corruption) tasks. In our experiments, finetuning reduces EER by 10-15 percentage points at 10-0 dB SNR across backbones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。