通过帧级监督后训练,提升语音大模型对深度伪造的鲁棒检测能力
Supervised Post-training of Speech Foundation Models for Robust Adaptation in Speech Deepfake Detection

- 在预训练模型基础上引入局部扰动与帧级标注,引导模型关注伪造特征
- ASVspoof5上达到4.50%的EER,无需数据增强即达顶尖水平
- 在不同伪造类型间表现均衡,适合实际部署的鲁棒检测场景
大型语音基础模型在语音深度伪造检测中展现出强大潜力,但直接微调受限于自监督预训练目标与伪造特定伪影之间的不匹配。为此,我们提出一种混合帧后训练策略,生成局部化的伪造导向扰动,并采用帧级监督促使自监督模型学习关键的局部不一致性。在ASVspoof5上,单模型未使用数据增强即实现4.50%的EER,达到当前最优水平;在ASVspoof2021 LA/DF上,LA与DF之间的绝对EER差距仅为0.16%,表明模型在不同失真条件下具有强且均衡的鲁棒性。结果表明,有监督后训练为语音基础模型的鲁棒深度伪造检测提供了有效且实用的适配路径。
原文摘要 · Abstract (English)
Large speech foundation models have shown strong potential for speech deepfake detection, but direct fine-tuning is limited by a mismatch between self-supervised pre-training objectives and spoof-specific artifacts. To address this, we propose a mix-frame post-training strategy to create localized spoof-oriented perturbations and use frame-level supervision to encourage the SSL model to learn local inconsistencies that are critical for robust spoof detection. On ASVspoof5, we achieve state-of-the-art EER 4.50% for a single model without data augmentation. On ASVspoof2021 LA/DF, it further achieves only 0.16\% absolute EER gap between LA and DF, indicating strong and balanced robustness across distinct distortion conditions. These results show that supervised post-training provides an effective and practical way to adapt speech foundation models for robust deepfake detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。