多语言数据融合提升语音伪造检测鲁棒性,获SAFE挑战赛双项第二。
Multilingual Dataset Integration Strategies for Robust Audio Deepfake Detection: A SAFE Challenge System
- 用WavLM大模型作前端,结合RawBoost增强,多语言数据训练
- 在未处理和伪装音频检测中均达第二,跨语言泛化能力强
- 适合做语音安全、反伪造系统研发的工程师参考
SAFE挑战赛评估三种场景下的合成语音检测:原始音频、含压缩伪影的音频,以及专为逃避检测设计的伪装音频。本文系统探索自监督学习(SSL)前端、训练数据构成及音频长度配置对鲁棒性的影响。基于AASIST的方法采用WavLM large前端与RawBoost增强,在涵盖9种语言、超过70个TTS系统的多语言数据集(256,600样本)上训练,数据来源包括CodecFake、MLAAD v5、SpoofCeleb、Famous Figures和MAILABS。通过对比不同SSL前端、三种数据版本及两种音频长度的实验,该方法在任务1(原始音频检测)和任务3(伪装音频检测)中均取得第二名,验证了其强泛化能力和鲁棒性。
原文摘要 · Abstract (English)
The SAFE Challenge evaluates synthetic speech detection across three tasks: unmodified audio, processed audio with compression artifacts, and laundered audio designed to evade detection. We systematically explore self-supervised learning (SSL) front-ends, training data compositions, and audio length configurations for robust deepfake detection. Our AASIST-based approach incorporates WavLM large frontend with RawBoost augmentation, trained on a multilingual dataset of 256,600 samples spanning 9 languages and over 70 TTS systems from CodecFake, MLAAD v5, SpoofCeleb, Famous Figures, and MAILABS. Through extensive experimentation with different SSL front-ends, three training data versions, and two audio lengths, we achieved second place in both Task 1 (unmodified audio detection) and Task 3 (laundered audio detection), demonstrating strong generalization and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。