构建真实社交平台深伪语音数据集,提升跨域检测能力
Fake Speech Wild: Detecting Deepfake Speech on Social Media Platform
- 构建包含254小时音视频的跨平台深伪语音数据集FSW
- 在多平台测试中实现平均3.54%等错误率,显著优于现有方法
- 通过数据增强提升模型对社交平台深伪语音的鲁棒性
语音生成技术的快速发展导致深伪语音在社交媒体上广泛传播。尽管现有深伪语音检测方法在公开数据集上表现良好,但在跨域场景下性能显著下降。为推进真实场景下的检测能力,我们提出Fake Speech Wild(FSW)数据集,包含来自四个不同媒体平台的254小时真实与深伪语音。基于该数据集,我们建立基准评估体系,使用公共数据集和先进的自监督学习(SSL)检测方法,在真实场景中评估当前检测模型表现。同时,研究数据增强策略对提升检测鲁棒性的效果。最终,通过扩充公共数据集并融合FSW训练集,显著提升了真实世界深伪语音检测性能,所有评测集上平均等错误率(EER)达3.54%。
原文摘要 · Abstract (English)
The rapid advancement of speech generation technology has led to the widespread proliferation of deepfake speech across social media platforms. While deepfake audio countermeasures (CMs) achieve promising results on public datasets, their performance degrades significantly in cross-domain scenarios. To advance CMs for real-world deepfake detection, we first propose the Fake Speech Wild (FSW) dataset, which includes 254 hours of real and deepfake audio from four different media platforms, focusing on social media. As CMs, we establish a benchmark using public datasets and advanced selfsupervised learning (SSL)-based CMs to evaluate current CMs in real-world scenarios. We also assess the effectiveness of data augmentation strategies in enhancing CM robustness for detecting deepfake speech on social media. Finally, by augmenting public datasets and incorporating the FSW training set, we significantly advanced real-world deepfake audio detection performance, achieving an average equal error rate (EER) of 3.54% across all evaluation sets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。