构建首个针对扩散模型语音伪造的专用数据集,解决现有检测器失效问题。
DiffSSD: A Diffusion-Based Dataset For Speech Forensics
- 构建包含200小时语音的DiffSSD数据集,涵盖8个开源和2个商用扩散生成器。
- 实测显示现有检测器在新生成语音上准确率不足50%,暴露检测盲区。
- 适合语音安全、反伪造研究者使用,推动检测技术适配新型生成工具。
基于扩散的语音生成器已广泛应用,可生成高质量合成语音,近期多起事件揭示其被恶意利用的风险。为应对这一挑战,已有合成语音检测方法被提出,但多数训练数据未包含扩散类生成器。本文表明,基于ASVspoof2019数据集训练的现有检测器,在识别最新扩散模型生成的语音时表现不佳。为此,我们提出了扩散型合成语音数据集(DiffSSD),包含约200小时标注语音,涵盖8个开源与2个商业扩散生成器。我们进一步在封闭集与开放集场景下评估了现有检测器在DiffSSD上的性能。结果凸显该数据集对检测现代生成器产出语音的重要性。
原文摘要 · Abstract (English)
Diffusion-based speech generators are ubiquitous. These methods can generate very high quality synthetic speech and several recent incidents report their malicious use. To counter such misuse, synthetic speech detectors have been developed. Many of these detectors are trained on datasets which do not include diffusion-based synthesizers. In this paper, we demonstrate that existing detectors trained on one such dataset, ASVspoof2019, do not perform well in detecting synthetic speech from recent diffusion-based synthesizers. We propose the Diffusion-Based Synthetic Speech Dataset (DiffSSD), a dataset consisting of about 200 hours of labeled speech, including synthetic speech generated by 8 diffusion-based open-source and 2 commercial generators. We also examine the performance of existing synthetic speech detectors on DiffSSD in both closed-set and open-set scenarios. The results highlight the importance of this dataset in detecting synthetic speech generated from recent open-source and commercial speech generators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。