构建首个融合对抗攻击的语音伪造检测数据集,支持真实场景下攻防研究。
ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech
- 通过众包收集2000+说话人、多样声学环境数据,生成32类语音伪造攻击。
- 首次纳入对抗攻击,涵盖传统与现代语音合成/变声模型混合生成方式。
- 提供7个独立分区及辅助数据,适合语音安全、反欺诈领域研究人员使用。
ASVspoof 5是该系列挑战的第五届,旨在推动语音伪造与深度伪造攻击的研究及检测方法设计。本文介绍全新构建的ASVspoof 5数据库,其数据通过众包方式在多样化声学环境下采集,覆盖约2000名说话人(较早期版本的100人显著增加)。数据库包含32种由众包生成的攻击算法,部分已使用新型替代检测模型进行优化。攻击类型涵盖传统与现代文本转语音及语音转换模型的混合生成,首次引入对抗攻击。协议设计包含七个说话人互斥的数据分区:两组用于训练不同攻击模型,两组用于替代检测模型的开发与评估,另三组构成标准训练、开发与评估集。额外提供来自3万+说话人的辅助数据集,可用于训练攻击算法所需的说话人编码器。实验验证了新数据库的有效性,使用自动说话人验证与伪造/深度伪造基线检测器进行测试。除攻击生成协议与工具外,其余资源已于2024年挑战中开放,现向社区免费提供。
原文摘要 · Abstract (English)
ASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake attacks as well as the design of detection solutions. We introduce the ASVspoof 5 database which is generated in a crowdsourced fashion from data collected in diverse acoustic conditions (cf. studio-quality data for earlier ASVspoof databases) and from ~2,000 speakers (cf. ~100 earlier). The database contains attacks generated with 32 different algorithms, also crowdsourced, and optimised to varying degrees using new surrogate detection models. Among them are attacks generated with a mix of legacy and contemporary text-to-speech synthesis and voice conversion models, in addition to adversarial attacks which are incorporated for the first time. ASVspoof 5 protocols comprise seven speaker-disjoint partitions. They include two distinct partitions for the training of different sets of attack models, two more for the development and evaluation of surrogate detection models, and then three additional partitions which comprise the ASVspoof 5 training, development and evaluation sets. An auxiliary set of data collected from an additional 30k speakers can also be used to train speaker encoders for the implementation of attack algorithms. Also described herein is an experimental validation of the new ASVspoof 5 database using a set of automatic speaker verification and spoof/deepfake baseline detectors. With the exception of protocols and tools for the generation of spoofed/deepfake speech, the resources described in this paper, already used by participants of the ASVspoof 5 challenge in 2024, are now all freely available to the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。