arXiv:2409.17285cs.SDcs.AI2024-09被引 60

构建百万级真实场景语音伪造数据集,支持语音伪造检测与鲁棒说话人识别。

SpoofCeleb: Speech Deepfake Detection and SASV In The Wild

  • 基于真实世界数据自动构建语音伪造样本,实现自然环境下的深度伪造生成。
  • 包含250万条语音,覆盖1251名说话人,覆盖多种噪声和环境条件。
  • 开源完整数据集与基线模型,适合语音安全与对抗性研究者使用。

本文提出SpoofCeleb,一个用于语音深度伪造检测(SDD)和抗欺骗自动说话人验证(SASV)的公开数据集。该数据集基于真实世界录音构建,利用在相同真实数据上训练的文本转语音(TTS)系统生成语音伪造攻击。为提升模型鲁棒性,需在多样声学环境与不同噪声水平下进行训练,但现有数据集多为高质量、清洁录音,难以满足需求。为此,我们开发了全自动处理流程,对VoxCeleb1数据集进行转换,适配TTS训练要求。在此基础上,训练了23个主流TTS模型。SpoofCeleb包含超过250万条语音,来自1251位唯一说话人,均在自然真实条件下采集。数据集配备精心划分的训练、验证与测试集,以及标准化实验协议。文中报告了SDD与SASV任务的基线性能。所有数据、协议与基线已公开发布于https://jungjee.github.io/spoofceleb。

原文摘要 · Abstract (English)

This paper introduces SpoofCeleb, a dataset designed for Speech Deepfake Detection (SDD) and Spoofing-robust Automatic Speaker Verification (SASV), utilizing source data from real-world conditions and spoofing attacks generated by Text-To-Speech (TTS) systems also trained on the same real-world data. Robust recognition systems require speech data recorded in varied acoustic environments with different levels of noise to be trained. However, current datasets typically include clean, high-quality recordings (bona fide data) due to the requirements for TTS training; studio-quality or well-recorded read speech is typically necessary to train TTS models. Current SDD datasets also have limited usefulness for training SASV models due to insufficient speaker diversity. SpoofCeleb leverages a fully automated pipeline we developed that processes the VoxCeleb1 dataset, transforming it into a suitable form for TTS training. We subsequently train 23 contemporary TTS systems. SpoofCeleb comprises over 2.5 million utterances from 1,251 unique speakers, collected under natural, real-world conditions. The dataset includes carefully partitioned training, validation, and evaluation sets with well-controlled experimental protocols. We present the baseline results for both SDD and SASV tasks. All data, protocols, and baselines are publicly available at https://jungjee.github.io/spoofceleb.

语音伪造数据集说话人识别TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。