构建45分钟长音频数据集,助力伪造语音检测与声纹识别研究
Descriptor:: Extended-Length Audio Dataset for Synthetic Voice Detection and Speaker Recognition (ELAD-SVDSR)
- 采集36人45分钟长音频,五种麦克风录制,覆盖多样语音特征
- 生成20个深度伪造语音,提升检测模型训练挑战性
- 适合语音安全、声纹认证和音频取证方向的研究者使用
本文提出用于合成语音检测与说话人识别的长时音频数据集ELAD-SVDSR,包含36名参与者在受控环境下朗读报纸文章的45分钟音频,每段录音由五种不同质量的麦克风采集。该数据集聚焦于长时语音,捕捉更丰富的语调、音高轮廓与表达细节,使生成的合成语音更具真实感与连贯性。基于此,已创建20个深度伪造语音并加入数据集,用以构建更具挑战性的训练与评估样本。数据集附带匿名化说话人人口统计信息,旨在推动音频取证、生物特征安全与语音认证技术的发展。
原文摘要 · Abstract (English)
This paper introduces the Extended Length Audio Dataset for Synthetic Voice Detection and Speaker Recognition (ELAD SVDSR), a resource specifically designed to facilitate the creation of high quality deepfakes and support the development of detection systems trained against them. The dataset comprises 45 minute audio recordings from 36 participants, each reading various newspaper articles recorded under controlled conditions and captured via five microphones of differing quality. By focusing on extended duration audio, ELAD SVDSR captures a richer range of speech attributes such as pitch contours, intonation patterns, and nuanced delivery enabling models to generate more realistic and coherent synthetic voices. In turn, this approach allows for the creation of robust deepfakes that can serve as challenging examples in datasets used to train and evaluate synthetic voice detection methods. As part of this effort, 20 deepfake voices have already been created and added to the dataset to showcase its potential. Anonymized metadata accompanies the dataset on speaker demographics. ELAD SVDSR is expected to spur significant advancements in audio forensics, biometric security, and voice authentication systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。