首个大规模环境音深伪检测数据集,提升真实场景下音频伪造识别能力
EnvSDD: Benchmarking Environmental Sound Deepfake Detection
- 基于预训练音频基础模型构建检测系统
- 在45.25小时真实与316.74小时伪造音视频上验证效果
- 支持未见生成模型和数据集,适合实际部署场景
音频生成系统如今能创造极为逼真的声景,虽可助力媒体制作,但也带来潜在风险。现有研究多聚焦于语音或歌唱声的深伪检测,但环境音具有不同特征,使原有方法在真实声音场景中效果下降。此外,当前环境音深伪检测数据集规模小、类型有限。为此,我们提出EnvSDD——首个专为该任务设计的大规模标注数据集,包含45.25小时真实音频与316.74小时伪造音频。测试集涵盖多种未见生成模型和数据集,用于评估模型泛化能力。我们还基于预训练音频基础模型提出一套深伪检测系统。在EnvSDD上的实验表明,该系统性能优于语音与歌唱领域当前最佳方法。
原文摘要 · Abstract (English)
Audio generation systems now create very realistic soundscapes that can enhance media production, but also pose potential risks. Several studies have examined deepfakes in speech or singing voice. However, environmental sounds have different characteristics, which may make methods for detecting speech and singing deepfakes less effective for real-world sounds. In addition, existing datasets for environmental sound deepfake detection are limited in scale and audio types. To address this gap, we introduce EnvSDD, the first large-scale curated dataset designed for this task, consisting of 45.25 hours of real and 316.74 hours of fake audio. The test set includes diverse conditions to evaluate the generalizability, such as unseen generation models and unseen datasets. We also propose an audio deepfake detection system, based on a pre-trained audio foundation model. Results on EnvSDD show that our proposed system outperforms the state-of-the-art systems from speech and singing domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。