arXiv:2508.04529cs.SD2025-08被引 13

构建首个大规模环境音深伪检测数据集,推动真实场景下音频伪造识别研究。

ESDD 2026: Environmental Sound Deepfake Detection Challenge Evaluation Plan

  • 构建45.25小时真实音与316.7小时伪造音的大型数据集EnvSDD。
  • 设立未见生成器和低资源黑盒两个挑战赛道,覆盖真实应用场景。
  • 面向音频安全、虚假信息防控领域研究者,助力对抗深度伪造技术。

近年来,音频生成系统的发展使高度逼真且沉浸式的音景得以实现,广泛应用于影视与虚拟现实领域。然而,这些生成技术也带来滥用风险,如用于伪造视频的欺骗性音频内容及误导性信息传播。现有环境音深伪检测(ESDD)数据集在规模和音频类型上均显不足。为此,我们提出了首个专为ESDD设计的大规模精心标注数据集EnvSDD,包含45.25小时真实音频与316.7小时伪造音频。基于该数据集,我们将启动环境音深伪检测挑战赛。具体设置两个赛道:未见生成器下的ESDD与黑盒低资源ESDD,涵盖实际应用中的多种挑战。比赛将于2026年IEEE国际声学、语音与信号处理会议(ICASSP 2026)期间举行。

原文摘要 · Abstract (English)

Recent advances in audio generation systems have enabled the creation of highly realistic and immersive soundscapes, which are increasingly used in film and virtual reality. However, these audio generators also raise concerns about potential misuse, such as generating deceptive audio content for fake videos and spreading misleading information. Existing datasets for environmental sound deepfake detection (ESDD) are limited in scale and audio types. To address this gap, we have proposed EnvSDD, the first large-scale curated dataset designed for ESDD, consisting of 45.25 hours of real and 316.7 hours of fake sound. Based on EnvSDD, we are launching the Environmental Sound Deepfake Detection Challenge. Specifically, we present two different tracks: ESDD in Unseen Generators and Black-Box Low-Resource ESDD, covering various challenges encountered in real-life scenarios. The challenge will be held in conjunction with the 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026).

音频伪造数据集深度伪造声学安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。