arXiv:2512.24140cs.SD2025-12综述被引 8

首个大规模环境音深伪检测挑战赛,助力识别伪造音频。

Environmental Sound Deepfake Detection Challenge: An Overview

  • 构建首个大规模环境音深伪检测数据集EnvSDD。
  • 通过ICASSP 2026挑战赛验证检测方法有效性。
  • 适合音频安全、内容验证领域研究者参考。

近期音频生成模型的发展使得高度逼真且沉浸式的声景成为可能,广泛应用于影视与虚拟现实相关领域。然而,这些音频生成器也带来了潜在滥用风险,例如为伪造视频生成欺骗性音频或传播误导信息。因此,亟需发展有效的虚假环境声音检测方法。现有的环境声音深伪检测(ESDD)数据集在规模和声类多样性方面仍显不足。为填补这一空白,我们推出了首个面向ESDD的大规模人工标注数据集EnvSDD。基于该数据集,我们发起了被列为ICASSP 2026三大主挑战之一的ESDD Challenge。本文对本次挑战赛进行了全面概述,包括对参赛结果的详细分析。

原文摘要 · Abstract (English)

Recent progress in audio generation models has made it possible to create highly realistic and immersive soundscapes, which are now widely used in film and virtual-reality-related applications. However, these audio generators also raise concerns about potential misuse, such as producing deceptive audio for fabricated videos or spreading misleading information. Therefore, it is essential to develop effective methods for detecting fake environmental sounds. Existing datasets for environmental sound deepfake detection (ESDD) remain limited in both scale and the diversity of sound categories they cover. To address this gap, we introduced EnvSDD, the first large-scale curated dataset designed for ESDD. Based on EnvSDD, we launched the ESDD Challenge, recognized as one of the ICASSP 2026 Grand Challenges. This paper presents an overview of the ESDD Challenge, including a detailed analysis of the challenge results.

音频安全深伪检测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。