arXiv:2603.04865cs.SD2026-03中稿 · Interspeech 2026被引 3

首个环境音深伪检测挑战赛,揭示音频伪造风险与防御策略

The First Environmental Sound Deepfake Detection Challenge: Benchmarking Robustness, Evaluation, and Insights

  • 构建首个环境音深伪检测数据集与评测框架
  • 97支队伍参与,1748份提交,验证了检测系统的鲁棒性边界
  • 揭示主流模型设计规律,为安全音频技术提供方向

近年来,音频生成技术的进展使得高保真的环境声音场景易于生成,可能被用于制造虚假警报、枪声和人群声等欺骗性内容,引发公共安全与信任危机。尽管语音和歌唱音深伪检测已广泛研究,环境音深伪检测(ESDD)仍处于起步阶段。为此,首届ESDD挑战赛启动,吸引97支队伍参与,收到1748份有效提交。本文介绍了任务设定、数据集构建、评估协议、基线系统及挑战结果的关键洞察。进一步分析了高性能系统中的常见架构选择与训练策略。最后,讨论了未来研究方向,明确了该领域关键机遇与开放问题。

原文摘要 · Abstract (English)

Recent progress in audio generation has made it increasingly easy to create highly realistic environmental soundscapes, which can be misused to produce deceptive content, such as fake alarms, gunshots, and crowd sounds, raising concerns for public safety and trust. While deepfake detection for speech and singing voice has been extensively studied, environmental sound deepfake detection (ESDD) remains underexplored. To advance ESDD, the first edition of the ESDD challenge was launched, attracting 97 registered teams and receiving 1,748 valid submissions. This paper presents the task formulation, dataset construction, evaluation protocols, baseline systems, and key insights from the challenge results. Furthermore, we analyze common architectural choices and training strategies among top-performing systems. Finally, we discuss potential future research directions for ESDD, outlining key opportunities and open problems to guide subsequent studies in this field.

音频安全深伪检测环境音挑战赛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。