arXiv:2606.10791cs.SD2026-06中稿 · 2026 ICME workshop综述被引 4

评测环境感知语音与声音伪造检测系统,揭示高效设计策略。

Overview of ESDD2: Environment-Aware Speech and Sound Deepfake Detection Challenge

论文配图:Overview of ESDD2: Environment-Aware Speech and Sound Deepfake Detection Challenge
图 1 · 摘自论文原文
  • 采用模块化分解与自监督编码器提升检测能力
  • 最优模型在测试集上宏F1达0.8775,显著优于基线
  • 适合关注音频伪造检测与鲁棒性研究者参考

2026年ICME会议同期举办的环境感知语音与声音深度伪造检测挑战赛(ESDD2),评估了五类音频欺骗检测任务,其中语音与环境音可独立或联合被篡改。挑战赛共吸引来自16个国家的94个注册团队,经验证后保留13支队伍参与最终分析。测试集上最佳系统取得0.8775的宏F1分数,远超分离增强联合学习基线(0.6327)。顶尖方案普遍采用模块化任务拆解、跨域自监督编码器、针对性数据增强及选择性集成,而非单纯扩大模型规模。辅助EER分析显示,对篡改环境音的检测仍存困难,且在未见生成器上的泛化能力不足。本文总结挑战结果并为未来环境感知深度伪造检测研究提供洞见。CompSpoofV2数据集与基线代码已公开以保障可复现性。

原文摘要 · Abstract (English)

The Environment-Aware Speech and Sound Deepfake Detection Challenge (ESDD2), held in conjunction with ICME 2026, evaluated systems for five component-level audio spoofing detection, where speech and environmental sounds may be manipulated independently or jointly. After the challenge concludes, we analyze the final leaderboard and summarize effective design choices from the top-performing submissions. The challenge attracted 94 registrations from 16 countries; after verification of submission requirements and metadata, 13 teams were retained for the final analysis. On the test set, the best system achieved a Macro-F1 score of 0.8775, substantially outperforming the separation-enhanced joint learning baseline (0.6327). Top systems consistently benefited from modular task decomposition, cross-domain self-supervised encoders, targeted data augmentation, and selective ensembling rather than simple model scaling. At the same time, auxiliary EER analyses reveal persistent difficulty in detecting the spoofed environmental component and in generalizing to unseen generators in the test set. This paper reports challenge results and provides insights for future environment-aware deepfake detection research. The CompSpoofV2 dataset and baseline code remain publicly available for reproducibility.

音频伪造深度伪造检测挑战自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。