面向各类合成音频的检测挑战,推动更通用的反伪造技术发展。
AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan

- 设计双赛道评估体系,覆盖语音与非语音类合成音频
- 引入真实场景干扰和新型生成方法,检验检测器鲁棒性
- 适合音频安全、媒体验证与内容治理领域的研究者
音频大模型的快速发展使语音与非语音音频(如音效、人声演唱、音乐)的低成本、高保真生成与篡改成为可能。尽管这促进了内容创作,但也带来严重的安全与信任问题。现有音频深度伪造检测方法多集中于语音,依赖特定声学痕迹,对真实世界失真不鲁棒,且难以泛化到异构音频类型和新兴欺骗手段。为此,我们提出为ACM Multimedia 2026设计的全类型音频深度伪造检测(AT-ADD)挑战赛,包含两个赛道:(1)鲁棒语音伪造检测,评估检测器在真实场景下对未知先进语音生成方法的应对能力;(2)全类型音频伪造检测,将检测范围扩展至语音、音效、人声演唱与音乐等多样未知音频类型,推动类型无关的泛化能力。通过提供标准化数据集、严格评估协议和可复现基线,AT-ADD旨在加速鲁棒、通用音频取证技术的发展,支持安全通信、可靠媒体验证与负责任的治理。
原文摘要 · Abstract (English)
The rapid advancement of Audio Large Language Models (ALLMs) has enabled cost-effective, high-fidelity generation and manipulation of both speech and non-speech audio, including sound effects, singing voices, and music. While these capabilities foster creativity and content production, they also introduce significant security and trust challenges, as realistic audio deepfakes can now be generated and disseminated at scale. Existing audio deepfake detection (ADD) countermeasures (CMs) and benchmarks, however, remain largely speech-centric, often relying on speech-specific artifacts and exhibiting limited robustness to real-world distortions, as well as restricted generalization to heterogeneous audio types and emerging spoofing techniques. To address these gaps, we propose the All-Type Audio Deepfake Detection (AT-ADD) Grand Challenge for ACM Multimedia 2026, designed to bridge controlled academic evaluation with practical multimedia forensics. AT-ADD comprises two tracks: (1) Robust Speech Deepfake Detection, which evaluates detectors under real-world scenarios and against unseen, state-of-the-art speech generation methods; and (2) All-Type Audio Deepfake Detection, which extends detection beyond speech to diverse, unknown audio types and promotes type-agnostic generalization across speech, sound, singing, and music. By providing standardized datasets, rigorous evaluation protocols, and reproducible baselines, AT-ADD aims to accelerate the development of robust and generalizable audio forensic technologies, supporting secure communication, reliable media verification, and responsible governance in an era of pervasive synthetic audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。