arXiv:2608.23437cs.SD2026-08

构建首个覆盖全类型音频的深度伪造检测基准,提升真实场景下的检测鲁棒性。

AT-ADD: A Benchmark and Challenge for Robust and All-Type Audio Deepfake Detection

论文配图:AT-ADD: A Benchmark and Challenge for Robust and All-Type Audio Deepfake Detection
图 1 · 摘自论文原文
  • 设计双赛道:语音检测与跨类型泛化检测,覆盖多种生成器和录制环境。
  • 挑战赛最优系统在两类任务上分别达90.71%和96.10%宏观F1,显著超越基线。
  • 揭示自监督表示与多裁剪推理是泛化关键,跨类型一致性仍是难题。

近期音频生成模型可合成高保真语音、环境音、人声演唱及音乐,带来多媒体信任新风险。现有音频深度伪造检测(ADD)基准多聚焦语音,且难以反映真实信道变化与多样音频类型。本文提出AT-ADD,一个大规模基准与挑战,用于评估鲁棒语音检测与全类型音频伪造检测能力。Track 1评估在未知生成器、多样录制条件、信号扰动及回放效应下的二分类语音检测性能;Track 2评估在语音、声音、演唱、音乐等类型未知时的通用真假检测能力。本文详述数据集构建、评估协议与可复现基线,并分析提交至ACM Multimedia 2026 Grand Challenge的最终系统。最强官方基线在Track 1和Track 2上分别取得76.73%和79.47%宏观F1,而优胜系统分别达到90.71%和96.10%。除总体排名外,对前五名提交系统的样本级分析揭示了生成器与类型层面的难度差异、跨系统错误互补性及排名稳定性。结果表明,大规模自监督表征、条件感知增强、多裁剪推理及结构化融合或路由是泛化核心,但生成器特异性鲁棒性与跨类型一致性仍待解决。

原文摘要 · Abstract (English)

Recent audio generation models can synthesize high-fidelity speech, environmental sound, singing voice, and music, creating new risks for multimedia trust. Existing audio deepfake detection (ADD) benchmarks remain predominantly speech-centric and often underrepresent realistic channel variation and diverse audio types. This paper presents AT-ADD, a large-scale benchmark and challenge designed to evaluate both robust speech deepfake detection and all-type audio deepfake detection. Track 1 evaluates binary speech detection under unseen generators, diverse recording conditions, signal perturbations, and replay effects. Track 2 evaluates type-agnostic real/fake detection over speech, sound, singing, and music when the audio type is unknown at test time. We detail the dataset construction, evaluation protocol, and reproducible baselines, and analyze the final systems submitted to the ACM Multimedia 2026 Grand Challenge. The strongest official baseline obtains 76.73% and 79.47% Macro-F1 on the Track 1 and Track 2 evaluation sets, respectively, whereas the winning challenge systems reach 90.71% and 96.10%. Beyond aggregate rankings, sample-level analysis of the top five submissions examines generator- and type-level difficulty, cross-system error complementarity, and ranking stability. The results show that large-scale self-supervised representations, condition-aware augmentation, multi-crop inference, and structured fusion or routing are central to generalization, while generator-specific robustness and consistent performance across diverse audio types remain unresolved.

音频伪造检测基准自监督多类型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。