arXiv:2608.14249cs.SD2026-08中稿 · ACM MM 2026

针对各类音频伪造的检测挑战,提出通用检测框架并验证其效果。

AT-ADD: All-Type Audio Deepfake Detection Challenge Summary

论文配图:AT-ADD: All-Type Audio Deepfake Detection Challenge Summary
图 1 · 摘自论文原文
  • 结合自监督表征与多裁剪推理,实现跨类型音频伪造检测。
  • 最佳系统在真实场景下达到90.71%宏F1,通用检测达96.10%。
  • 适合关注音频安全、伪造检测与模型泛化性的研究者。

本文总结了ACM多媒体2026年全类型音频深度伪造检测(AT-ADD)大挑战。该挑战包含两个赛道:在真实声学与信道变化下的鲁棒语音伪造检测,以及涵盖语音、环境音、演唱声和音乐的无类型依赖检测。我们介绍了任务设计、数据集与评测集构建、官方排行榜结果,以及参赛系统常见的设计模式。最佳Track 1系统在最终评测集上取得90.71%宏F1,最佳Track 2系统达96.10%宏F1。最终提交方案普遍采用大规模自监督音频表征、数据增强、多裁剪推理及结构化融合或路由策略。结果也揭示了对未见生成器的泛化能力、对真实语音域失真的鲁棒性,以及跨异构音频类型的性能均衡等仍存挑战。

原文摘要 · Abstract (English)

This paper summarizes the ACM Multimedia 2026 AT-ADD Grand Challenge on all-type audio deepfake detection. AT-ADD contains two tracks: robust speech deepfake detection under realistic acoustic and channel variations, and type-agnostic detection over speech, environmental sound, singing voice, and music. We describe the challenge tasks, dataset and evaluation-set design, official leaderboard results, and common design patterns observed in participating systems. The best Track 1 system achieved 90.71% Macro-F1 on the final evaluation set, while the best Track 2 system achieved 96.10% Macro-F1. The final submissions show that strong systems commonly combine large-scale self-supervised audio representations, data augmentation, multi-crop inference, and structured fusion or routing. The results also reveal remaining challenges in generalization to unseen generators, robustness to realistic speech-domain distortions, and balanced performance across heterogeneous audio types.

音频伪造深度伪造检测自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。