arXiv:2608.00493cs.SD2026-08中稿 · ACMMM 2026被引 1

先识别音频类型再检测真假,提升各类音频伪造检测效果。

Hidden-Domain Routing for All-Type Audio Deepfake Detection

  • 通过6秒音频判断类型,动态选择对应专家模型
  • 在语音、环境音、人声和音乐上分别达到88%~99%准确率
  • 适合需要统一检测多种音频的反伪造系统使用

全类型音频深度伪造检测需在语音、环境音、人声和音乐等不同音频类型上做出真伪判断,而推理时无法获知音频类型。在AT-ADD Track2中,这形成隐藏音频域条件:真实/伪造标签跨域共享,但表示结构与检测器得分行为随音频类型变化。本文提出闭合条件路由系统,先恢复隐藏音频域,再在选定分支内解释检测得分。AudioType-BEATs-6s路由模块从6秒窗口估计音频类型;语音由Speech-XLSR专家处理,环境音、人声和音乐则依赖基于EAT的通用音频专家,并采用分支本地得分解释策略。开发集表示分析、路由族对比及组件结果表明,音频域可有效分离,且不同类型检测器具有互补优势。在官方AT-ADD Track2最终评估中,系统取得96.10%的宏平均F1,位列榜首,各类型宏平均F1分别为:语音88.07%、环境音98.18%、人声99.07%、音乐99.08%。结果验证了在全类型音频伪造检测中,先恢复音频域再解释得分的有效性。

原文摘要 · Abstract (English)

All-type audio deepfake detection requires authenticity decisions across speech, environmental sound, singing voice, and music, while the audio type is unavailable at inference time. In AT-ADD Track2, this setting creates a hidden audio-domain condition: the binary real/fake label is shared across domains, but representation structure and detector-score behavior vary with audio type. We present a closed-condition routed system that first recovers the hidden audio domain and then interprets detector scores within the selected branch. The AudioType-BEATs-6s Router estimates audio type from a 6-second window; speech inputs are handled by the Speech-XLSR Expert, while sound, singing, and music rely on EAT-based general-audio experts with branch-local score interpretation. Development-set representation analysis, router-family comparisons, and component results show audio-domain separation and complementary detector strengths across audio types. On the official AT-ADD Track2 final evaluation, the system achieves 96.10% Track2 Macro-F1 and ranks first on the final leaderboard, with type-wise Macro-F1 scores of 88.07%, 98.18%, 99.07%, and 99.08% for speech, sound, singing, and music, respectively. These results support recovering the hidden audio domain before interpreting detector scores in all-type audio deepfake detection.

音频伪造多类型检测路由机制BEATs

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。