统一检测各类音频伪造,融合双域特征提升识别准确率。
Beyond Speech: Dual-Domain SSL Fusion for Unified All-Type Audio Deepfake Detection

- 用双域自监督模型融合声学与事件级特征,生成统一判真伪空间。
- 在AT-ADD Track 2上达95.58%宏平均F1,排名第二。
- 适合需跨类型音频伪造检测的场景,如语音、音乐、环境音等。
统一全类型音频伪造检测旨在判断输入音频片段是否为伪造,其类型可能为语音、环境音、歌唱声或音乐。现有以语音为中心或依赖类型的方法在此场景下不足,因测试时音频类型未知,但输出仍需单一二分类结果。为此,本文提出一种双域自监督学习融合方法,将异构音频映射至共享二分类真实性空间。采用EAT-large和wav2vec 2.0 XLS-R-300M作为互补的自监督特征源,分别提供广泛的声学与事件级表示,以及波形级、人声及语音敏感表示。逐层加权融合整合不同变压器深度的多层级伪影特征,而令牌级融合则形成统一特征池,无需强制两路自监督流间的帧级对齐。融合后的令牌通过多头注意力统计池化,并由二分类MLP头进行分类。在统一核心检测器之上应用保守语音精炼后,所提交系统在AT-ADD Track 2评估集上取得95.58%的宏平均F1,位列挑战赛第二名。
原文摘要 · Abstract (English)
Unified all-type audio deepfake detection aims to determine whether an input clip is real or fake when its audio type may be speech, environmental sound, singing voice, or music. Existing speech-centric or type-dependent solutions are insufficient for this setting because the test-time audio type is unknown, while the required output is still a single binary decision. To address these issues, this paper proposes a dual-domain SSL fusion method that maps heterogeneous audio into a shared binary authenticity space. EAT-large and wav2vec 2.0 XLS-R-300M are used as complementary SSL feature sources, providing broad acoustic and event-level representations as well as waveform-level, vocal, and speech-sensitive representations. Layer-wise weighted fusion integrates multi-level artifacts from different transformer depths, while token-level fusion forms a unified feature pool without enforcing frame-level alignment between the two SSL streams. The fused tokens are summarized by multi-head attentive statistics pooling and classified with a binary MLP head. With conservative speech refinement applied on top of this unified core detector, the submitted system achieves 95.58% Macro-F1 on the AT-ADD Track 2 evaluation set and ranks second in the challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。