提出融合多层特征的音频伪造检测模型,提升对环境音伪造的识别能力。
BEAT2AASIST model with layer fusion for ESDD 2026 Challenge
- 通过分频/通道拆分BEATs特征,双分支AASIST处理增强表征
- 采用top-k Transformer层融合策略,提升特征丰富性
- 结合声码器数据增强,有效应对未知伪造手法
近年来音频生成技术的发展加剧了环境音伪造的风险,推动了首届大规模环境音深度伪造检测基准ESDD 2026挑战赛的设立。本文提出BEAT2AASIST模型,通过在频率或通道维度上拆分BEATs提取的特征,并由双AASIST分支分别处理。为丰富特征表示,引入top-k Transformer层融合机制,采用拼接、CNN门控和SE门控三种策略。此外,采用基于声码器的数据增强方法,提高模型对未见伪造手段的鲁棒性。在官方测试集上的实验结果表明,该方法在挑战赛各赛道中均取得具有竞争力的表现。
原文摘要 · Abstract (English)
Recent advances in audio generation have increased the risk of realistic environmental sound manipulation, motivating the ESDD 2026 Challenge as the first large-scale benchmark for Environmental Sound Deepfake Detection (ESDD). We propose BEAT2AASIST which extends BEATs-AASIST by splitting BEATs-derived representations along frequency or channel dimension and processing them with dual AASIST branches. To enrich feature representations, we incorporate top-k transformer layer fusion using concatenation, CNN-gated, and SE-gated strategies. In addition, vocoder-based data augmentation is applied to improve robustness against unseen spoofing methods. Experimental results on the official test sets demonstrate that the proposed approach achieves competitive performance across the challenge tracks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。