arXiv:2512.08319eess.AS2025-12被引 3

用自监督模型融合检测未知生成器的环境音深伪,效果优异。

BUT Systems for Environmental Sound Deepfake Detection in the ESDD 2026 Challenge

  • 采用多种自监督模型+轻量注意力后端,提升特征判别力。
  • 单系统在开发集、评测集上误报率低至0.00%和4.80%。
  • 通过分布不确定性增强,有效应对未知频谱失真问题。

本文介绍布根大学(BUT)在ESDD 2026挑战赛中的提交方案,聚焦于轨道1:针对未知生成器的环境音深伪检测。为应对模型泛化至未见合成算法音频的挑战,提出一种基于多样自监督学习(SSL)模型的鲁棒集成框架。我们全面分析了通用音频SSL模型(如BEATs、EAT、Dasheng)及语音专用SSL模型,并将其与轻量级多头因子化注意力(MHFA)后端结合,以捕捉判别性表征。此外,引入基于分布不确定性建模的特征域增强策略,提升模型对未知频谱失真的鲁棒性。所有模型均仅使用官方EnvSDD数据训练,未依赖任何外部资源。实验结果表明该方法有效性:最优单系统在开发集、进展集(轨道1)、最终评估集上的等错误率(EER)分别为0.00%、4.60%和4.80%;融合系统进一步提升泛化能力,对应EER降至0.00%、3.52%和4.38%。

原文摘要 · Abstract (English)

This paper describes the BUT submission to the ESDD 2026 Challenge, specifically focusing on Track 1: Environmental Sound Deepfake Detection with Unseen Generators. To address the critical challenge of generalizing to audio generated by unseen synthesis algorithms, we propose a robust ensemble framework leveraging diverse Self-Supervised Learning (SSL) models. We conduct a comprehensive analysis of general audio SSL models (including BEATs, EAT, and Dasheng) and speech-specific SSLs. These front-ends are coupled with a lightweight Multi-Head Factorized Attention (MHFA) back-end to capture discriminative representations. Furthermore, we introduce a feature domain augmentation strategy based on distribution uncertainty modeling to enhance model robustness against unseen spectral distortions. All models are trained exclusively on the official EnvSDD data, without using any external resources. Experimental results demonstrate the effectiveness of our approach: our best single system achieved Equal Error Rates (EER) of 0.00\%, 4.60\%, and 4.80\% on the Development, Progress (Track 1), and Final Evaluation sets, respectively. The fusion system further improved generalization, yielding EERs of 0.00\%, 3.52\%, and 4.38\% across the same partitions.

深伪检测自监督学习音频安全模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。