提出跨注意力模型,提升复杂场景下的环境音深伪检测能力。
Technical Report of Nomi Team in the Environmental Sound Deepfake Detection Challenge 2026
- 设计音频-文本交叉注意力机制,融合多模态信息增强判别力。
- 在未见生成器和低资源黑盒场景下,误报率降低至基线以下。
- 适合关注音频安全与对抗样本防御的研究者参考。
本文介绍我们参与 ICASSP 2026 环境音深伪检测(ESDD)挑战赛的工作。挑战基于大规模 EnvSDD 数据集,包含多种合成环境声音。针对未见生成器和低资源黑盒场景的复杂性,我们提出一种音频-文本交叉注意力模型。通过单独及联合使用文本-音频模型的实验,结果表明该方法在等错误率(EER)上优于挑战基线(BEATs + AASIST 模型),展现出良好泛化能力与鲁棒性。
原文摘要 · Abstract (English)
This paper presents our work for the ICASSP 2026 Environmental Sound Deepfake Detection (ESDD) Challenge. The challenge is based on the large-scale EnvSDD dataset that consists of various synthetic environmental sounds. We focus on addressing the complexities of unseen generators and low-resource black-box scenarios by proposing an audio-text cross-attention model. Experiments with individual and combined text-audio models demonstrate competitive EER improvements over the challenge baseline (BEATs+AASIST model).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。