arXiv:2509.13878eess.AScs.LG2025-09被引 4

用多个低秩专家模型提升语音伪造检测的泛化能力。

Mixture of Low-Rank Adapter Experts in Generalizable Audio Deepfake Detection

  • 引入多专家低秩适配器,动态选择专用模块应对新攻击
  • 在跨域场景下将平均错误率从8.55%降至6.08%
  • 适合关注对抗性攻击防御与模型鲁棒性的研究者

以Wav2Vec2为代表的基座模型在语音任务中表现优异,但在固定语料上微调后难以泛化到未见的伪造方法。为此,本文提出混合低秩适配器专家(MoE-LoRA)方法,将多个低秩适配器嵌入模型注意力层,并通过路由机制选择性激活特定专家,增强对新型伪造攻击的适应能力。实验表明,该方法在域内和域外场景下均优于标准微调,显著降低等错误率(EER)。最佳模型将平均域外EER从8.55%降至6.08%,验证了其在通用语音伪造检测中的有效性。

原文摘要 · Abstract (English)

Foundation models such as Wav2Vec2 excel at representation learning in speech tasks, including audio deepfake detection. However, after being fine-tuned on a fixed set of bonafide and spoofed audio clips, they often fail to generalize to novel deepfake methods not represented in training. To address this, we propose a mixture-of-LoRA-experts approach that integrates multiple low-rank adapters (LoRA) into the model's attention layers. A routing mechanism selectively activates specialized experts, enhancing adaptability to evolving deepfake attacks. Experimental results show that our method outperforms standard fine-tuning in both in-domain and out-of-domain scenarios, reducing equal error rates relative to baseline models. Notably, our best MoE-LoRA model lowers the average out-of-domain EER from 8.55\% to 6.08\%, demonstrating its effectiveness in achieving generalizable audio deepfake detection.

音频伪造泛化检测LoRA专家混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。