用低秩专家自适应融合,提升音频伪造检测在真实干扰下的鲁棒性。
Adaptive Mixture of Low-Rank Experts for Robust Audio Spoofing Detection
- 引入攻击特异性专家与低秩微调,仅用1.13%参数实现高效适配。
- 自适应融合策略在混合攻击下准确率显著优于传统方法。
- 适合需要应对复杂现实后处理攻击的语音安全场景使用。
在音频伪造检测中,多数研究依赖干净数据集,导致模型易受真实世界后处理攻击(如信道压缩、噪声)影响。为此,本文提出自适应低秩专家混合框架(AMULET),通过攻击特异性专家(ASEs)与低秩微调(LoRA)结合,使每个专家仅需全量微调1.13%的参数即可聚焦特定后处理模式。进一步设计自适应专家融合(AEF)机制,动态选择并整合专家知识,增强检测鲁棒性。实验表明,相比全量微调模型,AMULET在噪声环境中的鲁棒性显著提升,且对未见后处理方法具有更强适应能力。在多种混合攻击下,其性能超越单专家及其它专家聚合策略,展现出优异的综合鲁棒性与适应性。
原文摘要 · Abstract (English)
In audio spoofing detection, most studies rely on clean datasets, making models susceptible to real-world post-processing attacks, such as channel compression and noise. To overcome this challenge, we propose the Adaptive MixtUre Low-rank ExperTs (AMULET) framework, which enhances resilience by leveraging attack-specific knowledge and dynamically adapting to varied attack conditions. Specifically, AMULET employs Attack-Specific Experts (ASEs) fine-tuned with Low-Rank Adaptation (LoRA), allowing each expert to focus on distinct post-processing patterns using just 1.13\% of the parameters required for full fine-tuning. Furthermore, we introduce Adaptive Expert Fusion (AEF), which adaptively selects and integrates expert knowledge to enhance the robustness of spoofing detection. Experimental results demonstrate that AMULET significantly enhances robustness by improving noise resilience and exhibiting greater adaptability to unseen post-processing methods compared to models trained with full fine-tuning. Additionally, our framework outperforms both single expert and other expert aggregation strategies under various mixed attacks, demonstrating its superior robustness and adaptability in managing complex real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。