arXiv:2601.02944eess.AS2026-01ACL被引 1

混合Mamba与注意力机制,提升语音伪造检测精度与稳定性。

XLSR-MamBo: Scaling the Hybrid Mamba-Attention Backbone for Audio Deepfake Detection

论文配图:XLSR-MamBo: Scaling the Hybrid Mamba-Attention Backbone for Audio Deepfake Detection
图 1 · 摘自论文原文
  • 用XLSR前端+混合Mamba-注意力后端捕捉语音伪造特征。
  • 在ASVspoof 2021和DFADD数据集上达到领先性能。
  • 深度扩展有效缓解浅层模型的不稳定性,适合安全场景应用。

先进语音合成技术已能生成高度逼真的语音,带来安全风险,推动了语音伪造检测(ADD)的研究。尽管状态空间模型(SSMs)具有线性复杂度优势,但纯因果SSM架构常难以捕捉全局频域伪影所需的内容检索能力。为此,我们提出XLSR-MamBo,一个模块化框架,结合XLSR前端与协同的Mamba-Attention后端,系统评估四种拓扑结构:Mamba、Mamba2、Hydra和Gated DeltaNet。实验表明,MamBo-3-Hydra-N3配置在ASVspoof 2021 LA、DF和In-the-Wild基准上表现优于或媲美现有最优系统。该性能得益于Hydra原生双向建模能力,比先前启发式双分支策略更高效地捕获整体时序依赖。此外,在DFADD数据集上的评估显示其对未见的扩散与流匹配合成方法具有良好泛化能力。关键分析发现,增加后端深度可有效缓解浅层模型中观察到的性能波动与不稳定性。结果证明,该混合框架能有效捕捉伪造语音信号中的伪影,为ADD提供高效解决方案。

原文摘要 · Abstract (English)

Advanced speech synthesis technologies have enabled highly realistic speech generation, posing security risks that motivate research into audio deepfake detection (ADD). While state space models (SSMs) offer linear complexity, pure causal SSMs architectures often struggle with the content-based retrieval required to capture global frequency-domain artifacts. To address this, we explore the scaling properties of hybrid architectures by proposing XLSR-MamBo, a modular framework integrating an XLSR front-end with synergistic Mamba-Attention backbones. We systematically evaluate four topological designs using advanced SSM variants, Mamba, Mamba2, Hydra, and Gated DeltaNet. Experimental results demonstrate that the MamBo-3-Hydra-N3 configuration achieves competitive performance compared to other state-of-the-art systems on the ASVspoof 2021 LA, DF, and In-the-Wild benchmarks. This performance benefits from Hydra's native bidirectional modeling, which captures holistic temporal dependencies more efficiently than the heuristic dual-branch strategies employed in prior works. Furthermore, evaluations on the DFADD dataset demonstrate robust generalization to unseen diffusion- and flow-matching-based synthesis methods. Crucially, our analysis reveals that scaling backbone depth effectively mitigates the performance variance and instability observed in shallower models. These results demonstrate the hybrid framework's ability to capture artifacts in spoofed speech signals, providing an effective method for ADD.

语音伪造检测混合架构Mamba音频安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。