arXiv:2505.13930cs.SDeess.AS2025-05中稿 · terspeech 2025被引 15

用双向Mamba捕捉语音伪造的细微特征,检测更准。

BiCrossMamba-ST: Speech Deepfake Detection with Bidirectional Mamba Spectro-Temporal Cross-Attention

  • 双分支结构分别处理频谱和时序信息,再通过交叉注意力融合
  • 在ASVSpoof数据集上比顶尖方法提升67.74%(LA21)和6.80%(DF21)
  • 直接用原始特征,适合需要高精度语音伪造检测的场景

我们提出BiCrossMamba-ST,一种基于双向Mamba块与互交叉注意力的双分支频谱-时序架构,用于鲁棒的语音深度伪造检测。该框架分别处理频谱子带与时间区间,再融合其表示,有效捕捉合成语音的细微线索。同时,引入基于卷积的2D注意力图,聚焦关键频谱-时序区域,提升检测能力。直接作用于原始特征,性能显著优于现有方法:在ASVSpoof LA21和DF21基准上,相对AASIST分别提升67.74%和26.3%,在DF21上比RawBMamba提升6.80%。代码与模型将公开。

原文摘要 · Abstract (English)

We propose BiCrossMamba-ST, a robust framework for speech deepfake detection that leverages a dual-branch spectro-temporal architecture powered by bidirectional Mamba blocks and mutual cross-attention. By processing spectral sub-bands and temporal intervals separately and then integrating their representations, BiCrossMamba-ST effectively captures the subtle cues of synthetic speech. In addition, our proposed framework leverages a convolution-based 2D attention map to focus on specific spectro-temporal regions, enabling robust deepfake detection. Operating directly on raw features, BiCrossMamba-ST achieves significant performance improvements, a 67.74% and 26.3% relative gain over state-of-the-art AASIST on ASVSpoof LA21 and ASVSpoof DF21 benchmarks, respectively, and a 6.80% improvement over RawBMamba on ASVSpoof DF21. Code and models will be made publicly available.

语音伪造Mamba注意力机制深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。