arXiv:2409.11909cs.SDeess.AS2024-09被引 37

用专家混合模型融合多层特征,无须微调即可高效检测假音频。

Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0

  • 基于门控网络的专家混合机制,自动选择关键层特征
  • 在ASVspoof2019/2021上达到与微调方法相当的检测精度
  • 全程冻结wav2vec 2.0,适合快速应对新型合成语音

语音合成技术对说话人验证系统构成了严重威胁。当前最有效的假音频检测方法依赖预训练模型,并通过融合模型各层特征进一步提升性能。然而,大多数已有融合方法需要微调预训练模型,导致训练时间过长,难以快速迭代以应对新型语音合成技术。为此,本文提出一种基于专家混合(Mixture of Experts)的特征融合方法,利用最后一层特征作为门控信号,从多层特征中提取并融合与假音频检测相关的信息,同时保持预训练模型完全冻结。在ASVspoof2019和ASVspoof2021数据集上的实验表明,该方法在无需微调的情况下,性能与需微调的方法相当。

原文摘要 · Abstract (English)

Speech synthesis technology has posed a serious threat to speaker verification systems. Currently, the most effective fake audio detection methods utilize pretrained models, and integrating features from various layers of pretrained model further enhances detection performance. However, most of the previously proposed fusion methods require fine-tuning the pretrained models, resulting in excessively long training times and hindering model iteration when facing new speech synthesis technology. To address this issue, this paper proposes a feature fusion method based on the Mixture of Experts, which extracts and integrates features relevant to fake audio detection from layer features, guided by a gating network based on the last layer feature, while freezing the pretrained model. Experiments conducted on the ASVspoof2019 and ASVspoof2021 datasets demonstrate that the proposed method achieves competitive performance compared to those requiring fine-tuning.

假音频检测专家混合wav2vec 2.0冻结模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。