arXiv:2508.01034eess.AS2025-08中稿 · APSIPA ASC 2025被引 2

融合声调谱与自监督语音特征,提升假语音检测的跨域泛化能力。

Fusion of Modulation Spectrogram and SSL with Multi-head Attention for Fake Speech Detection

  • 用自监督嵌入和声调谱融合生成新语音表征
  • 在多个数据集上相对基线提升37%以上
  • 适合需要跨语种、跨场景检测假语音的研究者

假语音检测系统已成为应对语音深度伪造的重要手段。现有系统因训练数据多样性不足,在跨领域语音样本上泛化能力差。本文提出一种结合自监督(SSL)语音嵌入与声调谱(MS)特征的新语音表征方法,并设计融合策略作为分类任务的前端。该融合表征输入AASIST后端网络。在单语和多语假语音数据集上进行实验,评估模型在跨数据集和多语言场景下的性能。所提模型在ASVspoof 2019和MLAAD数据集上的域内设置中分别取得37%和20%的相对性能提升。在域外场景下,于ASVspoof 2019上训练的模型在MLAAD上评估时实现36%的相对改进。所有测试语言中,模型均持续优于基线,表明其具备更强的域泛化能力。

原文摘要 · Abstract (English)

Fake speech detection systems have become a necessity to combat against speech deepfakes. Current systems exhibit poor generalizability on out-of-domain speech samples due to lack to diverse training data. In this paper, we attempt to address domain generalization issue by proposing a novel speech representation using self-supervised (SSL) speech embeddings and the Modulation Spectrogram (MS) feature. A fusion strategy is used to combine both speech representations to introduce a new front-end for the classification task. The proposed SSL+MS fusion representation is passed to the AASIST back-end network. Experiments are conducted on monolingual and multilingual fake speech datasets to evaluate the efficacy of the proposed model architecture in cross-dataset and multilingual cases. The proposed model achieves a relative performance improvement of 37% and 20% on the ASVspoof 2019 and MLAAD datasets, respectively, in in-domain settings compared to the baseline. In the out-of-domain scenario, the model trained on ASVspoof 2019 shows a 36% relative improvement when evaluated on the MLAAD dataset. Across all evaluated languages, the proposed model consistently outperforms the baseline, indicating enhanced domain generalization.

假语音检测自监督学习声调谱跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。