arXiv:2411.17349cs.SDeess.AS2024-11被引 8

对比ASR模型在语音深度伪造检测中的表现,发现更强的语音识别能力未必带来更好的伪造检测效果。

Comparative Analysis of ASR Methods for Speech Deepfake Detection

  • 用Whisper和Wav2Vec 2.0作为特征提取器,适配语音伪造检测任务
  • 多版本模型测试显示,ASR性能提升不等于伪造检测能力同步增强
  • 为构建高效伪造检测系统提供选型参考,尤其适合关注模型迁移机制的研究者

近年来,语音深度伪造检测常依赖预训练自监督模型。这些模型最初用于自动语音识别(ASR),已被证明能有效提取语音信号的有意义表征,对多种任务(包括深度伪造检测)有益。在此框架下,预训练模型作为特征提取器,从输入语音中提取嵌入表示,再输入二分类检测器。该方法取得显著准确率,暗示了ASR与语音深度伪造检测之间可能存在关联。然而,这种联系尚未明确:是否ASR性能提升必然带来更高的伪造检测能力?本文通过系统分析回答此问题。我们选取两种不同预训练自监督ASR模型——Whisper与Wav2Vec 2.0,采用其多个版本(参数量递增、ASR性能逐步提升),将其适配至语音深度伪造检测任务。实验探究了ASR性能改进是否伴随伪造检测性能提升。结果揭示了两任务间的关系,并为开发更有效的语音深度伪造检测系统提供了重要指导。

原文摘要 · Abstract (English)

Recent techniques for speech deepfake detection often rely on pre-trained self-supervised models. These systems, initially developed for Automatic Speech Recognition (ASR), have proved their ability to offer a meaningful representation of speech signals, which can benefit various tasks, including deepfake detection. In this context, pre-trained models serve as feature extractors and are used to extract embeddings from input speech, which are then fed to a binary speech deepfake detector. The remarkable accuracy achieved through this approach underscores a potential relationship between ASR and speech deepfake detection. However, this connection is not yet entirely clear, and we do not know whether improved performance in ASR corresponds to higher speech deepfake detection capabilities. In this paper, we address this question through a systematic analysis. We consider two different pre-trained self-supervised ASR models, Whisper and Wav2Vec 2.0, and adapt them for the speech deepfake detection task. These models have been released in multiple versions, with increasing number of parameters and enhanced ASR performance. We investigate whether performance improvements in ASR correlate with improvements in speech deepfake detection. Our results provide insights into the relationship between these two tasks and offer valuable guidance for the development of more effective speech deepfake detectors.

语音伪造ASR模型特征提取深度伪造检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。