arXiv:2501.18919cs.SDcs.AI2025-01中稿 · ICASSP,2025被引 5

用Whisper编码检测歌声深度伪造,有效识别伪造唱段。

Deepfake Detection of Singing Voices With Whisper Encodings

  • 利用Whisper模型的噪声敏感编码提取声学特征
  • 在不同模型尺寸下实现低于5%的错误等率(%EER)
  • 适合音乐版权保护与音频真实性验证场景

歌声深度伪造对音乐产业中的艺术家构成严重威胁。本文提出一种歌声深度伪造检测(SVDD)系统,采用开源AI模型Whisper的噪声变异性编码作为特征表示。尽管Whisper本身具备抗噪能力,其编码仍富含非语音信息且对噪声敏感,因而适合作为SVDD任务的输入特征。本研究在人声与混音两种信号上进行检测,评估了不同Whisper模型规模及两类分类器(CNN与ResNet34)在多种测试条件下的表现,以%EER作为核心评价指标。

原文摘要 · Abstract (English)

The deepfake generation of singing vocals is a concerning issue for artists in the music industry. In this work, we propose a singing voice deepfake detection (SVDD) system, which uses noise-variant encodings of open-AI's Whisper model. As counter-intuitive as it may sound, even though the Whisper model is known to be noise-robust, the encodings are rich in non-speech information, and are noise-variant. This leads us to evaluate Whisper encodings as feature representations for the SVDD task. Therefore, in this work, the SVDD task is performed on vocals and mixtures, and the performance is evaluated in \%EER over varying Whisper model sizes and two classifiers- CNN and ResNet34, under different testing conditions.

深度伪造检测歌声识别Whisper音频安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。