arXiv:2505.12994cs.SDeess.AS2025-05中稿 · Interspeech 2025被引 10

通过音频编码器分类法,追踪深度伪造语音的生成模型来源。

Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy

  • 基于神经音频编码器的结构差异,建立分类框架识别生成模型
  • 在CodecFake+数据集上实现初步溯源,准确率验证可行性
  • 适合语音安全、反欺诈领域研究人员参考

基于神经音频编码器的语音生成(CoSG)模型近期实现了高度逼真的音频深度伪造。我们将由CoSG系统生成的深度伪造语音称为编码器型深度伪造(CodecFake)。尽管现有反欺骗研究多聚焦于验证音频真实性,但几乎未关注生成这些伪造语音所用的CoSG模型溯源问题。在CodecFake生成过程中,语音到单元编码、离散单元建模及单元到语音解码等步骤均源于神经音频编码器。受此启发,我们提出通过神经音频编码器分类法实现CodecFake源追溯,该方法分解神经音频编码器以定位生成模型。在CodecFake+数据集上的实验结果表明,该方法具备初步可行性,同时揭示了若干有待深入研究的挑战。

原文摘要 · Abstract (English)

Recent advances in neural audio codec-based speech generation (CoSG) models have produced remarkably realistic audio deepfakes. We refer to deepfake speech generated by CoSG systems as codec-based deepfake, or CodecFake. Although existing anti-spoofing research on CodecFake predominantly focuses on verifying the authenticity of audio samples, almost no attention was given to tracing the CoSG used in generating these deepfakes. In CodecFake generation, processes such as speech-to-unit encoding, discrete unit modeling, and unit-to-speech decoding are fundamentally based on neural audio codecs. Motivated by this, we introduce source tracing for CodecFake via neural audio codec taxonomy, which dissects neural audio codecs to trace CoSG. Our experimental results on the CodecFake+ dataset provide promising initial evidence for the feasibility of CodecFake source tracing while also highlighting several challenges that warrant further investigation.

深度伪造语音溯源音频编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。