arXiv:2508.04143eess.AScs.CL2025-08中稿 · Interspeech SPSC 2…被引 11

首个多语言语音深度伪造溯源基准,可识别生成模型来源。

Multilingual Source Tracing of Speech Deepfakes: A First Benchmark

  • 构建多语言语音伪造溯源基准,覆盖同语种与跨语种场景。
  • 发现跨语言迁移时,微调语言影响模型泛化性能。
  • 适合语音安全、反伪造研究者参考,开源数据代码可用。

生成式AI的进展使得仅需几秒音频即可生成逼真的语音深度伪造。尽管此类工具具有实用价值,但也引发严重风险,使多种语言的逼真假语音频成为可能。现有研究主要聚焦于检测假语音,却很少关注溯源生成模型。本文首次提出多语言语音深度伪造源追溯基准,涵盖同语种与跨语种场景。我们比较了基于DSP与自监督学习(SSL)的建模方法;探究不同语言微调的SSL表征对跨语言泛化的影响;并评估模型在未见语言和说话人上的泛化能力。研究结果提供了关于训练与推理语言不一致时识别生成模型挑战的首个全面洞察。数据集、评测协议与代码已开源:https://github.com/xuanxixi/Multilingual-Source-Tracing。

原文摘要 · Abstract (English)

Recent progress in generative AI has made it increasingly easy to create natural-sounding deepfake speech from just a few seconds of audio. While these tools support helpful applications, they also raise serious concerns by making it possible to generate convincing fake speech in many languages. Current research has largely focused on detecting fake speech, but little attention has been given to tracing the source models used to generate it. This paper introduces the first benchmark for multilingual speech deepfake source tracing, covering both mono- and cross-lingual scenarios. We comparatively investigate DSP- and SSL-based modeling; examine how SSL representations fine-tuned on different languages impact cross-lingual generalization performance; and evaluate generalization to unseen languages and speakers. Our findings offer the first comprehensive insights into the challenges of identifying speech generation models when training and inference languages differ. The dataset, protocol and code are available at https://github.com/xuanxixi/Multilingual-Source-Tracing.

语音伪造源溯源多语言自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。