arXiv:2505.14188cs.SDeess.AS2025-05中稿 · INTERSPEECH 2025被引 13

判断语音伪造是否来自同一生成模型,助力声纹溯源与取证。

Source Verification for Speech Deepfakes

  • 借鉴说话人验证思路,用嵌入向量距离判断音频是否同源。
  • 在多语言、跨说话人、经后处理场景下仍保持较高识别准确率。
  • 首次提出源验证任务,为语音伪造司法鉴定提供新方法。

随着语音深度伪造生成器的泛滥,不仅需要评估合成音频的真实性,还需追溯其来源。尽管源归属模型试图解决此问题,但在面对未见过的生成器时表现不佳。本文提出源验证任务,受说话人验证启发,判断测试音频是否与一组参考信号来自同一生成模型。方法利用源归属分类器训练出的嵌入向量,计算音频间的距离得分以判断是否同源。我们在多种场景下评估多个模型,分析了说话人多样性、语言不匹配及后处理操作的影响。本工作首次探索源验证任务,揭示其潜力与脆弱性,为实际司法取证提供重要参考。

原文摘要 · Abstract (English)

With the proliferation of speech deepfake generators, it becomes crucial not only to assess the authenticity of synthetic audio but also to trace its origin. While source attribution models attempt to address this challenge, they often struggle in open-set conditions against unseen generators. In this paper, we introduce the source verification task, which, inspired by speaker verification, determines whether a test track was produced using the same model as a set of reference signals. Our approach leverages embeddings from a classifier trained for source attribution, computing distance scores between tracks to assess whether they originate from the same source. We evaluate multiple models across diverse scenarios, analyzing the impact of speaker diversity, language mismatch, and post-processing operations. This work provides the first exploration of source verification, highlighting its potential and vulnerabilities, and offers insights for real-world forensic applications.

语音伪造源验证深度伪造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。