arXiv:2510.02864cs.SD2025-10中稿 · @ ACM IH&MMSec 202…被引 1

判断语音伪造音频是否来自同一生成模型。

Forensic Similarity for Speech Deepfakes

  • 构建双分支网络,比对音频的伪造特征相似性。
  • 在未知模型上仍能准确识别来源一致性。
  • 适用于语音篡改检测,适合数字取证研究者。

本文提出语音伪造检测领域的「司法相似性」概念,旨在判断两个音频片段是否具有相同的伪造痕迹。受图像领域启发,我们设计了一个两阶段深度学习框架:基于孪生网络的特征提取器与核心决策模块,统称为相似性网络。该系统通过比较音频的司法特征,评估其是否源自同一生成源。实验表明,该方法在源验证任务中可有效判断两个语音样本是否由相同模型生成,并在音频拼接检测中展现良好扩展性。结果证明,该方法对未见过的伪造痕迹具备强泛化能力,体现了其鲁棒性、灵活性和实际应用价值。

原文摘要 · Abstract (English)

In this paper, we introduce the concept of forensic similarity in the speech deepfake detection domain, which aims to determine whether two audio segments share the same underlying forensic traces. Our approach is inspired by prior work in the image domain. To transfer this idea to the audio domain, we propose a two-stage deep learning framework consisting of a Siamese-based feature extractor and a core decision module, referred to as the similarity network. The system goal to assess whether two speech samples originate from the same source by comparing their forensic characteristics. In practice, the model maps pairs of audio segments to a similarity score indicating whether they contain identical or different forensic traces. We evaluate the proposed method on the emerging task of source verification, demonstrating its ability to determine whether two speech samples were generated by the same model. In addition, we explore its applicability to audio splicing detection as a complementary use case. Experimental results show that the proposed approach generalizes well to previously unseen forensic traces, highlighting its robustness, flexibility, and practical relevance for digital audio forensics.

语音伪造司法鉴定相似性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。