arXiv:2505.19644cs.SDcs.AI2025-05被引 11

构建首个系统化变体的伪造语音数据集,助力溯源与检测。

STOPA: A Database of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution

  • 系统性控制8种声学模型、6种声码器等生成参数。
  • 覆盖700万样本,涵盖13个不同合成器的多样变体。
  • 适合做语音伪造溯源、深度伪造检测与模型透明性研究。

深度伪造语音检测中的关键问题之一是源溯源——确定合成语句的来源。这可能涉及识别声学模型(AM)、声码器模型(VM)或其他生成特定参数。然而,进展受限于缺乏专门且系统编排的数据集。为此,我们提出STOPA,一个系统化变异且元数据丰富的深度伪造语音源溯源数据集,涵盖8种声学模型、6种声码器及多种参数设置,共700万样本,来自13个不同的合成器。与现有数据集相比,停用词较少、元数据稀疏不同,STOPA提供更广泛的生成因素系统控制,如声码器选择、声学模型或预训练权重,从而提升溯源可靠性。这种控制显著提高溯源准确率,有助于司法取证、深度伪造检测及生成模型透明性研究。

原文摘要 · Abstract (English)

A key research area in deepfake speech detection is source tracing - determining the origin of synthesised utterances. The approaches may involve identifying the acoustic model (AM), vocoder model (VM), or other generation-specific parameters. However, progress is limited by the lack of a dedicated, systematically curated dataset. To address this, we introduce STOPA, a systematically varied and metadata-rich dataset for deepfake speech source tracing, covering 8 AMs, 6 VMs, and diverse parameter settings across 700k samples from 13 distinct synthesisers. Unlike existing datasets, which often feature limited variation or sparse metadata, STOPA provides a systematically controlled framework covering a broader range of generative factors, such as the choice of the vocoder model, acoustic model, or pretrained weights, ensuring higher attribution reliability. This control improves attribution accuracy, aiding forensic analysis, deepfake detection, and generative model transparency.

语音伪造源溯源数据集深度伪造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。