arXiv:2601.07958cs.SDcs.AI2026-01

LJ-Spoof构建了多样化的语音伪造数据集,助力精准识别伪造语音来源。

LJ-Spoof: A Generatively Varied Corpus for Audio Anti-Spoofing and Synthesis Source Tracing

  • 系统性地改变语音合成中的声调、声码器和生成参数,构建多样化数据
  • 包含超300万条语音,覆盖30种文本转语音模型与10种真实处理变体
  • 适合语音安全研究者用于训练和评估反伪造模型,提升识别精度

说话人特定的反伪造与合成源溯源是音频反伪造的核心挑战。进展受限于缺乏系统性变化模型架构、合成流程和生成参数的数据集。为此,我们提出LJ-Spoof,一个说话人特定、生成多样化的语料库,系统性地变化韵律、声码器、生成超参数、真实语音输入源、训练方式及神经后处理。该语料库涵盖1位说话人(含工作室级录音)、30种文本转语音家族、500个生成变异子集、10种真实神经处理变体,以及超过300万条语音。这种高密度变异设计支持鲁棒的说话人条件反伪造与细粒度合成源溯源。本研究还将该数据集定位为实用的训练资源与基准评估套件,用于反伪造与源追溯任务。

原文摘要 · Abstract (English)

Speaker-specific anti-spoofing and synthesis-source tracing are central challenges in audio anti-spoofing. Progress has been hampered by the lack of datasets that systematically vary model architectures, synthesis pipelines, and generative parameters. To address this gap, we introduce LJ-Spoof, a speaker-specific, generatively diverse corpus that systematically varies prosody, vocoders, generative hyperparameters, bona fide prompt sources, training regimes, and neural post-processing. The corpus spans one speakers-including studio-quality recordings-30 TTS families, 500 generatively variant subsets, 10 bona fide neural-processing variants, and more than 3 million utterances. This variation-dense design enables robust speaker-conditioned anti-spoofing and fine-grained synthesis-source tracing. We further position this dataset as both a practical reference training resource and a benchmark evaluation suite for anti-spoofing and source tracing.

语音伪造反伪造数据集合成溯源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。