arXiv:2409.08731cs.SDeess.AS2024-09中稿 · IEEE SLT 2024被引 35

构建首个基于扩散与流匹配的语音伪造数据集,揭示现有反伪造模型脆弱性。

DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset

  • 基于扩散与流匹配模型生成高保真语音伪造样本
  • 实测显示当前反伪造模型对这类语音攻击泛化能力差
  • 为研发更鲁棒的反伪造技术提供关键数据资源

主流零样本文本转语音系统如 Voicebox 和 Seed-TTS 分别利用流匹配和扩散模型实现了接近人类水平的语音合成。然而,高质量语音合成也带来了身份滥用与信息安全风险。尽管已有诸多反伪造模型被提出,但现有顶尖反伪造模型在应对基于扩散与流匹配的语音合成系统所产生的语音伪造时的效能仍不明确。本文提出了扩散与流匹配基语音伪造(DFADD)数据集,收集了基于先进扩散与流匹配文本转语音模型生成的语音伪造样本。此外,我们发现当前反伪造模型在面对由扩散与流匹配文本转语音系统生成的高保真人声时缺乏足够鲁棒性。所提出的 DFADD 数据集填补了这一空白,为开发更稳健的反伪造模型提供了宝贵资源。

原文摘要 · Abstract (English)

Mainstream zero-shot TTS production systems like Voicebox and Seed-TTS achieve human parity speech by leveraging Flow-matching and Diffusion models, respectively. Unfortunately, human-level audio synthesis leads to identity misuse and information security issues. Currently, many antispoofing models have been developed against deepfake audio. However, the efficacy of current state-of-the-art anti-spoofing models in countering audio synthesized by diffusion and flowmatching based TTS systems remains unknown. In this paper, we proposed the Diffusion and Flow-matching based Audio Deepfake (DFADD) dataset. The DFADD dataset collected the deepfake audio based on advanced diffusion and flowmatching TTS models. Additionally, we reveal that current anti-spoofing models lack sufficient robustness against highly human-like audio generated by diffusion and flow-matching TTS systems. The proposed DFADD dataset addresses this gap and provides a valuable resource for developing more resilient anti-spoofing models.

语音伪造反伪造扩散模型流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。