arXiv:2410.06796cs.CRcs.AI2024-10被引 14

构建扩散模型生成的深度伪造语音数据集,评估其检测难度与真实度。

Diffuse or Confuse: A Diffusion Deepfake Speech Dataset

  • 用预训练模型构建扩散生成语音数据集
  • 扩散生成语音与传统方法质量相当,检测难度相似
  • 适合研究者测试新型深度伪造检测算法

人工智能与机器学习的进步显著提升了合成语音的生成能力。本文探索了扩散模型这一新型合成语音方法,利用现有工具和预训练模型构建了一个扩散语音数据集。同时,本研究评估了扩散生成的深度伪造语音与非扩散方法在生成质量及对现有检测系统威胁方面的表现。结果表明,扩散生成语音的检测难度总体上与非扩散方法相当,具体表现受检测器架构影响存在一定差异。使用扩散声码器重编码对语音质量影响极小,整体语音质量可与非扩散方法相媲美。

原文摘要 · Abstract (English)

Advancements in artificial intelligence and machine learning have significantly improved synthetic speech generation. This paper explores diffusion models, a novel method for creating realistic synthetic speech. We create a diffusion dataset using available tools and pretrained models. Additionally, this study assesses the quality of diffusion-generated deepfakes versus non-diffusion ones and their potential threat to current deepfake detection systems. Findings indicate that the detection of diffusion-based deepfakes is generally comparable to non-diffusion deepfakes, with some variability based on detector architecture. Re-vocoding with diffusion vocoders shows minimal impact, and the overall speech quality is comparable to non-diffusion methods.

语音生成扩散模型深度伪造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。