arXiv:2501.03805cs.SDcs.CL2025-01

新数据集让无缝语音篡改更难被发现,但仍有方法可检测。

Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits

  • 用Voicebox生成无缝语音编辑,避免传统拼接痕迹。
  • 人眼难以识别,但自监督模型仍能精准定位篡改处。
  • 适合研究语音伪造检测与防御的学者使用。

神经语音编辑技术的发展引发了对伪造攻击的担忧。传统部分编辑语料库主要关注剪切拼接编辑,虽保持说话人一致性,但常引入可检测的不连续性。近期方法如A³T和Voicebox通过利用上下文信息改善过渡效果。为推动伪造检测研究,我们构建了基于Voicebox的语音填充编辑(SINE)数据集,并详细说明了重实现Voicebox训练及数据集创建的过程。主观评估表明,使用该新技术编辑的语音比传统剪切拼接方法更难被察觉。尽管人类识别困难,实验结果表明自监督检测器在检测、定位及跨方法泛化方面表现优异。相关数据集与模型将公开发布。

原文摘要 · Abstract (English)

Neural speech editing advancements have raised concerns about their misuse in spoofing attacks. Traditional partially edited speech corpora primarily focus on cut-and-paste edits, which, while maintaining speaker consistency, often introduce detectable discontinuities. Recent methods, like A\textsuperscript{3}T and Voicebox, improve transitions by leveraging contextual information. To foster spoofing detection research, we introduce the Speech INfilling Edit (SINE) dataset, created with Voicebox. We detailed the process of re-implementing Voicebox training and dataset creation. Subjective evaluations confirm that speech edited using this novel technique is more challenging to detect than conventional cut-and-paste methods. Despite human difficulty, experimental results demonstrate that self-supervised-based detectors can achieve remarkable performance in detection, localization, and generalization across different edit methods. The dataset and related models will be made publicly available.

语音伪造检测方法自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。