对比三种先进语音合成模型,发现单一检测方法难应对新型音频伪造。
Audio Deepfake Detection in the Age of Advanced Text-to-Speech models
- 测试三种主流语音合成模型的伪造音频特征
- 多视角联合检测在所有模型上表现更稳定
- 适合关注语音伪造防御的研究者与安全工程师
近期文本转语音(TTS)系统的发展显著提升了合成语音的真实感,给音频深度伪造检测带来新挑战。本文对三种最先进的TTS模型——Dia2、Maya1和MeloTTS(分别代表流式、基于大语言模型和非自回归架构)进行了对比评估。利用Daily-Dialog数据集生成了12,000个合成音频样本,并在四个检测框架下进行测试,涵盖语义、结构和信号级分析。结果表明,检测器性能在不同生成机制间差异显著:针对某一TTS架构有效的模型,在其他架构上可能失效,尤其在处理基于大语言模型的合成语音时。相比之下,结合多分析层面的多视角检测方法在所有模型上均表现出鲁棒性。研究揭示了单范式检测器的局限性,强调需采用集成策略应对不断演进的音频伪造威胁。
原文摘要 · Abstract (English)
Recent advances in Text-to-Speech (TTS) systems have substantially increased the realism of synthetic speech, raising new challenges for audio deepfake detection. This work presents a comparative evaluation of three state-of-the-art TTS models--Dia2, Maya1, and MeloTTS--representing streaming, LLM-based, and non-autoregressive architectures. A corpus of 12,000 synthetic audio samples was generated using the Daily-Dialog dataset and evaluated against four detection frameworks, including semantic, structural, and signal-level approaches. The results reveal significant variability in detector performance across generative mechanisms: models effective against one TTS architecture may fail against others, particularly LLM-based synthesis. In contrast, a multi-view detection approach combining complementary analysis levels demonstrates robust performance across all evaluated models. These findings highlight the limitations of single-paradigm detectors and emphasize the necessity of integrated detection strategies to address the evolving landscape of audio deepfake threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。