arXiv:2410.04324cs.SDcs.AI2024-10被引 25

构建首个统一评测框架,系统评估AI语音伪造检测能力

Where are we in audio deepfake detection? A systematic analysis over generative and detection models

  • 提出SONAR框架,统一评测传统与大模型检测方法
  • 发现大模型检测器泛化能力强,跨语言表现稳定
  • 验证少样本微调可提升定制化检测效率,适合特定场景

近年来,基于生成式人工智能的文本转语音(TTS)和语音转换(VC)技术已能生成高质量、逼真的类人语音,带来身份冒用、欺诈、虚假信息传播等风险。但现有检测方法进展滞后,难以在不同数据集间泛化。本文提出SONAR——一个合成AI音频检测框架与基准测试平台,包含来自9个不同语音合成平台的全新评测数据集,涵盖主流TTS服务商与前沿TTS模型。这是首个统一评估传统与基础模型检测系统的框架。实验表明:(1) 现有检测方法存在局限性,而基础模型展现出更强泛化能力,可能源于其模型规模及预训练数据的规模与质量;(2) 语音基础模型具备稳健的跨语言泛化能力,在仅以英语微调的情况下仍保持多语言高性能,表明检测挑战主要来自合成音频的真实度而非语言特性;(3) 探索了少样本微调在提升泛化能力中的有效性,揭示其在个性化检测系统(如特定人物或机构专用)中的应用潜力。

原文摘要 · Abstract (English)

Recent advances in Text-to-Speech (TTS) and Voice-Conversion (VC) using generative Artificial Intelligence (AI) technology have made it possible to generate high-quality and realistic human-like audio. This poses growing challenges in distinguishing AI-synthesized speech from the genuine human voice and could raise concerns about misuse for impersonation, fraud, spreading misinformation, and scams. However, existing detection methods for AI-synthesized audio have not kept pace and often fail to generalize across diverse datasets. In this paper, we introduce SONAR, a synthetic AI-Audio Detection Framework and Benchmark, aiming to provide a comprehensive evaluation for distinguishing cutting-edge AI-synthesized auditory content. SONAR includes a novel evaluation dataset sourced from 9 diverse audio synthesis platforms, including leading TTS providers and state-of-the-art TTS models. It is the first framework to uniformly benchmark AI-audio detection across both traditional and foundation model-based detection systems. Through extensive experiments, (1) we reveal the limitations of existing detection methods and demonstrate that foundation models exhibit stronger generalization capabilities, likely due to their model size and the scale and quality of pretraining data. (2) Speech foundation models demonstrate robust cross-lingual generalization capabilities, maintaining strong performance across diverse languages despite being fine-tuned solely on English speech data. This finding also suggests that the primary challenges in audio deepfake detection are more closely tied to the realism and quality of synthetic audio rather than language-specific characteristics. (3) We explore the effectiveness and efficiency of few-shot fine-tuning in improving generalization, highlighting its potential for tailored applications, such as personalized detection systems for specific entities or individuals.

语音伪造检测框架大模型少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。