构建首个覆盖多种情感表达的语音伪造检测基准,揭示现有模型严重失效
AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks

- 构建涵盖TTS、VC、情绪化VC和大模型生成攻击的综合语音伪造数据集
- 260小时数据验证下,多数模型在情感语音上接近随机判断
- 发现情感状态与语音风格显著影响检测效果,适用于鲁棒性研究
语音深度伪造检测(SDD)系统在传统基准上表现优异,但现有数据集对情感丰富且近期的大音频-语言模型(LALM)攻击覆盖有限。现有情感伪造数据集规模小、攻击类型单一,通常仅包含语音转换(VC)或文本转语音(TTS)攻击。本文提出AffectDF,是首个涵盖情感表达的语音伪造检测综合基准,涵盖TTS、VC、情绪化VC及基于LALM的伪造攻击,覆盖五种情感状态下的表演式与自然情感语音。AffectDF包含约260小时由21种伪造技术生成的语音。我们评估了先进SDD系统在常规与情感伪造条件下的表现,包括仅推理提示与监督微调两种方式下的LALM检测器。实验表明,仅在常规数据集上训练的模型在AffectDF上性能急剧下降,部分系统接近随机水平。令人意外的是,大规模情感训练并未稳定提升跨域鲁棒性,说明当前SDD系统无法在情感与语调变化下学习通用伪造表征。鲁棒性在不同情感状态、攻击类型及表演/自然语音间差异显著。这些发现暴露了当前SDD系统的根本局限,并确立AffectDF作为开发更鲁棒检测模型的基准。
原文摘要 · Abstract (English)
Speech deepfake detection (SDD) systems achieve strong performance on conventional benchmarks; however, existing datasets provide limited coverage of emotionally expressive and recent large audio-language model (LALM)-based attacks. Existing emotional spoofing datasets are also limited in scale and attack diversity, typically covering only voice conversion (VC) or text-to-speech (TTS) attacks. We introduce AffectDF, the most comprehensive benchmark for emotionally expressive speech deepfakes, spanning TTS, VC, emotional VC, and LALM-based spoofing attacks across both acted and spontaneous emotional speech. AffectDF contains approximately 260 hours of speech generated using 21 spoofing attacks across five emotional states. We benchmark state-of-the-art SDD systems under conventional and emotional spoofing conditions, including LALM-based detectors evaluated with both inference-only prompting and supervised fine-tuning. Our experiments reveal severe robustness degradation when models trained on conventional benchmarks are evaluated on AffectDF, with several systems approaching near-random performance. Surprisingly, even large-scale emotional training does not consistently improve cross-domain robustness, indicating that current SDD systems fail to learn generalized spoof representations under emotional and prosodic variability. Robustness further varies substantially across emotional states, attack families, and acted vs spontaneous emotional speech conditions. These findings expose fundamental limitations of current SDD systems and establish AffectDF as a benchmark for developing more robust spoof detection models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。