arXiv:2607.11706cs.SDcs.AI2026-07中稿 · InterSpeech 2026

新基准VoxENES 2026揭示语音伪造检测器在现代合成语音前的严重失效

VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion

论文配图:VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
图 1 · 摘自论文原文
  • 构建包含5.3万条音频的新基准,覆盖10种现代语音合成与10种后处理
  • 最先进检测器在新数据上误报率高达28.98%,多数接近随机水平
  • 适合安全研究者和语音对抗防御开发者使用

当前由大语言模型驱动的文本转语音(TTS)和语音转换(VC)系统生成的合成语音,与许多传统伪造检测基准中的生成方式存在差异。这种不匹配导致了时间泛化差距,可能高估检测器在真实后处理条件下的鲁棒性。为此,我们提出VoxENES 2026,一个包含53,628个音频样本的双语(英语和西班牙语)基准,采用10种现代语音合成方法生成,并在10种标准化后处理条件下评估。基于该基准,我们对8个未经微调的预训练检测器进行评测,发现性能显著下降:最佳模型总体误报率(EER)为28.98%,多数在现代生成器和扰动下表现接近或低于随机水平。结果表明,现有检测器严重依赖脆弱特征,VoxENES 2026可作为开发鲁棒语音伪造应对措施的实用测试平台。

原文摘要 · Abstract (English)

Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish) benchmark of 53,628 audio samples generated using 10 contemporary speech synthesis methods and evaluated under 10 standardized post-processing conditions. Using VoxENES 2026, we benchmark eight pretrained detectors without fine-tuning and observe substantial performance degradation: the best model achieves 28.98\% EER overall, while most perform near or below random chance across modern generators and perturbations. Our results highlight the reliance on brittle artifacts in current detectors and establish VoxENES 2026 as a practical testbed for developing robust audio spoofing countermeasures.

语音伪造检测基准大模型语音后处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。