arXiv:2510.06544cs.SDcs.CR2025-10

首次系统评估17种伪造语音生成与8种检测器的对抗关系,揭示检测短板。

Benchmarking Fake Voice Detection in the Fake Voice Generation Arms Race

  • 采用一对一测试协议,精细分析生成器与检测器的交互
  • 神经音频编码器和流匹配生成器可绕过顶级检测器
  • 无通用鲁棒检测器,性能高度依赖生成架构

伪造语音生成技术的快速进步引发了与检测系统之间的对抗竞赛,亟需保障音频生态安全。然而现有基准存在关键缺陷:通常将多种伪造语音样本合并为单一数据集评估,掩盖了方法特有的痕迹,模糊了检测器对不同生成范式的响应差异,难以深入理解其真实脆弱性。为此,我们提出首个生态级基准,通过创新的一对一评估协议,系统评测17种最先进的伪造语音生成器与8种领先检测器之间的互动关系。该细粒度分析揭示了传统聚合测试中被隐藏的漏洞与敏感性。我们还设计统一评分体系,量化生成器的逃避能力与检测器的鲁棒性,实现公平直接比较。跨域评估显示,基于神经音频编码器和流匹配的现代生成器持续规避顶尖检测器。发现无任何检测器具备普遍鲁棒性,其有效性随生成器架构显著变化,凸显当前防御体系的重大泛化差距。本研究提供了更真实的威胁图景,并为构建下一代检测系统提供可行洞见。

原文摘要 · Abstract (English)

The rapid advancement of fake voice generation technology has ignited a race with detection systems, creating an urgent need to secure the audio ecosystem. However, existing benchmarks suffer from a critical limitation: they typically aggregate diverse fake voice samples into a single dataset for evaluation. This practice masks method-specific artifacts and obscures the varying performance of detectors against different generation paradigms, preventing a nuanced understanding of their true vulnerabilities. To address this gap, we introduce the first ecosystem-level benchmark that systematically evaluates the interplay between 17 state-of-the-art fake voice generators and 8 leading detectors through a novel one-to-one evaluation protocol. This fine-grained analysis exposes previously hidden vulnerabilities and sensitivities that are missed by traditional aggregated testing. We also propose unified scoring systems to quantify both the evasiveness of generators and the robustness of detectors, enabling fair and direct comparisons. Our extensive cross-domain evaluation reveals that modern generators, particularly those based on neural audio codecs and flow matching, consistently evade top-tier detectors. We found that no single detector is universally robust; their effectiveness varies dramatically depending on the generator's architecture, highlighting a significant generalization gap in current defenses. This work provides a more realistic assessment of the threat landscape and offers actionable insights for building the next generation of detection systems.

语音伪造检测对抗生成模型安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。