语音深度伪造检测在真实场景中难以通用,因环境变化导致检测盲区扩大。
Why Speech Deepfake Detectors Won't Generalize: The Limits of Detection in an Open World
- 检测器在真实多变环境中面临覆盖债务,新条件叠加使盲区指数增长。
- 新型合成技术抹去旧有痕迹,对话类场景(如会议、社交)最难防护。
- 高风险决策应依赖多层防御,仅靠检测不可靠。
语音深度伪造检测器常在干净、标准的基准条件下评估,但实际部署面临设备、采样率、编码方式、环境及攻击类型不断变化的开放世界挑战。这造成‘覆盖债务’:每新增一种条件便与已有条件组合,生成的数据盲区增速远超数据收集速度。由于攻击者可针对未覆盖区域发起攻击,最差情况下的性能(而非平均得分)才决定安全性。通过分析近期跨测试框架的结果,我们发现两个规律:较新的合成器会消除旧检测器依赖的遗留特征;而对话类场景(如远程会议、访谈、社交媒体)始终是最难防护的。研究结果表明,仅依赖检测无法保障高风险决策安全。检测器应作为多层次防御体系中的辅助信号,配合来源验证、身份凭证与政策管控使用。
原文摘要 · Abstract (English)
Speech deepfake detectors are often evaluated on clean, benchmark-style conditions, but deployment occurs in an open world of shifting devices, sampling rates, codecs, environments, and attack families. This creates a ``coverage debt" for AI-based detectors: every new condition multiplies with existing ones, producing data blind spots that grow faster than data can be collected. Because attackers can target these uncovered regions, worst-case performance (not average benchmark scores) determines security. To demonstrate the impact of the coverage debt problem, we analyze results from a recent cross-testing framework. Grouping performance by bona fide domain and spoof release year, two patterns emerge: newer synthesizers erase the legacy artifacts detectors rely on, and conversational speech domains (teleconferencing, interviews, social media) are consistently the hardest to secure. These findings show that detection alone should not be relied upon for high-stakes decisions. Detectors should be treated as auxiliary signals within layered defenses that include provenance, personhood credentials, and policy safeguards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。