arXiv:2608.17585cs.SDcs.CR2026-08中稿 · the 6th Symposium …

真实语音检测面临部署难题,需建立可商用数据标准与易懂评估体系

The Last Mile of Deepfake Speech Detection: An Industry-Academia Experience Report

  • 基于工业界合作经验,指出现有评测与真实场景脱节
  • 实际输入多为长时、降码率、部分合成的音频,非理想测试集
  • 提出需开发普通人能理解的评估指标,推动行业协作

当前合成语音检测基准在特定领域评估中错误率已低于1%,但在未见攻击、信道不匹配和分布偏移下性能显著下降。基于与商业语音识别厂商Phonexia为期三年的合作经验,我们报告了构建和部署检测系统所遇到的实际障碍:许多公开基准不允许用于商业模型开发;真实输入并非四秒干净片段,而是长时间、经编解码压缩、有时部分合成的录音;当校准系统输出对数似然比为2.5时,客户无法理解其决策意义。本文不提出新模型,而是将这些挑战转化为具体的研究与协作建议:建立可商用的数据集共享标准、设计贴近真实部署的评测基准,以及开发非专家可操作的评分体系。这些观察来自单一项目,需在其他场景中验证。

原文摘要 · Abstract (English)

Synthetic speech detection benchmarks now report sub-1% error rates on some in-domain evaluations, yet performance degrades under unseen attacks, channel mismatch, and distribution shift. Based on a three-year effort with Phonexia, a commercial speaker-recognition vendor, we report barriers encountered while building and deploying a detector. Many public benchmarks are not licensed for commercial model development. Real inputs are not four-second clean clips but long, codec-degraded, sometimes partially synthetic recordings. And when a calibrated system returns a log-likelihood ratio of 2.5, no one can tell the customer what it means for their decision. Rather than proposing a new model, we connect these barriers to concrete research and coordination proposals: shared standards for commercially usable datasets, realistic deployment benchmarks, and scores that non-experts can act on. These observations come from one project and should be tested in other settings.

语音检测工业落地评估标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。