现代语音合成骗过人类听力,连警告也无济于事。
Tracking the Trend in How Speech Synthesizers Deceive People

- 用2019、2022、2024年三款合成工具测试人类辨识能力
- 全段伪造时人类识别准确率从90%降至48%,部分篡改仅9%准确
- 人类与检测器互补失效,适合安全与验证系统研究者关注
语音合成技术进步使深度伪造音频极为逼真。早期研究显示人类识别准确率达70%-80%,但主要基于较旧合成工具。本文对比2019、2022、2024年发布的三种语音合成工具,在82名信息技术专业人士中测试其对伪造语音的识别能力,并与六个预训练检测器在同一数据上进行基准对比。对于完全合成语音(全伪造),F1得分从RTVC和YourTTS的约90%降至ElevenLabs的48%,即便听众已被告知存在深度伪造。在部分伪造场景中,即仅修改一句话,严格准确率降至9%,且77%情况下听者将合成句误判为真实。人类与检测器表现互补性失效,均无法可靠定位短时篡改。此外,听者越来越将真实语音误标为伪造,削弱了对未篡改音频的信任。结果表明,在选定的现代与部分伪造条件下,仅依赖人类感知不可靠,亟需程序化验证、来源追溯、水印技术和逐段检测。
原文摘要 · Abstract (English)
Advances in speech synthesis have made deepfake audio highly realistic. Earlier studies reported 70-80% human detection accuracy, but relied primarily on older synthesizers. We compare human detection for three selected voice synthesis tools released in 2019, 2022, and 2024 with 82 IT professionals, and benchmark humans against six pretrained detectors on the same material. For fully synthetic speech (full spoofs), the F1 score drops from about 90% for RTVC and YourTTS to 48% for ElevenLabs, although listeners were explicitly warned that deepfakes were present. For partial spoofing, where only one sentence of an utterance is altered, strict accuracy falls to 9%, and listeners classify the synthetic sentence as bona fide 77% of the time. Humans and detectors fail in complementary ways, and neither reliably localizes short manipulations. Additionally, listeners increasingly mislabel bona fide speech as fake, eroding trust in unmanipulated audio. These findings show that human perception alone is unreliable for the selected modern and partial-spoof conditions and motivate procedural verification, provenance, watermarking, and segment-level detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。