真实语音可信度下降:人对合成语音判断力未变,但开始怀疑真声音。
Eroding Trust in Real Speech: A Large-Scale Study of Human Audio Deepfake Perception
- 138种语音系统生成音频,1768人参与听觉测试。
- 真人语音识别准确率从72.7%降至64.1%,假语音识别率仅微降。
- 商业级与自回归模型生成的音频最难识别(61.3%-65.9%)。
近年来语音深度伪造技术快速进步,但其对人类对真实语音信任的影响尚未被研究。本文开展了迄今最大规模的语音深伪感知听觉实验,收集了来自1,768名参与者对138种文本到语音及语音转换系统的35,532次判断。核心发现为信任偏移:与2021年基线相比,人们对伪造样本的识别准确率几乎不变(72.9% → 71.2%),但对真实样本的识别准确率从72.7%降至64.1%。人们并非更难发现合成痕迹,而是越来越不信任真实语音。商业系统和自回归语言模型生成的音频最难检测(61.3%–65.9%),而传统seq2seq与流匹配模型生成的音频仍较易识别(75.4%–76.8%)。作为基准的机器学习检测器在所有条件下准确率均超过94.5%。结果表明,现代深度伪造的主要威胁或非单纯欺骗,而是对真实音频信任的侵蚀。
原文摘要 · Abstract (English)
Audio deepfakes have improved rapidly recently, yet their effect on human trust in real speech remains unstudied. We present the largest listening study on audio deepfake perception to date, collecting 35,532 judgments from 1,768 participants across 138 text-to-speech and voice conversion systems. Our central finding is a skepticism shift: compared to a 2021 baseline, human accuracy on fake samples barely changed (72.9% to 71.2%), but accuracy on real samples dropped from 72.7% to 64.1%. Participants are not worse at detecting synthesis artifacts; rather, they increasingly distrust authentic speech. Samples generated by commercial and autoregressive language model systems proved hardest to detect (61.3 - 65.9%), while those from traditional seq2seq and flow-matching models remain easier to spot (75.4 - 76.8%). An ML detector that served as a reference point maintained over 94.5% accuracy across all conditions. Our results suggest that the primary threat posed by modern deepfakes may not be mere deception, but the erosion of trust in genuine audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。