用真人误判率评估语音合成真实度,发现多数系统仍难骗过人类。
The State Of TTS: A Case Study with Human Fooling Rates
- 提出真人误判率(HFR)衡量机器语音骗人能力
- 商业模型零样本下接近真人水平,开源系统仍有差距
- 建议用高表现参考数据集评测,避免低标准误导
尽管近年来主观评价显示语音合成(TTS)技术快速进步,但当前系统能否在类图灵测试中真正骗过人类?我们引入真人误判率(HFR),直接衡量机器生成语音被误认为真人语音的频率。对开源与商用TTS模型的大规模评估揭示关键发现:(i)基于CMOS的“人类水平”宣称在欺骗测试中常不成立;(ii)TTS进展应以人类表现高的数据集为基准,若使用单调或表达力弱的参考样本,评测标准过低;(iii)商用模型在零样本场景下接近人类欺骗水平,而开源系统在自然对话语音上仍存困难;(iv)在高质量数据上微调可提升真实感,但无法完全弥合差距。研究强调需结合更真实、以人为中心的评估方式,补充现有主观测试。
原文摘要 · Abstract (English)
While subjective evaluations in recent years indicate rapid progress in TTS, can current TTS systems truly pass a human deception test in a Turing-like evaluation? We introduce Human Fooling Rate (HFR), a metric that directly measures how often machine-generated speech is mistaken for human. Our large-scale evaluation of open-source and commercial TTS models reveals critical insights: (i) CMOS-based claims of human parity often fail under deception testing, (ii) TTS progress should be benchmarked on datasets where human speech achieves high HFRs, as evaluating against monotonous or less expressive reference samples sets a low bar, (iii) Commercial models approach human deception in zero-shot settings, while open-source systems still struggle with natural conversational speech; (iv) Fine-tuning on high-quality data improves realism but does not fully bridge the gap. Our findings underscore the need for more realistic, human-centric evaluations alongside existing subjective tests.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。