arXiv:2505.11200cs.SDcs.AI2025-05被引 8

用真人听辨测试评估中文语音合成系统,更真实可靠。

Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese

  • 设计多维中文语音数据集,模拟真实对话场景
  • 通过是否像真人声音判断,区分模型表现差异
  • 自动生成评分工具,适合快速模型迭代

大语言模型的进展显著提升了语音合成系统在语调、自然度和情感表达上的控制能力,使系统接近人类水平。尽管平均意见分(MOS)仍是标准评估方法,但其主观性强、环境差异大且解释性差。现有数据集也缺乏多维度设计,常忽略说话风格、上下文多样性及陷阱句式,尤其在中文语音评估中更为明显。为此,我们提出音频图灵测试(Audio Turing Test, ATT),配套构建多维中文语料库ATT-Corpus,并采用类图灵测试的评估协议。评估者只需判断语音是否像真人,无需复杂评分,降低偏差,提升鲁棒性。为加速开发,我们还基于人工标注数据微调Qwen2-Audio-Instruct,实现自动评估工具Auto-ATT。实验表明,ATT能有效区分模型在特定能力维度的表现;Auto-ATT与人工评估高度一致,具备快速可靠的评估价值。白盒数据集与工具已发布于Hugging Face集合(https://huggingface.co/collections/meituan/audio-turing-test-682446320368164faeaf38a4)。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have significantly improved text-to-speech (TTS) systems, enhancing control over speech style, naturalness, and emotional expression, which brings TTS Systems closer to human-level performance. Although the Mean Opinion Score (MOS) remains the standard for TTS System evaluation, it suffers from subjectivity, environmental inconsistencies, and limited interpretability. Existing evaluation datasets also lack a multi-dimensional design, often neglecting factors such as speaking styles, context diversity, and trap utterances, which is particularly evident in Chinese TTS evaluation. To address these challenges, we introduce the Audio Turing Test (ATT), a multi-dimensional Chinese corpus dataset ATT-Corpus paired with a simple, Turing-Test-inspired evaluation protocol. Instead of relying on complex MOS scales or direct model comparisons, ATT asks evaluators to judge whether a voice sounds human. This simplification reduces rating bias and improves evaluation robustness. To further support rapid model development, we also finetune Qwen2-Audio-Instruct with human judgment data as Auto-ATT for automatic evaluation. Experimental results show that ATT effectively differentiates models across specific capability dimensions using its multi-dimensional design. Auto-ATT also demonstrates strong alignment with human evaluations, confirming its value as a fast and reliable assessment tool. The white-box ATT-Corpus and Auto-ATT can be found in ATT Hugging Face Collection (https://huggingface.co/collections/meituan/audio-turing-test-682446320368164faeaf38a4).

语音合成图灵测试中文评估自动评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。