arXiv:2511.07931cs.SDcs.AI2025-11被引 26

构建首个大规模语音自然度判断数据集与模型,提升生成语音的人类感知对齐度。

SpeechJudge: Towards Human-Level Judgment for Speech Naturalness

  • 基于9.9万组语音对构建人类偏好数据集,覆盖多语言与多种语音风格。
  • 提出新奖励模型在自然度判断任务中达77.2%准确率,显著超越现有方法。
  • 适用于语音合成模型后训练对齐人类偏好,尤其适合追求高自然度的应用场景。

将大生成模型与人类反馈对齐是关键挑战,尤其在语音合成领域,因缺乏大规模人类偏好数据集而尤为突出。为此,我们提出SpeechJudge,包含一个数据集、一个基准测试和一个以自然度为核心的奖励模型。首先,我们构建了SpeechJudge-Data,一个包含99,000组语音对的大规模人类反馈语料库,涵盖多种先进零样本文本到语音(TTS)模型、多样语音风格及多语言场景,并由人工标注智能可懂度与自然度偏好。基于此,我们建立SpeechJudge-Eval基准测试,评估发现现有指标与AudioLLMs在此任务上表现不佳,领先模型Gemini-2.5-Flash与人类判断一致性不足70%,凸显巨大提升空间。为此,我们开发了基于Qwen2.5-Omni-7B的生成式奖励模型SpeechJudge-GRM,通过两阶段微调:带思维链的监督微调(SFT)与在难点样本上的强化学习(GRPO)。在SpeechJudge-Eval上,该模型达到77.2%准确率(推理时扩展至10次采样后达79.4%),优于经典Bradley-Terry奖励模型(72.7%)。此外,该模型可作为奖励函数用于语音生成模型的后训练,促进其与人类偏好对齐。

原文摘要 · Abstract (English)

Aligning large generative models with human feedback is a critical challenge. In speech synthesis, this is particularly pronounced due to the lack of a large-scale human preference dataset, which hinders the development of models that truly align with human perception. To address this, we introduce SpeechJudge, a comprehensive suite comprising a dataset, a benchmark, and a reward model centered on naturalness--one of the most fundamental subjective metrics for speech synthesis. First, we present SpeechJudge-Data, a large-scale human feedback corpus of 99K speech pairs. The dataset is constructed using a diverse set of advanced zero-shot text-to-speech (TTS) models across diverse speech styles and multiple languages, with human annotations for both intelligibility and naturalness preference. From this, we establish SpeechJudge-Eval, a challenging benchmark for speech naturalness judgment. Our evaluation reveals that existing metrics and AudioLLMs struggle with this task; the leading model, Gemini-2.5-Flash, achieves less than 70% agreement with human judgment, highlighting a significant gap for improvement. To bridge this gap, we develop SpeechJudge-GRM, a generative reward model (GRM) based on Qwen2.5-Omni-7B. It is trained on SpeechJudge-Data via a two-stage post-training process: Supervised Fine-Tuning (SFT) with Chain-of-Thought rationales followed by Reinforcement Learning (RL) with GRPO on challenging cases. On the SpeechJudge-Eval benchmark, the proposed SpeechJudge-GRM demonstrates superior performance, achieving 77.2% accuracy (and 79.4% after inference-time scaling @10) compared to a classic Bradley-Terry reward model (72.7%). Furthermore, SpeechJudge-GRM can be also employed as a reward function during the post-training of speech generation models to facilitate their alignment with human preferences.

语音合成自然度评估奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。