arXiv:2602.13891cs.SDcs.AI2026-02被引 2

用可解释的推理机制提升语音自然度评估,让机器像人一样评判语音质量。

GSRM: Generative Speech Reward Model for Speech RLHF

  • 将语音自然度评估拆解为特征提取与基于特征的推理链,实现可解释判断。
  • 在31,000条专家评分数据上训练,预测相关性接近人类评分者一致性。
  • 可作为在线强化学习中语音生成的验证器,提升语音模型自然度。

近期语音语言模型(如 GPT-4o Voice Mode 和 Gemini Live)展现出出色的语音生成能力,但合成语音的审美自然度仍不及真人语音。提升生成质量需可靠评估工具。现有自然度评估多回归音频为标量分数,解释性差且泛化能力弱。受生成式奖励建模启发,本文提出生成式语音奖励模型(GSRM),一种面向语音的以推理为核心的奖励模型。GSRM 将自然度评估分解为可解释的声学特征提取与特征驱动的思维链推理,实现可解释判断。为此,我们构建了一个大规模人工反馈数据集(31,000 条专家评分)及一个真实用户-助手语音交互的域外基准。实验表明,GSRM 显著优于现有语音自然度预测模型,在自然度评分预测上的人机相关性接近人类评分者间一致性。进一步证明,GSRM 可作为在线强化学习微调中的有效验证器,提升语音大模型生成的自然度。

原文摘要 · Abstract (English)

Recent advances in speech language models, such as GPT-4o Voice Mode and Gemini Live, have demonstrated promising speech generation capabilities. Nevertheless, the aesthetic naturalness of the synthesized audio still lags behind that of human speech. Enhancing generation quality requires a reliable evaluator of speech naturalness. However, existing naturalness evaluators typically regress raw audio to scalar scores, offering limited interpretability of the evaluation and moreover fail to generalize to speech across different taxonomies. Inspired by recent advances in generative reward modeling, we propose the Generative Speech Reward Model (GSRM), a reasoning-centric reward model tailored for speech. The GSRM is trained to decompose speech naturalness evaluation into an interpretable acoustic feature extraction stage followed by feature-grounded chain-of-thought reasoning, enabling explainable judgments. To achieve this, we curated a large-scale human feedback dataset comprising 31k expert ratings and an out-of-domain benchmark of real-world user-assistant speech interactions. Experiments show that GSRM substantially outperforms existing speech naturalness predictors, achieving model-human correlation of naturalness score prediction that approaches human inter-rater consistency. We further show how GSRM can improve the naturalness of speech LLM generations by serving as an effective verifier for online RLHF.

语音生成奖励模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。