arXiv:2510.00743cs.SDcs.AI2025-10被引 3

用偏好比较重做语音质量评估,让模型更准更可复现。

From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling

  • 将多种MOS数据集转为偏好对比任务,统一评估标准。
  • 标量奖励模型准确率达74%以上,但对细微差异仍难区分。
  • 提出感知MOS差异的生成式奖励模型,提升细粒度判别能力。

评估合成语音的感知质量对指导语音生成模型的开发至关重要。传统方法依赖人工主观评分如均值意见分(MOS),但存在标注成本高、标准不一、可复现性差等问题。为此,本文提出MOS-RMBench,将多样化的MOS数据集统一转化为偏好对比任务,实现跨数据集的严格评估。基于此,系统构建并评估了三种奖励建模范式:标量奖励模型、半标量奖励模型和生成式奖励模型(GRMs)。实验发现:(1) 标量模型整体表现最佳,准确率持续超过74%;(2) 多数模型在合成语音上的表现显著低于人类语音;(3) 所有模型在MOS差异极小的样本对上均表现不佳。为改善这一问题,提出一种感知MOS差异的GRM,引入基于MOS差值的奖励函数,使模型能根据样本对难度自适应调整奖励。实验表明,该方法显著提升细粒度质量判别能力,并缩小与标量模型在最困难案例上的差距。本工作旨在建立基准与方法框架,推动自动语音质量评估研究更严谨、可扩展。

原文摘要 · Abstract (English)

Assessing the perceptual quality of synthetic speech is crucial for guiding the development and refinement of speech generation models. However, it has traditionally relied on human subjective ratings such as the Mean Opinion Score (MOS), which depend on manual annotations and often suffer from inconsistent rating standards and poor reproducibility. To address these limitations, we introduce MOS-RMBench, a unified benchmark that reformulates diverse MOS datasets into a preference-comparison setting, enabling rigorous evaluation across different datasets. Building on MOS-RMBench, we systematically construct and evaluate three paradigms for reward modeling: scalar reward models, semi-scalar reward models, and generative reward models (GRMs). Our experiments reveal three key findings: (1) scalar models achieve the strongest overall performance, consistently exceeding 74% accuracy; (2) most models perform considerably worse on synthetic speech than on human speech; and (3) all models struggle on pairs with very small MOS differences. To improve performance on these challenging pairs, we propose a MOS-aware GRM that incorporates an MOS-difference-based reward function, enabling the model to adaptively scale rewards according to the difficulty of each sample pair. Experimental results show that the MOS-aware GRM significantly improves fine-grained quality discrimination and narrows the gap with scalar models on the most challenging cases. We hope this work will establish both a benchmark and a methodological framework to foster more rigorous and scalable research in automatic speech quality assessment.

语音生成奖励建模评估基准偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。