UniSRM统一评估语音生成质量,支持多维度细粒度判断。
UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment

- 构建两阶段模型,通过推理机制实现细粒度语音评估
- 在多种任务上表现优于现有方法,与人工评分高度一致
- 适合需要规模化、可复现语音评估的研究与工业场景
语音生成的评估仍严重依赖人工打分(如平均意见分MOS),成本高、主观性强且难以大规模复现。尽管已有研究探索基于AudioLLM的评判模型,但多数仅针对单一场景(如话语级质量或单轮对话),覆盖任务和评估维度有限。本文提出UniSRM,一种统一的语音评分模型,支持多维度、可解释的奖励信号,并具备可靠的推理能力。为支持训练与评估,我们构建了UniSRM-Data和UniSRM-Bench数据集,涵盖从话语级质量到上下文连贯性在内的多样化语音评估任务。基于此,我们设计了两阶段流水线的UniSRM模型,实现基于推理的细粒度评估。此外,引入一致性推理奖励以提升推理可靠性。实验表明,UniSRM在广泛语音评估任务中均能提供更可靠、更符合人类判断的结果,为语音质量的可扩展、统一评估提供了实用基础。
原文摘要 · Abstract (English)
Evaluating speech generation still relies heavily on human judgments, such as Mean Opinion Score (MOS), which are expensive, subjective, and difficult to reproduce at scale. While a few recent studies have begun to explore AudioLLM-based judge models, existing efforts typically target only a narrow set of scenarios (e.g., utterance-level quality or single-turn dialogue) and provide limited coverage of diverse speech generation tasks and evaluation dimensions. In this work, we propose UniSRM, a unified speech reward model that can support multi-dimensional, interpretable reward signals with reliable reasoning. To support training and evaluation, we introduce UniSRM-Data and UniSRM-Bench, covering speech evaluation tasks from utterance-level quality to context-level coherence. Based on this dataset, we present the unified speech reward model, UniSRM, with a two-stage pipeline that enables reasoning-based fine-grained assessment. Furthermore, we introduce Reasoning-Consistent Rewards to improve the reliability of the reasoning process. Experiments show that UniSRM delivers more reliable and human-aligned judgments across a broad range of speech evaluation tasks, offering a practical foundation for scalable and unified evaluation of speech quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。