arXiv:2508.20916cs.CL2025-08AAAI被引 7

SageLM可多维度评估语音大模型,且解释性强。

SageLM: A Multi-aspect and Explainable Large Language Model for Speech Judgement

  • 联合评估语义与语音特征,避免传统分步方法的缺陷。
  • 在人工评估中达成82.79%一致率,显著优于基线模型。
  • 适用于语音交互系统评测,尤其适合需要可解释性的场景。

语音到语音(S2S)大语言模型是自然人机交互的基础,支持端到端口语对话系统。然而,对这些模型的评估仍是核心挑战。我们提出 exttt{SageLM},一种端到端、多维度且可解释的语音大模型,用于全面评估 S2S LLM。首先,不同于忽略声学特征的级联方法,SageLM 联合评估语义与声学维度。其次,它利用基于推理的监督提升可解释性,并引导模型学习,相比规则强化学习方法在评估结果对齐上表现更优。第三,我们引入 extit{SpeechFeedback} 合成偏好数据集,并采用两阶段训练范式以缓解语音偏好数据稀缺问题。在语义与声学双维度训练下,SageLM 与人类评估者达成 82.79% 的一致性,优于级联和 SLM 基线至少 7.42% 和 26.20%。

原文摘要 · Abstract (English)

Speech-to-Speech (S2S) Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling end-to-end spoken dialogue systems. However, evaluating these models remains a fundamental challenge. We propose \texttt{SageLM}, an end-to-end, multi-aspect, and explainable speech LLM for comprehensive S2S LLMs evaluation. First, unlike cascaded approaches that disregard acoustic features, SageLM jointly assesses both semantic and acoustic dimensions. Second, it leverages rationale-based supervision to enhance explainability and guide model learning, achieving superior alignment with evaluation outcomes compared to rule-based reinforcement learning methods. Third, we introduce \textit{SpeechFeedback}, a synthetic preference dataset, and employ a two-stage training paradigm to mitigate the scarcity of speech preference data. Trained on both semantic and acoustic dimensions, SageLM achieves an 82.79\% agreement rate with human evaluators, outperforming cascaded and SLM-based baselines by at least 7.42\% and 26.20\%, respectively.

语音评估大模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。