arXiv:2510.14664cs.SDeess.AS2025-10ACL被引 25

用大模型做语音质量评估,能解释还能跨语言通用。

SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation

  • 让大模型像人一样分析语音质量并给出理由。
  • 在32,207段多语言语音上测试,表现优于传统方法。
  • 适合需要可解释评估的语音生成、反诈检测场景。

生成式语音技术发展迅速,但合成语音的感知质量评估仍是核心挑战。现有方法通常依赖单一评分或二分类判断,缺乏可解释性且难以跨任务和语言泛化。本文提出SpeechLLM-as-Judges新范式,使大语言模型(LLMs)能够进行结构化、带解释的语音质量评估。为此,我们构建了SpeechEval数据集,包含32,207个多语言语音片段及128,754条标注,覆盖质量评估、成对比较、改进建议和深度伪造检测四类任务。基于此,我们训练了具备语音质量感知能力的SQ-LLM模型,采用思维链推理与奖励优化策略提升性能。实验表明,SQ-LLM在多任务、多语言场景下表现优异,展现出该范式在推动语音质量评估方面的重要潜力。相关代码、模型与数据已开源。

原文摘要 · Abstract (English)

Generative speech technologies are progressing rapidly, but evaluating the perceptual quality of synthetic speech remains a core challenge. Existing methods typically rely on scalar scores or binary decisions, which lack interpretability and generalization across tasks and languages. We present SpeechLLM-as-Judges, a new paradigm for enabling large language models (LLMs) to conduct structured and explanation-based speech quality evaluation. To support this direction, we introduce SpeechEval, a large-scale dataset containing 32,207 multilingual speech clips and 128,754 annotations spanning four tasks: quality assessment, pairwise comparison, improvement suggestion, and deepfake detection. Based on this resource, we develop SQ-LLM, a speech-quality-aware LLM trained with chain-of-thought reasoning and reward optimization to improve capability. Experimental results show that SQ-LLM delivers strong performance across tasks and languages, revealing the potential of this paradigm for advancing speech quality evaluation. The relevant code, models, and data are publicly available at https://github.com/NKU-HLT/SpeechLLM-as-Judges.

语音评估大模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。