arXiv:2501.17202cs.SDcs.CL2025-01ICLR被引 46

让语音大模型学会评估语音质量,还能像人一样描述优劣

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators

  • 用真人评分构建首个自然语言语音质量评估语料库
  • 模型预测语音质量误差仅0.17,判断两段语音优劣准确率达98.6%
  • 生成的评价描述质量优于专用模型,适合开发智能语音助手

理想的多模态智能体应能感知输入模态的质量。尽管大语言模型(LLM)已具备处理语音任务的能力,但多数音频LLM仍无法识别所处理语音的质量。这一局限源于语音质量评估缺乏合适数据集,难以纳入多任务训练。为此,我们构建了首个基于自然语言的语音质量评估语料库,源自真实人类评分。该语料库不仅包含整体平均意见分(MOS),还提供多维度详细分析及质量退化原因识别,并支持类似人类的语音样本对比(A/B测试)。基于此语料库,我们提出一种音频LLM对齐方法(ALLD),通过大模型蒸馏引导其从原始语音中提取关键信息并生成有意义描述。实验表明,ALLD在MOS预测上均方误差为0.17,A/B测试准确率达98.6%;生成文本在两项任务上的BLEU分数分别为25.8和30.2,超越专用模型。该研究推动了音频LLM对语音信号的全面感知,助力现实世界听觉与感官智能体的发展。

原文摘要 · Abstract (English)

An ideal multimodal agent should be aware of the quality of its input modalities. Recent advances have enabled large language models (LLMs) to incorporate auditory systems for handling various speech-related tasks. However, most audio LLMs remain unaware of the quality of the speech they process. This limitation arises because speech quality evaluation is typically excluded from multi-task training due to the lack of suitable datasets. To address this, we introduce the first natural language-based speech evaluation corpus, generated from authentic human ratings. In addition to the overall Mean Opinion Score (MOS), this corpus offers detailed analysis across multiple dimensions and identifies causes of quality degradation. It also enables descriptive comparisons between two speech samples (A/B tests) with human-like judgment. Leveraging this corpus, we propose an alignment approach with LLM distillation (ALLD) to guide the audio LLM in extracting relevant information from raw speech and generating meaningful responses. Experimental results demonstrate that ALLD outperforms the previous state-of-the-art regression model in MOS prediction, with a mean square error of 0.17 and an A/B test accuracy of 98.6%. Additionally, the generated responses achieve BLEU scores of 25.8 and 30.2 on two tasks, surpassing the capabilities of task-specific models. This work advances the comprehensive perception of speech signals by audio LLMs, contributing to the development of real-world auditory and sensory intelligent agents.

语音质量大模型评估LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。