arXiv:2409.16644eess.AScs.CL2024-09中稿 · ICASSP 2025被引 31

用听觉大模型统一评估语音质量,支持多维度预测与自然语言描述。

Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation

  • 通过任务提示微调听觉大模型,实现端到端的语音质量评估。
  • 在NISQA、BVCC等数据集上达到主流小模型水平的MOS和相似度预测性能。
  • 可生成噪音、失真等质量缺陷的自然语言解释,适合需要可解释性的场景。

语音质量评估通常需从平均意见分(MOS)、说话人相似度(SIM)等多个维度进行,单一任务的小模型难以全面覆盖。本文提出利用近期出现的听觉大语言模型(Auditory LLMs)实现自动语音质量评估。通过设计特定任务提示,对听觉LLMs进行微调,使其能预测MOS、SIM及A/B测试结果,这些指标常用于文本转语音系统评估。此外,微调后的听觉LLM还能生成关于噪音、失真、不连续性及整体质量的自然语言描述,提升输出可解释性。在NISQA、BVCC、SOMOS和VoxSim等语音质量数据集上进行了大量实验,采用SALMONN、Qwen-Audio和Qwen2-Audio等开源听觉大模型,自然语言描述任务还评估了商用模型Google Gemini 1.5 Pro。结果表明,听觉大模型在预测MOS和SIM方面表现媲美现有最先进小模型,同时在A/B测试和自然语言描述任务中也展现出良好效果。相关数据处理脚本与微调模型权重已公开于https://github.com/bytedance/SALMONN。

原文摘要 · Abstract (English)

Speech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) \etc., which can be challenging to cover using one small model designed for a single task. In this paper, we propose leveraging recently introduced auditory large language models (LLMs) for automatic speech quality assessment. By employing task-specific prompts, auditory LLMs are finetuned to predict MOS, SIM and A/B testing results, which are commonly used for evaluating text-to-speech systems. Additionally, the finetuned auditory LLM is able to generate natural language descriptions assessing aspects like noisiness, distortion, discontinuity, and overall quality, providing more interpretable outputs. Extensive experiments have been performed on the NISQA, BVCC, SOMOS and VoxSim speech quality datasets, using open-source auditory LLMs such as SALMONN, Qwen-Audio, and Qwen2-Audio. For the natural language descriptions task, a commercial model Google Gemini 1.5 Pro is also evaluated. The results demonstrate that auditory LLMs achieve competitive performance compared to state-of-the-art task-specific small models in predicting MOS and SIM, while also delivering promising results in A/B testing and natural language descriptions. Our data processing scripts and finetuned model checkpoints can be found at https://github.com/bytedance/SALMONN.

语音评估听觉LLM可解释性多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。