用大模型融合多种信号,少样本下实现更准的语音质量评估
Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models

- 让大模型当裁判,综合声学特征和伪标签做判断
- 少样本时性能超越DNSMOS等传统方法,尤其在匹配数据上提升显著
- 适合标注数据稀缺的语音质量评估场景
本文提出GatherMOS,一种利用大语言模型(LLM)作为元评估器的新框架,通过整合轻量级声学描述符与DNSMOS和VQScore生成的伪标签,使LLM能够对异构输入进行推理,推断感知均值意见分(MOS)。我们进一步探索零样本和少样本上下文学习设置,结果显示,零样本GatherMOS在多种条件下表现稳定,而当支持样本与测试条件匹配时,少样本引导可带来显著性能提升。在VoiceBank-DEMAND数据集上的实验表明,GatherMOS在有限标注数据条件下,持续优于DNSMOS、VQScore、简单平均法,甚至超过基于学习的模型如CNN-BLSTM和MOS-SSL。这些结果突显了基于大模型聚合在非侵入式语音质量评估中的实际潜力。
原文摘要 · Abstract (English)
In this paper, we introduce GatherMOS, a novel framework that leverages large language models (LLM) as meta-evaluators to aggregate diverse signals into quality predictions. GatherMOS integrates lightweight acoustic descriptors with pseudo-labels from DNSMOS and VQScore, enabling the LLM to reason over heterogeneous inputs and infer perceptual mean opinion scores (MOS). We further explore both zero-shot and few-shot in-context learning setups, showing that zero-shot GatherMOS maintains stable performance across diverse conditions, while few-shot guidance yields large gains when support samples match the test conditions. Experiments on the VoiceBank-DEMAND dataset demonstrate that GatherMOS consistently outperforms DNSMOS, VQScore, naive score averaging, and even learning-based models such as CNN-BLSTM and MOS-SSL when trained under limited labeled-data conditions. These results highlight the potential of LLM-based aggregation as a practical strategy for non-intrusive speech quality evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。