arXiv:2409.09914eess.AScs.SD2024-09中稿 · IEEE ICASSP 2025被引 12

用大模型实现无需训练的语音质量评估,效果优于现有方法。

A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models

  • 将Whisper转文本,再用提示工程评估语音自然度
  • 对语音可懂度预测准确率高,与识别错误率相关性强
  • 零样本无需训练,适合快速评估新语音数据

本文研究两种基于大语言模型的零样本非侵入式语音评估策略。首先探索GPT-4o的音频分析能力;其次提出GPT-Whisper,利用Whisper进行音频转写,并通过定向提示工程评估文本自然度。实验对比GPT-4o与GPT-Whisper的预测指标,考察其与人工质量/可懂度评分及自动语音识别字符错误率(CER)的相关性。结果表明,GPT-4o单独使用效果较差,而GPT-Whisper预测准确率更高,与语音质量和可懂度有中等相关性,与CER相关性更强。相比SpeechLMScore和DNSMOS,GPT-Whisper在可懂度评估上表现更优,但在质量估计上略逊于SpeechLMScore。此外,其在Whisper CER的Spearman秩相关性上优于监督模型MOS-SSL和MTI-Net。结果验证了GPT-Whisper在无需额外训练数据下实现零样本语音评估的潜力。

原文摘要 · Abstract (English)

This work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which uses Whisper as an audio-to-text module and evaluates the naturalness of text via targeted prompt engineering. We evaluate the assessment metrics predicted by GPT-4o and GPT-Whisper, examining their correlation with human-based quality and intelligibility assessments and the character error rate (CER) of automatic speech recognition. Experimental results show that GPT-4o alone is less effective for audio analysis, while GPT-Whisper achieves higher prediction accuracy, has moderate correlation with speech quality and intelligibility, and has higher correlation with CER. Compared to SpeechLMScore and DNSMOS, GPT-Whisper excels in intelligibility metrics, but performs slightly worse than SpeechLMScore in quality estimation. Furthermore, GPT-Whisper outperforms supervised non-intrusive models MOS-SSL and MTI-Net in Spearman's rank correlation for CER of Whisper. These findings validate GPT-Whisper's potential for zero-shot speech assessment without requiring additional training data.

语音评估大模型零样本Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。