arXiv:2510.07437cs.CLcs.AI2025-10EMNLP被引 6

用大模型改进语音识别评分,更懂语义细微差别。

LASER: An LLM-based ASR Scoring and Evaluation Rubric

  • 基于大模型上下文学习,从示例中自动学得评分规则。
  • 印度语种测试中与人工标注相关性达94%。
  • 小模型微调后可精准判断错误应扣多少分。

标准语音识别评估指标(如词错率WER)常对形态和句法差异过度惩罚,而这些差异并不显著影响语义。我们提出基于大模型的评分框架LASER,利用先进大模型的上下文学习能力,通过包含详细示例的提示(prompt)进行学习。在使用Gemini 2.5 Pro时,针对印地语的LASER评分与人工标注的相关性高达94%。提示中使用的印地语示例也有效用于分析马拉地语、卡纳达语和马拉雅拉姆语中的错误。此外,我们还展示如何将较小的Llama 3模型在参考文本与语音识别输出的词对示例上进行微调,以预测应施加的惩罚程度,准确率接近89%。

原文摘要 · Abstract (English)

Standard ASR evaluation metrics like Word Error Rate (WER) tend to unfairly penalize morphological and syntactic nuances that do not significantly alter sentence semantics. We introduce an LLM-based scoring rubric LASER that leverages state-of-the-art LLMs' in-context learning abilities to learn from prompts with detailed examples. Hindi LASER scores using Gemini 2.5 Pro achieved a very high correlation score of 94% with human annotations. Hindi examples in the prompt were also effective in analyzing errors in other Indian languages such as Marathi, Kannada and Malayalam. We also demonstrate how a smaller LLM like Llama 3 can be finetuned on word-pair examples derived from reference and ASR predictions to predict what kind of penalty should be applied with close to 89% accuracy.

语音识别大模型评分体系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。