研究语音识别错误对大模型应用的影响,提出新评估方法。
An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications
- 分析大模型纠正语音识别错误的能力,改进评估思路。
- 发现不同错误类型对下游任务影响差异显著。
- 适合关注语音交互系统优化的研究者与工程师。
自动语音识别(ASR)在人机交互中扮演关键角色,是众多应用的接口。传统上,ASR性能通过词错误率(WER)评估,该指标量化生成转录中的插入、删除和替换数量。然而,随着大型语言模型(LLM)作为核心处理组件在各类应用中日益普及,不同类型的ASR错误对下游任务的影响亟需深入探讨。本文分析了LLM纠正ASR引入错误的能力,并提出一种针对LLM驱动应用的新型ASR性能评估方法。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) plays a crucial role in human-machine interaction and serves as an interface for a wide range of applications. Traditionally, ASR performance has been evaluated using Word Error Rate (WER), a metric that quantifies the number of insertions, deletions, and substitutions in the generated transcriptions. However, with the increasing adoption of large and powerful Large Language Models (LLMs) as the core processing component in various applications, the significance of different types of ASR errors in downstream tasks warrants further exploration. In this work, we analyze the capabilities of LLMs to correct errors introduced by ASRs and propose a new measure to evaluate ASR performance for LLM-powered applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。