arXiv:2603.16411cs.CLeess.AS2026-03

用多候选假设和大模型,自动修复语音识别中的罕见实体错误。

RECOVER: Robust Entity Correction via agentic Orchestration of hypothesis Variants for Evidence-based Recovery

  • 通过多种策略融合语音识别的多个候选结果作为证据
  • 实体词错误率降低8%-46%,召回率最高提升22个百分点
  • 适合金融、医疗等专业领域,需高精度实体识别的场景

在自动语音识别(ASR)中,稀有和领域特定术语的实体识别极具挑战性。在金融、医学和航空管制等领域,这类错误代价高昂。若实体完全未出现在ASR输出中,事后纠正将极为困难。为此,本文提出RECOVER,一种作为工具使用代理的智能纠错框架。它利用ASR生成的多个候选结果作为证据,检索相关实体,并在约束条件下应用大语言模型(LLM)进行纠正。采用1-Best、基于实体感知的选择、识别器输出投票误差减少(ROVER)集成以及LLM-Select四种策略融合候选结果。在五个不同数据集上的评估显示,其实体短语词错误率(E-WER)相对降低8%-46%,召回率最高提升22个百分点。其中LLM-Select策略在实体纠错上表现最佳,同时保持整体词错误率稳定。

原文摘要 · Abstract (English)

Entity recognition in Automatic Speech Recognition (ASR) is challenging for rare and domain-specific terms. In domains such as finance, medicine, and air traffic control, these errors are costly. If the entities are entirely absent from the ASR output, post-ASR correction becomes difficult. To address this, we introduce RECOVER, an agentic correction framework that serves as a tool-using agent. It leverages multiple hypotheses as evidence from ASR, retrieves relevant entities, and applies Large Language Model (LLM) correction under constraints. The hypotheses are used using different strategies, namely, 1-Best, Entity-Aware Select, Recognizer Output Voting Error Reduction (ROVER) Ensemble, and LLM-Select. Evaluated across five diverse datasets, it achieves 8-46% relative reductions in entity-phrase word error rate (E-WER) and increases recall by up to 22 percentage points. The LLM-Select achieves the best overall performance in entity correction while maintaining overall WER.

语音识别实体纠错大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。