arXiv:2509.14263cs.CLcs.SD2025-09

用细粒度指令优化语音识别后编辑,又快又准。

Context-Enhanced Granular Edit Representation for Efficient and Accurate ASR Post-editing

  • 用结构化指令代替重写,提升编辑效率
  • 在LibriSpeech上实现最低词错误率
  • 适合需要高效高精度的语音后处理场景

尽管语音识别(ASR)技术已广泛应用于工业界和大众,但系统仍常出现错误,需人工后编辑。虽然大语言模型(LLM)在后编辑中表现强大,但基线全重写模型存在推理效率低的问题,常重复生成相同冗余文本。紧凑编辑表示虽存在,但往往缺乏准确性和上下文信息。本文提出CEGER(Context-Enhanced Granular Edit Representation),一种用于高精度、高效语音识别后编辑的紧凑编辑表示方法。CEGER使LLM能生成一系列结构化、细粒度、富含上下文的指令,以修改原始ASR输出。一个独立的展开模块基于这些指令确定性地重建修正文本。在LibriSpeech数据集上的大量实验表明,CEGER在词错误率(WER)上优于全重写和以往紧凑表示方法,达到当前最佳水平。

原文摘要 · Abstract (English)

Despite ASR technology being full-scale adopted by industry and for large portions of the population, ASR systems often have errors that require editors to post-edit text quality. While LLMs are powerful post-editing tools, baseline full rewrite models have inference inefficiencies because they often generate the same redundant text over and over again. Compact edit representations have existed but often lack the efficacy and context required for optimal accuracy. This paper introduces CEGER (Context-Enhanced Granular Edit Representation), a compact edit representation that was generated for highly accurate, efficient ASR post-editing. CEGER allows LLMs to generate a sequence of structured, fine-grained, contextually rich commands to modify the original ASR output. A separate expansion module deterministically reconstructs the corrected text based on the commands. Extensive experiments on the LibriSpeech dataset that were conducted, CEGER achieves state-of-the-art accuracy, achieving the lowest word error rate (WER) versus full rewrite and prior compact representations.

语音识别后编辑大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。