用大模型评估语义,让语音识别能像人一样对话修正。
Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
- 用大模型当裁判,评估语音识别的语义准确性。
- 在多轮对话中不断修正识别结果,提升语义连贯性。
- 适合研究交互式语音识别与智能系统的人看。
近年来,自动语音识别(ASR)因模型架构和大规模训练数据的进步取得了显著进展。然而,仍有两个关键问题未被充分探索:其一,长期主导的词错误率(WER)对所有词汇同等对待,常无法反映句子层面的语义正确性;其二,交互式纠错——人类交流中的核心环节——在ASR研究中尚未得到系统性研究。本文在智能体(agentic)框架下整合这两方面,提出利用大模型作为评判者(LLM-as-a-Judge)来实现语义感知的评估,超越传统字符级准确率。同时,设计基于大模型驱动的智能体框架,模拟人类多轮交互,通过语义反馈实现识别输出的迭代优化。在GigaSpeech(英文)、WenetSpeech(中文)及ASRU 2019混语测试集上进行大量实验,客观与主观评估均验证了该框架在提升语义保真度和交互纠错能力方面的有效性。代码将开源,以推动交互式与智能体式语音识别的研究。
原文摘要 · Abstract (English)
Recent years have witnessed remarkable progress in automatic speech recognition (ASR), driven by advances in model architectures and large-scale training data. However, two important aspects remain underexplored. First, Word Error Rate (WER), the dominant evaluation metric for decades, treats all words equally and often fails to reflect the semantic correctness of an utterance at the sentence level. Second, interactive correction-an essential component of human communication-has rarely been systematically studied in ASR research. In this paper, we integrate these two perspectives under an agentic framework for interactive ASR. We propose leveraging LLM-as-a-Judge as a semantic-aware evaluation metric to assess recognition quality beyond token-level accuracy. Furthermore, we design an LLM-driven agent framework to simulate human-like multi-turn interaction, enabling iterative refinement of recognition outputs through semantic feedback. Extensive experiments are conducted on standard benchmarks, including GigaSpeech (English), WenetSpeech (Chinese), the ASRU 2019 code-switching test set. Both objective and subjective evaluations demonstrate the effectiveness of the proposed framework in improving semantic fidelity and interactive correction capability. We will release the code to facilitate future research in interactive and agentic ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。