arXiv:2606.13464cs.CLcs.AI2026-06

用知识图谱增强语音识别纠错,让长对话更准确。

Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations

论文配图:Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations
图 1 · 摘自论文原文
  • 构建动态更新的语义知识库,存储实体和潜在混淆词
  • 在10组实验中9次优于直接纠错,提升定位准确性
  • 适合长时交互场景下的语音转写优化

自动语音识别(ASR)纠错传统上聚焦于孤立话语或短上下文。然而,随着文本与语音在长交互中日益交织,ASR纠错需依赖对话级上下文证据。现有方法多依赖当前假设或拼接原始对话历史,在此类情境下,稀疏的纠错线索易被冗余和噪声掩盖。为此,我们提出一种面向长文本-语音交织对话的本体记忆增强型ASR纠错框架。该框架将前序交互历史组织为可动态更新的本体记忆,以可检索节点形式存储实体、术语、表面变体、潜在的ASR混淆项及语义关系,实现基于上下文的精准纠错。为评估此设定,我们构建了RAMC-Corr数据集,源自MAGIC-RAMC,专用于长距离ASR纠错且具备可追溯上下文。在RAMC-Corr上的实验表明,我们的方法在10组配对基线设置中,有9组优于直接纠错,并促使纠错更具选择性与上下文依据,有效应对上下文相关的识别错误。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) correction has traditionally focused on isolated utterances or short local contexts. However, as text and speech become increasingly interleaved in long interactions, ASR correction requires conversation-level contextual evidence. Existing ASR correction methods often rely on the current hypothesis or concatenate raw dialogue history. In such contexts, sparse correction evidence can be difficult to locate amid redundancy and noise. Addressing these challenges, we propose an ontology memory-augmented ASR correction framework for long text-speech interleaved conversations. The framework organizes preceding interaction history into a dynamically updatable ontology memory, where entities, terminology, surface variants, potential ASR confusions, and semantic relations are stored as retrievable nodes for context-grounded correction. To evaluate this setting, we construct RAMC-Corr, a dataset derived from MAGIC-RAMC for long-range ASR correction with grounded context. Experiments on RAMC-Corr show that our method improves over direct correction in 9 out of 10 paired backbone-setting combinations and encourages more selective and evidence-grounded corrections for context-dependent ASR errors.

语音识别纠错知识图谱长对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。