arXiv:2606.14391cs.CLcs.AI2026-06中稿 · Interspeech 2026

用持续学习让语音识别更懂口吃,保留信息不丢

Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR

论文配图:Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR
图 1 · 摘自论文原文
  • 引入显式口吃标记符,在预训练模型中稳定学习
  • 在不同口吃分布数据上持续训练,保持整体识别率
  • 发现标记学习与识别性能存在权衡,跨方法共享注意力机制

尽管大规模语音识别(ASR)取得进展,但口吃性语音仍是难题,因为现有系统常被优化为省略口吃内容,导致信息丢失和幻觉。以往工作聚焦于逐字转录和口吃标记整合,但在小数据集上适应时易引发灾难性遗忘。本文通过持续学习(CL)并引入显式口吃标记符来解决此问题。首先将这些标记符注入预训练的ASR模型,建立稳定的标记机制;随后在具有不同口吃分布的额外数据集上继续训练。通过对训练过程中的模型动态进行详细分析,我们发现标记学习与ASR性能之间存在权衡,并且跨多种持续学习方法存在一致的交叉注意力头机制。

原文摘要 · Abstract (English)

Despite advances in large-scale Automatic Speech Recognition (ASR), disfluent speech remains challenging, as state-of-the-art systems are often optimized to omit disfluencies, leading to information loss and hallucinations. Prior work has focused on verbatim transcription and the integration of disfluency markers, but adapting models on limited datasets can lead to catastrophic forgetting of general-domain knowledge. We address this gap by leveraging continual learning (CL) with explicit disfluency tokens. We first introduce these tokens into a pretrained ASR model to establish stable token mechanisms, and then continue training on additional datasets with varying disfluency distributions. Through a detailed analysis of model dynamics during training, we identify a trade-off between marker learning and ASR performance, and a consistent cross-attention head mechanism shared across CL methods.

语音识别持续学习口吃处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。