arXiv:2506.00853cs.SDeess.AS2025-06中稿 · Interspeech 2025被引 4

为口吃者定制语音识别模型,能显著降低识别错误率。

Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches

  • 为每位口吃者单独训练识别模型,提升个性化适配
  • 在自然对话中,错误率明显下降,尤其在即兴发言时
  • 适合开发更包容的语音交互产品,如虚拟助手

口吃表现为阻塞、延长和重复等非自愿不流畅现象,常被自动语音识别(ASR)系统误判,导致词错误率升高,使口吃者难以使用语音驱动技术。不同说话人和语境下的不流畅表现差异大,且标注的口吃语音数据有限,进一步增加了ASR训练难度。本文研究了针对口吃语音的ASR微调方法,比较了跨多人的通用模型与针对个体语音特征定制的个性化模型。在虚拟助手、视频访谈等多种语音-AI应用场景下评估个性化对转录准确率的影响。结果表明,个性化ASR显著降低了词错误率,尤其在自发性口语中效果突出,凸显了定制化模型在实现更包容语音技术方面的潜力。

原文摘要 · Abstract (English)

Stuttering -- characterized by involuntary disfluencies such as blocks, prolongations, and repetitions -- is often misinterpreted by automatic speech recognition (ASR) systems, resulting in elevated word error rates and making voice-driven technologies inaccessible to people who stutter. The variability of disfluencies across speakers and contexts further complicates ASR training, compounded by limited annotated stuttered speech data. In this paper, we investigate fine-tuning ASRs for stuttered speech, comparing generalized models (trained across multiple speakers) to personalized models tailored to individual speech characteristics. Using a diverse range of voice-AI scenarios, including virtual assistants and video interviews, we evaluate how personalization affects transcription accuracy. Our findings show that personalized ASRs significantly reduce word error rates, especially in spontaneous speech, highlighting the potential of tailored models for more inclusive voice technologies.

语音识别口吃个性化无障碍

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。