arXiv:2505.20163cs.CLeess.AS2025-05中稿 · INTERSPEECH 2025被引 5

用大模型生成纠错提升失语症语音识别准确率

Exploring Generative Error Correction for Dysarthric Speech Recognition

  • 两阶段框架:先识别再用大模型纠错
  • 在结构化与即兴语音上显著提升准确率
  • 适合语音无障碍项目、残障语音研究者

尽管端到端自动语音识别(ASR)取得显著进展,准确转录失语症语音仍是重大挑战。本文为INTERSPEECH 2025语音可及性项目提出一种两阶段框架,结合先进语音识别模型与基于大语言模型的生成式错误纠正(GER)。我们评估了不同模型规模与训练策略配置,引入特定假设选择机制以提升转写准确性。在语音可及性项目数据集上的实验表明,该方法在结构化和即兴语音上表现优异,但在单字识别方面仍存挑战。通过全面分析,我们揭示了声学与语言建模在失语症语音识别中的互补作用。

原文摘要 · Abstract (English)

Despite the remarkable progress in end-to-end Automatic Speech Recognition (ASR) engines, accurately transcribing dysarthric speech remains a major challenge. In this work, we proposed a two-stage framework for the Speech Accessibility Project Challenge at INTERSPEECH 2025, which combines cutting-edge speech recognition models with LLM-based generative error correction (GER). We assess different configurations of model scales and training strategies, incorporating specific hypothesis selection to improve transcription accuracy. Experiments on the Speech Accessibility Project dataset demonstrate the strength of our approach on structured and spontaneous speech, while highlighting challenges in single-word recognition. Through comprehensive analysis, we provide insights into the complementary roles of acoustic and linguistic modeling in dysarthric speech recognition

语音识别失语症大模型纠错

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。