arXiv:2506.22810cs.SDeess.AS2025-06中稿 · Interspeech 2025被引 4

用自训练提升Whisper对长段口吃语音的识别能力

A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition

  • 通过自训练扩充数据并适应不完整语音片段
  • 在SAP挑战赛中词错误率和语义得分均第二
  • 适合研究口吃语音识别与模型鲁棒性提升

口吃语音识别(DSR)提升了行动受限人群使用智能设备的便利性。以往研究受限于现有数据集多为孤立词汇、指令短语及少量句子,且仅涵盖少数说话人,导致研究局限于命令交互系统与说话人适配。Speech Accessibility Project(SAP)发布了一个大规模多样化的英语口吃语音数据集,推动了SAP挑战赛,旨在构建说话人与文本无关的DSR系统。本文通过一种新颖的自训练方法,显著提升了Whisper模型在长段口吃语音上的表现。该方法增加了训练数据量,并使模型能够适应推理时可能遇到的不完整语音片段。我们的系统在SAP挑战赛中以词错误率和语义得分两项指标均位列第二。

原文摘要 · Abstract (English)

Dysarthric speech recognition (DSR) enhances the accessibility of smart devices for dysarthric speakers with limited mobility. Previously, DSR research was constrained by the fact that existing datasets typically consisted of isolated words, command phrases, and a limited number of sentences spoken by a few individuals. This constrained research to command-interaction systems and speaker adaptation. The Speech Accessibility Project (SAP) changed this by releasing a large and diverse English dysarthric dataset, leading to the SAP Challenge to build speaker- and text-independent DSR systems. We enhanced the Whisper model's performance on long dysarthric speech via a novel self-training method. This method increased training data and adapted the model to handle potentially incomplete speech segments encountered during inference. Our system achieved second place in both Word Error Rate and Semantic Score in the SAP Challenge.

语音识别自训练口吃语音Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。