arXiv:2506.21622cs.CLcs.AI2025-06被引 5

用语义重连法轻量个性化语音识别,提升障碍人群语音转写准确率。

Adapting Foundation Speech Recognition Models to Impaired Speech: A Semantic Re-chaining Approach for Personalization of German Speech

  • 基于语义一致性扩充小规模障碍语音数据集
  • 在儿童结构性语音障碍数据上显著提升转录质量
  • 适合无障碍语音交互与个性化语音模型开发

由脑性麻痹或遗传疾病引起的语音障碍给自动语音识别(ASR)系统带来重大挑战。尽管近期取得进展,但像Whisper这样的ASR模型仍难以处理非典型语音,主要受限于训练数据不足以及收集和标注非典型语音样本的困难。本文提出一种实用且轻量的个性化管道,通过语义一致性筛选词语并扩充小规模障碍语音数据集。该方法应用于一名具有结构性语音障碍的儿童数据,显著提升了转录质量,展现出降低非典型语音群体沟通障碍的潜力。

原文摘要 · Abstract (English)

Speech impairments caused by conditions such as cerebral palsy or genetic disorders pose significant challenges for automatic speech recognition (ASR) systems. Despite recent advances, ASR models like Whisper struggle with non-normative speech due to limited training data and the difficulty of collecting and annotating non-normative speech samples. In this work, we propose a practical and lightweight pipeline to personalize ASR models, formalizing the selection of words and enriching a small, speech-impaired dataset with semantic coherence. Applied to data from a child with a structural speech impairment, our approach shows promising improvements in transcription quality, demonstrating the potential to reduce communication barriers for individuals with atypical speech patterns.

语音识别障碍语音个性化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。