用低秩贝叶斯微调提升残障语音识别准确率
Variational Low-Rank Adaptation for Personalized Impaired Speech Recognition
- 基于贝叶斯低秩适配实现小样本高效微调
- 在英德双语残障语音数据集上显著提升识别率
- 适合资源匮乏场景下个性化语音识别应用
由先天性疾病(如脑瘫、唐氏综合征、阿佩尔综合征)或后天脑损伤(如中风、外伤、肿瘤)引起的语音障碍,给自动语音识别(ASR)系统带来巨大挑战。尽管近期取得进展,但像Whisper这样的先进ASR模型在非标准语音上仍表现不佳,主要因训练数据有限且声学差异大。此外,收集和标注非标准语音极为困难:许多患者发音费力,而标注常需熟悉说话人的护理人员。本文提出一种基于贝叶斯低秩适配的新型ASR个性化方法,支持数据高效的微调。我们在英语UA-Speech数据集和新采集的德语儿童结构性语音障碍数据集BF-Sprache上验证了该方法。该数据集与方法设计均反映低资源环境下残障语音的真实挑战。实验表明,该方法显著提升了残障语音的识别准确率,同时保持数据与标注效率,为构建包容性ASR提供了可行路径。
原文摘要 · Abstract (English)
Speech impairments resulting from congenital disorders, such as cerebral palsy, down syndrome, or apert syndrome, as well as acquired brain injuries due to stroke, traumatic accidents, or tumors, present major challenges to automatic speech recognition (ASR) systems. Despite recent advancements, state-of-the-art ASR models like Whisper still struggle with non-normative speech due to limited training data availability and high acoustic variability. Moreover, collecting and annotating non-normative speech is burdensome: speaking is effortful for many affected individuals, while laborious annotation often requires caregivers familiar with the speaker. This work introduces a novel ASR personalization method based on Bayesian Low-rank Adaptation for data-efficient fine-tuning. We validate our method on the English UA-Speech dataset and a newly collected German speech dataset, BF-Sprache, from a child with structural speech impairment. The dataset and approach are designed to reflect the challenges of low-resource settings that include individuals with speech impairments. Our method significantly improves ASR accuracy for impaired speech while maintaining data and annotation efficiency, offering a practical path toward inclusive ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。