针对罕见失语者语音识别,提出零样本/单样本数据增强新方法。
Zero- and One-Shot Data Augmentation for Sentence-Level Dysarthric Speech Recognition in Constrained Scenarios
- 设计文本覆盖策略,实现无需目标语音即可生成适配数据。
- 在未见过的失语者上提升识别准确率,零样本下增益显著。
- 适合资源受限场景,如康复训练和日常沟通系统部署。
近年来,失语者语音识别(DSR)研究从单字理解发展到句子级表达解析,满足了失语人群的沟通需求。然而,可用数据稀缺仍是制约句级DSR系统发展的关键瓶颈。为此,失语者数据增强(DDA)成为重要方向。生成模型常用于为自动语音识别任务生成训练数据,但其效果依赖于合成数据对目标域的表征能力。由于失语者发音差异大,现有模型难以生成有效数据,尤其在零样本或单样本学习场景下。为此,本文提出一种专为文本匹配设计的文本覆盖策略,实现高效零/单样本DDA,显著提升面对未见失语者的语音识别性能。该成果对失语康复项目及日常通用句子通信场景具有重要应用价值。
原文摘要 · Abstract (English)
Dysarthric speech recognition (DSR) research has witnessed remarkable progress in recent years, evolving from the basic understanding of individual words to the intricate comprehension of sentence-level expressions, all driven by the pressing communication needs of individuals with dysarthria. Nevertheless, the scarcity of available data remains a substantial hurdle, posing a significant challenge to the development of effective sentence-level DSR systems. In response to this issue, dysarthric data augmentation (DDA) has emerged as a highly promising approach. Generative models are frequently employed to generate training data for automatic speech recognition tasks. However, their effectiveness hinges on the ability of the synthesized data to accurately represent the target domain. The wide-ranging variability in pronunciation among dysarthric speakers makes it extremely difficult for models trained on data from existing speakers to produce useful augmented data, especially in zero-shot or one-shot learning settings. To address this limitation, we put forward a novel text-coverage strategy specifically designed for text-matching data synthesis. This innovative strategy allows for efficient zero/one-shot DDA, leading to substantial enhancements in the performance of DSR when dealing with unseen dysarthric speakers. Such improvements are of great significance in practical applications, including dysarthria rehabilitation programs and day-to-day common-sentence communication scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。