arXiv:2606.19823eess.AScs.LG2026-06中稿 · Interspeech 2026, …

用零样本语音克隆生成假数据,解决失语症语音识别数据少难题。

Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning

论文配图:Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning
图 1 · 摘自论文原文
  • 用零样本克隆技术生成失语症语音,无需额外录音。
  • 克隆数据微调后词错误率仅26.00%,接近真实数据表现。
  • 适合研究语音识别与无障碍技术的开发者参考。

自动语音识别在失语症语音上仍不可靠,主要因数据稀缺和说话人差异大。传统合成数据方法常需大量特定说话人数据,重新引入收集瓶颈。本文探索零样本语音克隆作为低负担数据增强策略,使用Higgs Audio V2克隆TORGO数据集中的说话人。对克隆数据、真实数据及混合数据微调Whisper-medium,并在保留的真实语音上评估。相比零样本(31.62% WER),克隆微调(Clone FT)达到26.00% WER,几乎媲美真实数据微调(24.44%)和混合数据微调(25.12%)。值得注意的是,克隆和混合微调在中重度失语者上优于真实数据微调。在SAP-1102跨语料库评估中,克隆微调取得最佳结果(相对提升11.45%)。结果表明,零样本克隆可提供可扩展的训练数据,绕过昂贵的数据收集瓶颈。

原文摘要 · Abstract (English)

Automatic speech recognition remains unreliable for dysarthric speech due to data scarcity and high inter-speaker variability. While synthetic data can address these gaps, traditional methods often require extensive speaker-specific data, reintroducing the collection bottleneck. We investigate zero-shot voice cloning as a low-burden augmentation strategy, using Higgs Audio V2 to clone speakers in the TORGO dataset. We fine-tune (FT) Whisper-medium on cloned, real, and hybrid data and evaluate on held-out real speech. Compared to the zero-shot (31.62%), Clone FT achieved a competitive 26.00% WER, nearly matching the 24.44% and 25.12% seen with Real and Hybrid FT, respectively. Notably, Clone and Hybrid FT outperform Real FT for moderate-severe speakers. Clone FT achieves the best results (11.45% relative) in cross-corpus evaluation on the SAP-1102. These results suggest that zero-shot cloning provides scalable training data that circumvents the costly data collection bottleneck.

语音识别失语症数据增强零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。