arXiv:2505.12991cs.SDeess.AS2025-05中稿 · Interspeech 2025被引 13

用大模型生成的语音合成数据,提升发音障碍者语音识别效果。

Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition

  • 用大模型生成文本,训练语音合成器模拟发音障碍语音。
  • 结合个性化特征,使识别错误率降低23%以上。
  • 适合研究语音无障碍与个性化语音识别的人士。

本文针对发音障碍语音识别挑战赛提交方案。通过参数高效微调与潜在音频表示结合,改进编码器-解码器语音识别系统。利用大语言模型生成提示词,指导Parler-TTS微调生成语料一致的发音障碍语音合成数据。基于x向量的个性化微调持续降低词错误率(WER)。AdaLoRA适配器相比全量微调和标准低秩适配,分别实现约23%和22%的相对WER下降。引入wav2vec 2.0音频表征进一步提升性能,相对优化约5%。使用合成发音障碍语音训练,相较仅个性化微调,可带来最高约7%的相对WER改善。

原文摘要 · Abstract (English)

In this work, we present our submission to the Speech Accessibility Project challenge for dysarthric speech recognition. We integrate parameter-efficient fine-tuning with latent audio representations to improve an encoder-decoder ASR system. Synthetic training data is generated by fine-tuning Parler-TTS to mimic dysarthric speech, using LLM-generated prompts for corpus-consistent target transcripts. Personalization with x-vectors consistently reduces word error rates (WERs) over non-personalized fine-tuning. AdaLoRA adapters outperform full fine-tuning and standard low-rank adaptation, achieving relative WER reductions of ~23% and ~22%, respectively. Further improvements (~5% WER reduction) come from incorporating wav2vec 2.0-based audio representations. Training with synthetic dysarthric speech yields up to ~7% relative WER improvement over personalized fine-tuning alone.

语音识别发音障碍合成数据个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。