arXiv:2509.15516eess.AScs.SD2025-09被引 1

用元学习实现少样本语音个性化,零样本即可适配失语症患者语音。

The Universal Personalizer: Few-Shot Dysarthric Speech Recognition via Meta-Learning

  • 通过元学习融合上下文学习,单模型支持快速个性化
  • 在Euphonia上达到13.9%的词错误率,优于基线17.5%
  • 无需离线合并或定制分块,适合资源受限场景

个性化失语症语音识别受限于耗时的录音采集和用户专属训练。本文提出一种混合元训练方法,使单一模型可通过上下文学习(ICL)实现零样本与少样本的即时个性化。在Euphonia数据集上,词错误率(WER)达13.9%,优于独立说话人基线(17.5%)。在SAP Test-1上,5.3%的WER超越挑战赛优胜团队(5.97%)。Test-2中9.49%的WER仅次于冠军(8.11%),且无需依赖离线模型融合或自定义音频分块技术。通过数据筛选,使用同说话人随机样本可降低40%的错误率,验证了主动个性化有效性。静态文本筛选无法超越此基线,但理想相似度显示仍有巨大提升空间,凸显动态声学检索是下一前沿。数据消融实验证实模型可在低资源下快速适应新说话人,具备实际应用潜力。

原文摘要 · Abstract (English)

Personalizing dysarthric ASR is hindered by demanding enrollment collection and per-user training. We propose a hybrid meta-training method for a single model, enabling zero-shot and few-shot on-the-fly personalization via in-context learning (ICL). On Euphonia, it achieves 13.9% Word Error Rate (WER), surpassing speaker-independent baselines (17.5%). On SAP Test-1, our 5.3% WER outperforms the challenge-winning team (5.97%). On Test-2, our 9.49% trails only the winner (8.11%) but without relying on techniques like offline model-merging or custom audio chunking. Curation yields a 40% WER reduction using random same-speaker examples, validating active personalization. While static text curation fails to beat this baseline, oracle similarity reveals substantial headroom, highlighting dynamic acoustic retrieval as the next frontier. Data ablations confirm rapid low-resource speaker adaptation, establishing the model as a practical personalized solution.

语音识别元学习个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。