arXiv:2508.18732cs.SDcs.AI2025-08

多患者联合微调提升失语语音识别准确率

Cross-Learning Fine-Tuning Strategy for Dysarthric Speech Recognition Via CDSD database

  • 用多个失语者数据同时微调模型,避免单一患者过拟合
  • 在CDSD数据集上实现最高13.15%的词错误率降低
  • 适合需要低数据依赖的临床语音识别系统

失语语音识别面临严重程度差异和与正常语音的差异挑战。传统方法为每位患者单独微调预训练于正常语音的语音识别模型,以避免特征冲突。然而实验表明,同时对多位失语说话人进行多说话人微调,反而能提升个体语音模式的识别效果。该策略通过更广泛的病理特征学习增强泛化能力,缓解说话人特异性过拟合,减少对单个患者数据的依赖,并提高目标说话人识别准确率——在CDSD数据集上相较单说话人微调,词错误率(WER)最高降低13.15%。

原文摘要 · Abstract (English)

Dysarthric speech recognition faces challenges from severity variations and disparities relative to normal speech. Conventional approaches individually fine-tune ASR models pre-trained on normal speech per patient to prevent feature conflicts. Counter-intuitively, experiments reveal that multi-speaker fine-tuning (simultaneously on multiple dysarthric speakers) improves recognition of individual speech patterns. This strategy enhances generalization via broader pathological feature learning, mitigates speaker-specific overfitting, reduces per-patient data dependence, and improves target-speaker accuracy - achieving up to 13.15% lower WER versus single-speaker fine-tuning.

语音识别失语症微调策略多说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。