arXiv:2509.16718cs.SDcs.CL2025-09EMNLP被引 4

混合通用与个人特征,提升失语症语音识别准确率

Idiosyncratic Versus Normative Modeling of Atypical Speech Recognition: Dysarthric Case Studies

  • 先学通用模式再个性化适配,兼顾泛化与个体差异
  • 仅需128句训练数据,错误率36.43%,低于全个性化方法
  • 微调语音编码器效果最佳,错误率从71%降至32%

当前最先进的自动语音识别(ASR)模型(如Whisper)在失语症等异常语音上表现不佳。以往研究多聚焦完全个性化的(即个体特异性)模型,但结合泛化能力与个体差异的策略可能更有效。本文对比四种策略:(a)基于正常语音训练的规范模型,(b)完全个性化的个体模型,(c)基于其他失语症者数据训练的失语症规范模型,(d)先建模通用模式再适配个体的失语症-个体化模型。实验表明,失语症-个体化模型优于纯个性化方法,且所需个性化数据不足一半(36.43 WER,训练量128 vs 36.99,训练量256)。此外,仅微调语音编码器即可将平均词错误率从71%降至32%。结果表明,融合跨说话人共性与说话人特异性模式,能有效提升对代表性不足语音群体的识别性能。

原文摘要 · Abstract (English)

State-of-the-art automatic speech recognition (ASR) models like Whisper, perform poorly on atypical speech, such as that produced by individuals with dysarthria. Past works for atypical speech have mostly investigated fully personalized (or idiosyncratic) models, but modeling strategies that can both generalize and handle idiosyncracy could be more effective for capturing atypical speech. To investigate this, we compare four strategies: (a) $\textit{normative}$ models trained on typical speech (no personalization), (b) $\textit{idiosyncratic}$ models completely personalized to individuals, (c) $\textit{dysarthric-normative}$ models trained on other dysarthric speakers, and (d) $\textit{dysarthric-idiosyncratic}$ models which combine strategies by first modeling normative patterns before adapting to individual speech. In this case study, we find the dysarthric-idiosyncratic model performs better than idiosyncratic approach while requiring less than half as much personalized data (36.43 WER with 128 train size vs 36.99 with 256). Further, we found that tuning the speech encoder alone (as opposed to the LM decoder) yielded the best results reducing word error rate from 71% to 32% on average. Our findings highlight the value of leveraging both normative (cross-speaker) and idiosyncratic (speaker-specific) patterns to improve ASR for underrepresented speech populations.

语音识别失语症个性化建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。