用Whisper模型分析中风患者命名任务语音,提升语言功能自动评估能力
Application of Whisper in Clinical Practice: the Post-Stroke Speech Assessment during a Naming Task
- 微调Whisper模型以提升中风患者单字发音转录准确率
- 微调后词错误率降低87.72%(健康)和71.22%(患者)
- 模型可有效预测语音质量,适合临床语言康复评估场景
中风后语言障碍的详细评估仍是一项认知复杂且依赖医生的任务,限制了及时、可扩展的诊断。自动语音识别(ASR)基础模型为通过智能系统辅助人工评估提供了可能,但其在语言障碍语境下的有效性尚不明确。本研究评估了Whisper——一种先进ASR基础模型——在常用图片命名任务中对中风患者语音的转录与分析能力。我们同时考察了逐字转录准确率及模型支持下游语言功能预测的能力,后者对中风预后具有重要意义。结果表明,基线Whisper模型在单字语音上表现不佳。然而,微调后显著提升了转录准确率(健康语音词错误率下降87.72%,患者语音下降71.22%)。此外,模型学习到的表征可实现对语音质量的精准预测(健康组平均F1 Macro为0.74,患者组为0.75)。但在未见过的(TORGO)数据集上评估显示泛化能力有限,凸显Whisper无法在跨域临床语音上实现零样本单字转录,强调需针对特定临床人群适配模型。尽管跨领域泛化仍存挑战,这些发现表明,经适当微调的基础模型有望推动中风相关语言障碍的自动化评估与康复进展。
原文摘要 · Abstract (English)
Detailed assessment of language impairment following stroke remains a cognitively complex and clinician-intensive task, limiting timely and scalable diagnosis. Automatic Speech Recognition (ASR) foundation models offer a promising pathway to augment human evaluation through intelligent systems, but their effectiveness in the context of speech and language impairment remains uncertain. In this study, we evaluate whether Whisper, a state-of-the-art ASR foundation model, can be applied to transcribe and analyze speech from patients with stroke during a commonly used picture-naming task. We assess both verbatim transcription accuracy and the model's ability to support downstream prediction of language function, which has major implications for outcomes after stroke. Our results show that the baseline Whisper model performs poorly on single-word speech utterances. Nevertheless, fine-tuning Whisper significantly improves transcription accuracy (reducing Word Error Rate by 87.72% in healthy speech and 71.22% in speech from patients). Further, learned representations from the model enable accurate prediction of speech quality (average F1 Macro of 0.74 for healthy, 0.75 for patients). However, evaluations on an unseen (TORGO) dataset reveal limited generalizability, highlighting the inability of Whisper to perform zero-shot transcription of single-word utterances on out-of-domain clinical speech and emphasizing the need to adapt models to specific clinical populations. While challenges remain in cross-domain generalization, these findings highlight the potential of foundation models, when appropriately fine-tuned, to advance automated speech and language assessment and rehabilitation for stroke-related impairments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。