arXiv:2411.00980cs.CLcs.HC2024-11被引 4

针对语音障碍者优化语音识别,提升远程医疗沟通质量

Enhancing AAC Software for Dysarthric Speakers in e-Health Settings: An Evaluation Using TORGO

  • 用算法消除数据集中的语音重叠问题,改善模型训练公平性
  • 改进后模型对轻重度构音障碍者识别错误率显著降低
  • 结合大语言模型进行多模态纠错,适合医疗AI研究者参考

脑瘫(CP)和肌萎缩侧索硬化症(ALS)患者常因发音困难导致构音障碍,影响医疗沟通质量。现有先进语音识别(ASR)系统如Whisper和Wav2vec2.0在识别此类异常语音时表现不佳,主要因训练数据不足。本文针对英文构音障碍语音识别评估常用数据集TORGO中存在的提示重叠问题,提出算法予以解决。处理后,尽管使用SOTA ASR模型,轻度与重度构音障碍者的词错误率仍极高。为此,本文进一步研究了n-gram语言模型及基于大语言模型(LLM)的多模态生成式纠错方法(如Whispering-LLaMA),用于第二阶段语音识别。结果表明,当前技术仍难以满足无障碍医疗通信需求,亟需进一步改进。

原文摘要 · Abstract (English)

Individuals with cerebral palsy (CP) and amyotrophic lateral sclerosis (ALS) frequently face challenges with articulation, leading to dysarthria and resulting in atypical speech patterns. In healthcare settings, communication breakdowns reduce the quality of care. While building an augmentative and alternative communication (AAC) tool to enable fluid communication we found that state-of-the-art (SOTA) automatic speech recognition (ASR) technology like Whisper and Wav2vec2.0 marginalizes atypical speakers largely due to the lack of training data. Our work looks to leverage SOTA ASR followed by domain specific error-correction. English dysarthric ASR performance is often evaluated on the TORGO dataset. Prompt-overlap is a well-known issue with this dataset where phrases overlap between training and test speakers. Our work proposes an algorithm to break this prompt-overlap. After reducing prompt-overlap, results with SOTA ASR models produce extremely high word error rates for speakers with mild and severe dysarthria. Furthermore, to improve ASR, our work looks at the impact of n-gram language models and large-language model (LLM) based multi-modal generative error-correction algorithms like Whispering-LLaMA for a second pass ASR. Our work highlights how much more needs to be done to improve ASR for atypical speakers to enable equitable healthcare access both in-person and in e-health settings.

语音识别医疗AI构音障碍大模型纠错

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。