arXiv:2508.08684cs.CLcs.CY2025-08被引 2

评估主流语音识别模型在老年患者临床对话中的表现

Out of the Box, into the Clinic? Evaluating State-of-the-Art ASR for Clinical Applications for Older Adults

  • 对比通用多语言与专为老年人优化的荷兰语语音模型
  • 通用模型表现优于微调模型,且压缩后可提升速度与精度平衡
  • 发现高错误率输入场景,为临床应用提供改进方向

语音控制界面可在临床环境中辅助老年人,例如通过聊天机器人实现交互,但针对少数群体的可靠语音识别仍是瓶颈。本研究评估了当前最先进的语音识别(ASR)模型在荷兰老年群体语言使用上的表现,这些用户与专为老年群体设计的Welzijn.AI聊天机器人互动。我们对比了通用多语言模型与针对老年人荷兰语微调的模型,并考察其处理速度。结果表明,通用多语言模型优于微调模型,说明现代ASR模型具备良好的泛化能力,无需特定适配即可在真实场景中使用。此外,裁剪通用模型有助于在准确率与速度间取得更好平衡。然而,研究也识别出导致高词错误率的输入模式,并将其置于实际应用场景中进行分析。

原文摘要 · Abstract (English)

Voice-controlled interfaces can support older adults in clinical contexts -- with chatbots being a prime example -- but reliable Automatic Speech Recognition (ASR) for underrepresented groups remains a bottleneck. This study evaluates state-of-the-art ASR models on language use of older Dutch adults, who interacted with the Welzijn.AI chatbot designed for geriatric contexts. We benchmark generic multilingual ASR models, and models fine-tuned for Dutch spoken by older adults, while also considering processing speed. Our results show that generic multilingual models outperform fine-tuned models, which suggests recent ASR models can generalise well out of the box to real-world datasets. Moreover, our results indicate that truncating generic models is helpful in balancing the accuracy-speed trade-off. Nonetheless, we also find inputs which cause a high word error rate and place them in context.

语音识别老年医疗临床应用多语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。