arXiv:2509.13145cs.SD2025-09

用多模态大模型实现语音康复的实时发音反馈

UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model

  • 融合舌部超声与语音信号,实现发音动作精准分析
  • 构建高质量音-超声对话数据集,提升临床适应性
  • 适合语言障碍康复者及治疗师使用,支持交互式指导

言语康复对中风等神经损伤导致的言语障碍至关重要。传统人工和计算机辅助系统在实时可及性和发音动作反馈方面存在局限。近年来,多模态大语言模型(MLLMs)在医疗领域展现出巨大潜力,尤其体现在自适应评估与治疗反馈能力上。然而,发音信息获取与融合不足、发音器官运动轨迹解析不充分以及领域专用数据集匮乏等问题,制约了其在言语康复中的应用。为此,我们提出一种基于MLLM的言语康复辅助系统,利用舌部超声成像与语音信号,提供精确、交互式的发音动作反馈。我们构建了一个高质量的领域专用数据集,包含超声-语音对话对,用于模型微调以增强临床适应性。此外,本方法设计了超声视频与语音信号的时空融合训练策略,实现精细化发音异常分析,并生成可操作的反馈建议。实验结果表明,该模型在发音分析与临床评估中均表现有效。

原文摘要 · Abstract (English)

Speech therapy is essential for rehabilitating speech disorders caused by neurological impairments such as stroke. However, traditional manual and computer-assisted systems are limited in real-time accessibility and articulatory motion feedback. Recent advances in multimodal large language models (MLLMs) have demonstrated significant potential in healthcare, especially through their adaptive assessment and therapeutic feedback capabilities. Nevertheless, challenges including insufficient acquisition and fusion of articulatory information, inadequate parsing of articulatory organ motion trajectories, and the scarcity of domain-specific datasets hinder the application of MLLMs in speech therapy. To address these limitations, we propose an MLLM-based speech rehabilitation assistance system that leverages ultrasound tongue imaging and speech signals to deliver precise, interactive articulatory feedback. We construct a high-quality domain-specific dataset comprising ultrasound-speech dialogue pairs. This dataset facilitates fine-tuning to enhance the model's clinical adaptability. Furthermore, our method develops spatiotemporal fusion training strategy of ultrasound videos and speech signals, enabling fine-grained articulatory impairment analysis and ultimately generating actionable feedback. Experimental results demonstrate the effectiveness of our model in articulatory analysis and clinical assessment.

语音康复多模态模型超声成像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。