arXiv:2502.13983eess.AScs.AI2025-02被引 3

用手势增强语音识别,帮助语言障碍者更准确沟通

Gesture-Aware Zero-Shot Speech Recognition for Patients with Language Disorders

  • 融合手势与语音的多模态识别,零样本学习提升适应性
  • 引入手势信息后语义理解显著提升,弥补语音缺陷
  • 适合语言障碍人群,推动无障碍人机交互发展

语言障碍者因语言处理和理解能力受限,难以使用依赖语音识别(ASR)的语音助手。尽管当前ASR已能处理言语不流畅问题,但对非语言交流方式如手势的关注仍不足。本研究提出一种基于多模态大语言模型的零样本手势感知语音识别系统,旨在解读仅靠语音无法捕捉的视觉隐含语义。实验表明,加入手势信息可显著提升语义理解能力。该方法有助于开发更符合语言障碍者需求的高效通信技术。

原文摘要 · Abstract (English)

Individuals with language disorders often face significant communication challenges due to their limited language processing and comprehension abilities, which also affect their interactions with voice-assisted systems that mostly rely on Automatic Speech Recognition (ASR). Despite advancements in ASR that address disfluencies, there has been little attention on integrating non-verbal communication methods, such as gestures, which individuals with language disorders substantially rely on to supplement their communication. Recognizing the need to interpret the latent meanings of visual information not captured by speech alone, we propose a gesture-aware ASR system utilizing a multimodal large language model with zero-shot learning for individuals with speech impairments. Our experiment results and analyses show that including gesture information significantly enhances semantic understanding. This study can help develop effective communication technologies, specifically designed to meet the unique needs of individuals with language impairments.

语音识别手势识别无障碍设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。