arXiv:2503.21785eess.AScs.SD2025-03被引 2

用大模型零样本识别手语手势,不需训练就能提升听障者沟通效率

Lend a Hand: Semi Training-Free Cued Speech Recognition via MLLM-Driven Hand Modeling for Barrier-free Communication

  • 利用多模态大模型零样本识别手部动作,免去复杂训练
  • 在混合数据集上达到领先准确率,听障者数据也表现优异
  • 适合无障碍交流系统研发者,尤其关注低资源场景应用

有声手语(Cued Speech, CS)是一种融合口读与手势编码的视觉沟通系统,旨在提升听障人士的交流效率。自动有声手语识别(ACSR)通过人工智能自动识别手势与口部动作并转为文本。然而,以往方法依赖复杂的融合模块和训练技术,且因数据稀缺,手部特征提取与识别建模效果不佳,严重制约了性能。为此,本文创新性地探索多模态大模型(MLLM)在识别CS手形与位置上的能力,提出无需训练的半训练自由框架STF-ACSR。该方法通过中文有声手语提示模块(CCSPM),实现零样本手势识别,结合无训练关键帧筛选与定制化提示工程。随后通过极简融合模块(MFM)将结果整合至口读模型,显著提升识别效果。研究还补充了8位听障者的有声手语数据,构建新混合数据集。大量实验表明,STF-ACSR在正常及听障者数据上均显著优于现有方法。代码与模型已开源。

原文摘要 · Abstract (English)

Cued Speech (CS) is an innovative visual communication system that integrates lip-reading with hand coding, designed to enhance effective communication for individuals with hearing impairments. Automatic CS Recognition (ACSR) refers to the AI-driven process of automatically recognizing hand gestures and lip movements in CS, converting them into text. However, previous work often relies on complex fusion modules and training techniques. Additionally, due to the limited amount of data in CS, the extraction of hand features, as well as recognition modeling, has consistently been subpar, significantly limiting the effectiveness of ACSR. To address this issue, we have innovatively explored the capabilities of Multimodal large language models (MLLMs) in recognizing hand shapes and positions in CS. More precisely, we propose a new Semi Training-Free paradigm for ACSR, named STF-ACSR. This approach leverages zero-shot recognition of hand movements through the Chinese CS Prompt Module (CCSPM), which equipped a training-free keyframe filtering and customized prompt engineering based on MLLM. It then integrates the recognition results into the lip-reading model using a Minimalist Fusion Module (MFM), effectively achieving superior recognition results. Furthermore, specifically for this study, we have supplemented the existing dataset of 6 normal hearing CS cuers by recording additional data from 8 cuers with hearing impairments, resulting in a new mixed dataset. Extensive experiments have demonstrated that STF-ACSR significantly outperforms previous methods on both normal and hearing-impaired data. Implementation and checkpoints are available at https://github.com/DennisHgj/STF_ACSR.

语音识别手语识别大模型无障碍

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。