arXiv:2505.02304cs.CLcs.CV2025-05被引 1

用大模型生成手语描述并对比学习,提升手语识别准确率。

Generative Sign-description Prompts with Multi-positive Contrastive Learning for Sign Language Recognition

  • 用大模型生成多层级手语描述,结合检索增强生成技术。
  • 在中英文手语数据集上分别达到97.1%和97.07%准确率。
  • 适合做跨语言手语识别与无障碍通信系统开发。

手语识别(SLR)因手势与非手势信号的复杂性,难以获得精准标注。本文首次将生成式大语言模型(LLM)引入SLR任务,提出一种基于多正例对比学习的生成式手语描述提示方法(GSP-MC)。该方法利用领域专用大模型与检索增强生成(RAG),通过多步提示工程和专家验证的手语语料库,生成精确的多部分描述。GSP-MC采用双编码器架构,通过概率匹配实现骨架特征与全局、同义词及部件级文本描述的双向对齐。结合全局与部件级损失,优化KL散度,确保所有文本-骨架对间稳健对齐,同时捕捉手语整体语义与细节动态。实验表明,在中文SLR500(准确率97.1%)和土耳其AUTSL数据集(97.07%)上均达到当前最优性能,展现出良好的跨语言适用性,具备推动包容性通信技术发展的潜力。

原文摘要 · Abstract (English)

Sign language recognition (SLR) faces fundamental challenges in creating accurate annotations due to the inherent complexity of simultaneous manual and non-manual signals. To the best of our knowledge, this is the first work to integrate generative large language models (LLMs) into SLR tasks. We propose a novel Generative Sign-description Prompts Multi-positive Contrastive learning (GSP-MC) method that leverages retrieval-augmented generation (RAG) with domain-specific LLMs, incorporating multi-step prompt engineering and expert-validated sign language corpora to produce precise multipart descriptions. The GSP-MC method also employs a dual-encoder architecture to bidirectionally align hierarchical skeleton features with multiple text descriptions (global, synonym, and part level) through probabilistic matching. Our approach combines global and part-level losses, optimizing KL divergence to ensure robust alignment across all relevant text-skeleton pairs while capturing both sign-level semantics and detailed part dynamics. Experiments demonstrate state-of-the-art performance against existing methods on the Chinese SLR500 (reaching 97.1%) and Turkish AUTSL datasets (97.07% accuracy). The method's cross-lingual effectiveness highlight its potential for developing inclusive communication technologies.

手语识别大模型对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。