arXiv:2502.20171cs.CLcs.CV2025-02被引 1

用少量样本实现跨语言手语识别,提升技术可扩展性

Representing Signs as Signs: One-Shot ISLR to Facilitate Functional Sign Language Technologies

  • 基于关键特征预训练模型,通过向量搜索快速识别新手语
  • 在10235个陌生语言手语上达50.8%的一次性识别准确率
  • 面向聋哑群体设计,适合动态词汇库的实用场景

孤立手语识别(ISLR)对构建可扩展的手语技术至关重要,但现有方法多依赖特定语言,难以泛化。为此,我们提出一种一次性学习方法,可在不同语言间迁移并适应不断变化的词汇表。该方法先预训练模型以关键特征嵌入手语,再通过密集向量搜索实现对未见手语的快速精准识别。在包含10,235个不同语言手语的大型词典上,取得50.8%的一次性平均排名倒数(MRR)的最优表现。该方法在多种语言和支持集下均具鲁棒性,提供可扩展、可适应的解决方案。研究由聋哑及听力障碍(DHH)群体共同参与,契合真实应用需求,推动了规模化手语识别的发展。

原文摘要 · Abstract (English)

Isolated Sign Language Recognition (ISLR) is crucial for scalable sign language technology, yet language-specific approaches limit current models. To address this, we propose a one-shot learning approach that generalises across languages and evolving vocabularies. Our method involves pretraining a model to embed signs based on essential features and using a dense vector search for rapid, accurate recognition of unseen signs. We achieve state-of-the-art results, including 50.8% one-shot MRR on a large dictionary containing 10,235 unique signs from a different language than the training set. Our approach is robust across languages and support sets, offering a scalable, adaptable solution for ISLR. Co-created with the Deaf and Hard of Hearing (DHH) community, this method aligns with real-world needs, and advances scalable sign language recognition.

手语识别一次学习跨语言可扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。