arXiv:2603.19523cs.CV2026-03

构建大规模手语指拼数据集,提升连续手语中字母识别准确率

Recognising BSL Fingerspelling in Continuous Signing Sequences

  • 基于迭代标注框架构建23,000条连续手语指拼数据
  • 通过考虑双手协同与口部动作,将字符错误率降低50%
  • 适合手语识别与自动化标注研究者参考

指拼是英国手语(BSL)的关键组成部分,用于拼写专有名词、术语及无固定词汇的手势词。由于手势速度极快且母语使用者常省略字母,指拼识别极具挑战性。现有BSL指拼数据集规模小或存在时间与字母层面的标注误差。本文提出新数据集FS23K,采用迭代标注框架构建,涵盖23,000条连续手语序列。同时设计一种新型识别模型,显式建模双臂协作与口部动作线索。在精修标注数据上,该方法将字符错误率(CER)较先前最优结果降低50%。实验验证了方法有效性,展现了其在手语理解与可扩展自动标注中的应用潜力。项目页面见https://taeinkwon.com/projects/fs23k/。

原文摘要 · Abstract (English)

Fingerspelling is a critical component of British Sign Language (BSL), used to spell proper names, technical terms, and words that lack established lexical signs. Fingerspelling recognition is challenging due to the rapid pace of signing and common letter omissions by native signers, while existing BSL fingerspelling datasets are either small in scale or temporally and letter-wise inaccurate. In this work, we introduce a new large-scale BSL fingerspelling dataset, FS23K, constructed using an iterative annotation framework. In addition, we propose a fingerspelling recognition model that explicitly accounts for bi-manual interactions and mouthing cues. As a result, with refined annotations, our approach halves the character error rate (CER) compared to the prior state of the art on fingerspelling recognition. These findings demonstrate the effectiveness of our method and highlight its potential to support future research in sign language understanding and scalable, automated annotation pipelines. The project page can be found at https://taeinkwon.com/projects/fs23k/.

手语识别指拼数据集双臂协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。