arXiv:2604.07606cs.CV2026-04中稿 · CVPR被引 1

用AI自动生成手语标注,解决高质量数据稀缺难题。

Bootstrapping Sign Language Annotations with Sign Language Models

  • 结合手语识别模型与小样本大模型,自动推测手语视频中的时间区间和标签。
  • 在FSBoard上达到6.7%词错误率,在ASL Citizen上达74%准确率,性能领先。
  • 为手语研究提供近500段人工标注视频和超300小时伪标注数据,适合研究者使用。

基于AI的手语翻译受限于高质量标注数据的缺乏。新发布的ASL STEM Wiki和FLEURS-ASL数据集包含专业译员录制的数百小时视频,但仅部分完成标注,主要因大规模标注成本过高。本文提出一种伪标注流程,输入为手语视频与对应英文文本,输出包括词义、指拼词和手语分类符的时间区间预测。该流程利用手指拼写识别器和孤立手语识别器(ISR)的稀疏预测结果,结合K-Shot大语言模型方法生成候选标注。为此,我们构建了简单有效的基线手指拼写与ISR模型,在FSBoard上实现6.7%词错误率,在ASL Citizen数据集上达到74%的顶1准确率。为验证并提供金标准,专业译员对近500段来自ASL STEM Wiki的视频进行了序列级词义标注,涵盖词义、分类符和指拼词。这些人工标注及超过300小时的伪标注数据将作为补充材料发布。

原文摘要 · Abstract (English)

AI-driven sign language interpretation is limited by a lack of high-quality annotated data. New datasets including ASL STEM Wiki and FLEURS-ASL contain professional interpreters and 100s of hours of data but remain only partially annotated and thus underutilized, in part due to the prohibitive costs of annotating at this scale. In this work, we develop a pseudo-annotation pipeline that takes signed video and English as input and outputs a ranked set of likely annotations, including time intervals, for glosses, fingerspelled words, and sign classifiers. Our pipeline uses sparse predictions from our fingerspelling recognizer and isolated sign recognizer (ISR), along with a K-Shot LLM approach, to estimate these annotations. In service of this pipeline, we establish simple yet effective baseline fingerspelling and ISR models, achieving state-of-the-art on FSBoard (6.7% CER) and on ASL Citizen datasets (74% top-1 accuracy). To validate and provide a gold-standard benchmark, a professional interpreter annotated nearly 500 videos from ASL STEM Wiki with sequence-level gloss labels containing glosses, classifiers, and fingerspelling signs. These human annotations and over 300 hours of pseudo-annotations are being released in supplemental material.

手语识别伪标注数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。