arXiv:2609.03788cs.CVcs.CL2026-09

无需词典标注,用视频描述实现日语手语的开放词汇识别。

A Reverse Sign Language Dictionary: Open-Vocabulary Sign Recognition from Continuous Signing via Video Captioning and Description Retrieval

论文配图:A Reverse Sign Language Dictionary: Open-Vocabulary Sign Recognition from Continuous Signing via Video Captioning and Description Retrieval
图 1 · 摘自论文原文
  • 用视觉语言模型将手语片段转为自由描述,再通过多语言编码器匹配词库。
  • 已见手势识别准确率提升至49%(未微调仅4.5%),未见手势也从11.5%升至21.0%。
  • 首次实现无词典标注的日语手语开放词汇检索,适合手语技术研究者使用。

孤立手语识别通常基于封闭词表分类,无法泛化到训练中未见的手势,且依赖词典标注。本文提出一种反向手语词典:从连续手语中提取手势片段,先用开权重视觉语言模型将其生成自由形式的动作描述,再通过多语言句子编码器从目标描述词库中检索最相近条目。该方法无需词典监督,支持开放词汇。在1300个日语手语对话数据集的手势片段上测试,微调后已见类别的前10名检索率从4.5%提升至49%,接近标准监督分类器(I3D)表现;未见类别检索率从11.5%提升至21.0%(p=0.0094),而传统分类器无法参与此任务。句编码器对同义描述的召回率接近100%,表明差距主要来自描述生成质量。这是首个无需词典标注、基于描述的连续手语开放词汇检索,也是首个针对日语手语的工作。

原文摘要 · Abstract (English)

Isolated Sign Language Recognition (ISLR) is conventionally cast as closed-set classification over gloss labels, which cannot generalize to signs unseen in training and ties every deployment to a gloss-annotated lexicon. We instead recognize signs extracted from continuous signing by (1) captioning a sign-level clip into a free-form procedural description of the articulation with an open-weight vision-language model, and (2) retrieving the closest entry from a vocabulary of target descriptions with a multilingual sentence encoder: a reverse sign language dictionary that needs no gloss supervision and admits an open vocabulary. On 1,300 sign-level segments from a Japanese Sign Language (JSL) dialogue corpus annotated with procedural descriptions (against a 2% top-10 chance floor over the 503-entry target vocabulary), fine-tuning the captioner substantially improves seen-class retrieval: language and vision tower fine-tuning raises top-10 retrieval on seen classes from 4.5% (untrained) to 49%, becoming statistically indistinguishable from a standard supervised closed-set classifier (I3D) on two of the three test sets where a closed-set classifier can be evaluated at all. More importantly, unseen-class retrieval also improves significantly over the untrained pipeline (11.5% -> 21.0% top-10, p=0.0094), a regime in which the closed-set classifier cannot participate. A matcher-side empirical upper-bound analysis shows the sentence encoder already recovers close to 100% of paraphrased gold descriptions, locating a gap in captioning quality that we aim to address in future work. To our knowledge this is the first description-based, open-vocabulary sign lookup from continuous signing without gloss supervision, and the first for JSL.

手语识别开放词汇视频描述多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。