arXiv:2609.03695cs.CV2026-09

让手语字典检索突破签名者差异,实现跨语言零样本识别

SignSeek: Learning Transferable Representations for Sign Dictionary Retrieval

论文配图:SignSeek: Learning Transferable Representations for Sign Dictionary Retrieval
图 1 · 摘自论文原文
  • 通过关键动作部位掩码对比学习,提升手语表示的泛化能力
  • 在多个数据集上达到新纪录,零样本迁移至英式手语仍领先
  • 适用于手语检索、孤立手势识别和字幕对齐,无需微调

手语词典对学习者至关重要,但仅凭查询视频自动检索手语仍具挑战,因签名者间存在自然差异。现有手语表征学习方法多用于闭集识别,生成的嵌入无法泛化到开放集、与签名者无关的检索场景。本文提出SignSeek,通过显著性引导的动作部位掩码进行对比学习:对比目标使同一词义的手语在不同签名者间对齐;艺术化显著性引导掩码(ASGM)定位每种手语最核心的动作部位。这驱动两个互补目标——掩码对比对齐(MAC)损失,仅通过单一动作部位感知手语;掩码预测(MAP)损失,从时空上下文重建该部位的潜在表示。模型在26.6万条样本(约5700个词义)上预训练,覆盖多种手语语言,在不需下游微调的情况下,于ASL-Citizen、WLASL和NMFs-CSL数据集上实现跨语料库检索新基准。尤为突出的是,其在完全未见的英式手语(BSL)上实现零样本泛化,表现超越专门针对BSL训练的方法,并可无缝迁移到孤立手语识别与字幕对齐任务,优于先前基于骨架的方法。

原文摘要 · Abstract (English)

Sign language dictionaries are essential resources for sign language learners, yet automatically retrieving a sign from a dictionary, given only a query video, remains a challenging problem due to the natural variability between signers. Existing sign representation learning methods are built for closed-set recognition, producing embeddings that do not generalise to the open-set, signer-independent setting that retrieval demands. \textbf{SignSeek} closes this gap by contrastively learning sign representations with saliency-guided articulator masking. A contrastive objective aligns same-gloss signs across signers, while our Articulator Saliency-Guided Masking (ASGM) pinpoints the single most critical articulator per sign. This drives two complementary objectives, a masked contrastive alignment (MAC) loss that sees the sign through a single articulator and a masked prediction (MAP) loss that reconstructs it in latent space from the surrounding spatio-temporal context. Pretrained on 266K samples ($\sim$5,700 glosses) across multiple sign languages, \textbf{SignSeek} sets a new state-of-the-art performance in cross-corpus retrieval on ASL-Citizen, WLASL, and NMFs-CSL without any downstream fine-tuning. Strikingly, it achieves zero-shot generalisation to an entirely unseen British Sign Language (BSL), surpassing methods explicitly trained on BSL, and transfers seamlessly to isolated sign recognition and subtitle alignment, outperforming prior skeleton-based methods.

手语识别对比学习零样本迁移字典检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。