用视觉相似性匹配手语字典视频与连续手语视频
SignMatch: Matching Dictionary Signs to Continuous Sign Language Video

- 构建原型结构的符号嵌入空间,学习手形与动作特征
- 在三个手语数据集上实现跨语言高精度匹配
- 无需特定任务标注,可直接用于未见手语的识别
本文旨在将手语字典视频与连续手语视频中的对应手势进行匹配,匹配依据仅为视觉相似性——即手形与相对于身体的动作。为此,我们从标注了手势的连续视频中学习一种原型结构的符号嵌入空间,每个可学习的原型对应一个手语类别。孤立的手语字典视频被映射到该嵌入空间,从而实现字典示例与连续手语实例之间的匹配。该设计支持通过嵌入相似性实现直接字典引导的手语匹配,并可仅用字典示例自然扩展至未见手语。在ASL-Citizen字典检索、ChaLearn OSLWL字典到连续手语匹配,以及使用BOBSL的CSLR2评估自动手语标注的实验中,该方法展现出强泛化能力。在无基准特定监督的情况下,所学表示能有效跨美国、英国和西班牙手语迁移,在所有三个基准上均优于以往方法。
原文摘要 · Abstract (English)
The objective of this paper is to match dictionary sign videos to corresponding signs in continuous signing videos, where a match is defined by the visual similarity alone - the handshape and motion relative to the body. To achieve this, we learn a prototype-structured sign embedding space from continuous video annotated with signs, where each learnable prototype corresponds to a sign class. Isolated dictionary videos are then mapped into this sign space, enabling the matching between dictionary exemplars and continuous sign instances. This design supports direct dictionary-guided sign matching through embedding similarity and naturally extends to unseen signs using only dictionary exemplars. Experiments on ASL-Citizen dictionary retrieval, ChaLearn OSLWL dictionary-to-continuous sign matching, and using BOBSL's CSLR2 evaluation for automatic sign annotation demonstrate strong generalisation across datasets, tasks and sign languages. Without benchmark-specific supervision, the learned representation transfers effectively across American, British, and Spanish Sign Languages, outperforming prior methods on all three benchmarks. Project page: https://www.robots.ox.ac.uk/~vgg/research/signmatch/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。