arXiv:2506.20947cs.CVcs.MM2025-06

用分层子动作树融合手语词知识,提升连续手语识别精度。

Hierarchical Sub-action Tree for Continuous Sign Language Recognition

  • 构建分层子动作树结构,逐级对齐视觉与文本模态。
  • 在四个数据集上准确率显著提升,最高达8.2%增益。
  • 适合手语识别、跨模态学习研究者参考。

连续手语识别(CSLR)旨在将未修剪的视频转录为词素(glosses),即文本词汇。现有研究指出,缺乏大规模数据集和精确标注已成为制约CSLR发展的瓶颈。尽管已有工作提出跨模态方法以对齐视觉与文本信息,但通常仅从词素中提取文本特征,未能充分挖掘其知识。本文提出分层子动作树(HST)模型,称为HST-CSLR,高效融合词素知识与视觉表征学习。通过引入大语言模型中的词素特定知识,增强文本信息利用效率。具体地,我们构建用于文本信息表示的分层子动作树,逐步对齐视觉与文本模态,并借助树结构降低计算复杂度。此外,引入对比对齐增强机制,进一步弥合双模态差距。在四个数据集(PHOENIX-2014、PHOENIX-2014T、CSL-Daily、Sign Language Gesture)上的实验验证了所提方法的有效性。

原文摘要 · Abstract (English)

Continuous sign language recognition (CSLR) aims to transcribe untrimmed videos into glosses, which are typically textual words. Recent studies indicate that the lack of large datasets and precise annotations has become a bottleneck for CSLR due to insufficient training data. To address this, some works have developed cross-modal solutions to align visual and textual modalities. However, they typically extract textual features from glosses without fully utilizing their knowledge. In this paper, we propose the Hierarchical Sub-action Tree (HST), termed HST-CSLR, to efficiently combine gloss knowledge with visual representation learning. By incorporating gloss-specific knowledge from large language models, our approach leverages textual information more effectively. Specifically, we construct an HST for textual information representation, aligning visual and textual modalities step-by-step and benefiting from the tree structure to reduce computational complexity. Additionally, we impose a contrastive alignment enhancement to bridge the gap between the two modalities. Experiments on four datasets (PHOENIX-2014, PHOENIX-2014T, CSL-Daily, and Sign Language Gesture) demonstrate the effectiveness of our HST-CSLR.

手语识别跨模态分层结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。