通过语义增强预训练,提升骨骼数据在手语理解中的表现。
Sigma: Semantically Informative Pre-training for Skeleton-based Sign Language Understanding
- 引入语义感知的早期融合机制,加强视觉与文本模态交互。
- 采用分层对齐学习,同时捕捉细粒度动作和高层语义关系。
- 适合关注手语识别、翻译与跨模态学习的研究者。
预训练在手语理解(SLU)任务中已被证明能有效学习可迁移特征。近年来,基于骨骼的方法因能鲁棒处理个体与背景差异,不受外观或环境因素影响而受到越来越多关注。现有方法仍面临三大挑战:1)语义关联弱,模型通常仅捕捉骨骼数据中的低级运动模式,难以关联到语言意义;2)局部细节与全局上下文失衡,模型或过度关注细微动作,或忽略它们以追求整体结构;3)跨模态学习效率低,构建语义对齐的多模态表示仍具挑战。为此,我们提出Sigma,一个统一的骨骼基手语理解框架,包含:1)语义感知的早期融合机制,促进视觉与文本模态深度交互,用语言上下文丰富视觉特征;2)分层对齐学习策略,联合最大化不同层级配对特征间的一致性,有效捕捉细粒度与高层语义关系;3)融合对比学习、文本匹配与语言建模的统一预训练框架,提升语义一致性和泛化能力。Sigma在多个基准上实现了孤立手语识别、连续手语识别和无词签手语翻译的新最优性能,涵盖不同手语与口语,验证了语义信息丰富的预训练价值及骨骼数据作为独立解法的有效性。
原文摘要 · Abstract (English)
Pre-training has proven effective for learning transferable features in sign language understanding (SLU) tasks. Recently, skeleton-based methods have gained increasing attention because they can robustly handle variations in subjects and backgrounds without being affected by appearance or environmental factors. Current SLU methods continue to face three key limitations: 1) weak semantic grounding, as models often capture low-level motion patterns from skeletal data but struggle to relate them to linguistic meaning; 2) imbalance between local details and global context, with models either focusing too narrowly on fine-grained cues or overlooking them for broader context; and 3) inefficient cross-modal learning, as constructing semantically aligned representations across modalities remains difficult. To address these, we propose Sigma, a unified skeleton-based SLU framework featuring: 1) a sign-aware early fusion mechanism that facilitates deep interaction between visual and textual modalities, enriching visual features with linguistic context; 2) a hierarchical alignment learning strategy that jointly maximises agreements across different levels of paired features from different modalities, effectively capturing both fine-grained details and high-level semantic relationships; and 3) a unified pre-training framework that combines contrastive learning, text matching and language modelling to promote semantic consistency and generalisation. Sigma achieves new state-of-the-art results on isolated sign language recognition, continuous sign language recognition, and gloss-free sign language translation on multiple benchmarks spanning different sign and spoken languages, demonstrating the impact of semantically informative pre-training and the effectiveness of skeletal data as a stand-alone solution for SLU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。