用跨模态一致性提升手语识别,融合视觉与姿态信息
Cross-Modal Consistency Learning for Sign Language Recognition
- 通过自监督对比学习对齐RGB与姿态模态特征空间
- 在四个基准上达到领先性能,最高提升6.2%准确率
- 适合做手语识别或多模态表示学习的研究者参考
预训练已被证明能有效提升孤立手语识别(ISLR)的性能。现有方法仅关注紧凑的姿态数据,虽消除背景干扰但语义线索不足;而直接从原始RGB视频学习又受非手语视觉特征影响。为此,我们提出跨模态一致性学习框架CCL-SLR,基于自监督预训练利用RGB与姿态模态间的跨模态一致性。首先,通过单模态与跨模态对比学习,逐步对齐两模态特征空间,提取一致的手语表征。其次,引入运动保持掩码(MPM)和语义正样本挖掘(SPM)技术,分别从数据增强和样本相似性角度增强跨模态一致性。在四个ISLR基准上的大量实验表明,CCL-SLR表现优异,验证了其有效性。代码将公开。
原文摘要 · Abstract (English)
Pre-training has been proven to be effective in boosting the performance of Isolated Sign Language Recognition (ISLR). Existing pre-training methods solely focus on the compact pose data, which eliminates background perturbation but inevitably suffers from insufficient semantic cues compared to raw RGB videos. Nevertheless, learning representation directly from RGB videos remains challenging due to the presence of sign-independent visual features. To address this dilemma, we propose a Cross-modal Consistency Learning framework (CCL-SLR), which leverages the cross-modal consistency from both RGB and pose modalities based on self-supervised pre-training. First, CCL-SLR employs contrastive learning for instance discrimination within and across modalities. Through the single-modal and cross-modal contrastive learning, CCL-SLR gradually aligns the feature spaces of RGB and pose modalities, thereby extracting consistent sign representations. Second, we further introduce Motion-Preserving Masking (MPM) and Semantic Positive Mining (SPM) techniques to improve cross-modal consistency from the perspective of data augmentation and sample similarity, respectively. Extensive experiments on four ISLR benchmarks show that CCL-SLR achieves impressive performance, demonstrating its effectiveness. The code will be released to the public.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。