arXiv:2502.02196cs.CVcs.AI2025-02被引 17

用集成学习提升手语跨视角识别准确率

Exploiting Ensemble Learning for Cross-View Isolated Sign Language Recognition

  • 基于多维度视频Swin Transformer构建集成模型
  • 在RGB和RGB-D双赛道均获第三名
  • 适合关注跨视角手势识别的开发者

本文针对WWW 2025举办的跨视角孤立手语识别(CV-ISLR)挑战赛提出解决方案。传统孤立手语识别数据集多为正面视角,而实际场景中摄像头角度多样,导致模型泛化困难。为此,我们探索集成学习的优势,增强模型在多视角下的鲁棒性与泛化能力。方法基于多维视频Swin Transformer架构,通过集成策略实现优异性能。最终方案在基于RGB和基于RGB-D的两个赛道中均取得第三名,验证了该方法在处理跨视角识别挑战中的有效性。代码已开源:https://github.com/Jiafei127/CV_ISLR_WWW2025。

原文摘要 · Abstract (English)

In this paper, we present our solution to the Cross-View Isolated Sign Language Recognition (CV-ISLR) challenge held at WWW 2025. CV-ISLR addresses a critical issue in traditional Isolated Sign Language Recognition (ISLR), where existing datasets predominantly capture sign language videos from a frontal perspective, while real-world camera angles often vary. To accurately recognize sign language from different viewpoints, models must be capable of understanding gestures from multiple angles, making cross-view recognition challenging. To address this, we explore the advantages of ensemble learning, which enhances model robustness and generalization across diverse views. Our approach, built on a multi-dimensional Video Swin Transformer model, leverages this ensemble strategy to achieve competitive performance. Finally, our solution ranked 3rd in both the RGB-based ISLR and RGB-D-based ISLR tracks, demonstrating the effectiveness in handling the challenges of cross-view recognition. The code is available at: https://github.com/Jiafei127/CV_ISLR_WWW2025.

手语识别跨视角集成学习视频Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。