arXiv:2509.10710cs.CV2025-09中稿 · GCPR 2025

用可提示分割融合视觉与姿态信息,提升手语识别精度

SegSLR: Promptable Video Segmentation for Isolated Sign Language Recognition

  • 通过可提示零样本视频分割,精准保留手部形状和朝向细节
  • 在ChaLearn249 IsoGD数据集上超越现有最佳方法,准确率显著提升
  • 适合关注手语识别中多模态融合与细粒度特征提取的研究者

孤立手语识别(ISLR)通常依赖于RGB图像或签发者姿态信息。然而,将这两种模态结合时,常因边界框等不精确表示而丢失关键细节,如手形和朝向。为此,我们提出ISLR系统SegSLR,通过可提示的零样本视频分割融合RGB与姿态信息。基于姿态信息提供的手部与身体粗略定位,系统对视频进行分割以保持所有相关形状信息。随后,分割结果聚焦RGB数据处理,仅针对手语识别最相关的身体部位。该方法有效融合了两种模态信息。在复杂度高的ChaLearn249 IsoGD数据集上的评估显示,SegSLR优于现有最优方法。消融实验表明,聚焦签发者身体与手部显著提升性能,验证了设计合理性。

原文摘要 · Abstract (English)

Isolated Sign Language Recognition (ISLR) approaches primarily rely on RGB data or signer pose information. However, combining these modalities often results in the loss of crucial details, such as hand shape and orientation, due to imprecise representations like bounding boxes. Therefore, we propose the ISLR system SegSLR, which combines RGB and pose information through promptable zero-shot video segmentation. Given the rough localization of the hands and the signer's body from pose information, we segment the respective parts through the video to maintain all relevant shape information. Subsequently, the segmentations focus the processing of the RGB data on the most relevant body parts for ISLR. This effectively combines RGB and pose information. Our evaluation on the complex ChaLearn249 IsoGD dataset shows that SegSLR outperforms state-of-the-art methods. Furthermore, ablation studies indicate that SegSLR strongly benefits from focusing on the signer's body and hands, justifying our design choices.

手语识别视频分割多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。