arXiv:2512.14489cs.CV2025-12ICCV

构建首个意大利手语识别数据集,助力多模态手语研究

SignIT: A Comprehensive Dataset and Multimodal Analysis for Italian Sign Language Recognition

  • 收集644段视频,覆盖3.33小时,标注94类手语动作
  • 采用2D关键点与RGB帧融合,模型在复杂场景下仍表现有限
  • 适合手语识别、多模态分析及无障碍技术研究者使用

本文提出SignIT,一个用于研究意大利手语(LIS)识别的新数据集。数据集包含644个视频,总时长3.33小时。我们手工标注了属于5大类(动物、食物、颜色、情绪、家庭)的94种不同手语类别,并提取了用户手部、面部和身体的2D关键点。基于该数据集,我们构建了手语识别基准,采用多种先进模型,分析了时间信息、2D关键点与RGB帧对模型性能的影响。结果表明,现有模型在该具有挑战性的LIS数据集上仍存在明显局限。数据与标注已公开:https://fpv-iplab.github.io/SignIT/

原文摘要 · Abstract (English)

In this work we present SignIT, a new dataset to study the task of Italian Sign Language (LIS) recognition. The dataset is composed of 644 videos covering 3.33 hours. We manually annotated videos considering a taxonomy of 94 distinct sign classes belonging to 5 macro-categories: Animals, Food, Colors, Emotions and Family. We also extracted 2D keypoints related to the hands, face and body of the users. With the dataset, we propose a benchmark for the sign recognition task, adopting several state-of-the-art models showing how temporal information, 2D keypoints and RGB frames can be influence the performance of these models. Results show the limitations of these models on this challenging LIS dataset. We release data and annotations at the following link: https://fpv-iplab.github.io/SignIT/.

手语识别多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。