首个罗马尼亚手语识别数据集,推动低资源手语研究
RoCoISLR: A Romanian Corpus for Isolated Sign Language Recognition
- 构建包含9000+视频的罗马尼亚孤立手语数据集RoCoISLR
- Transformer模型表现优于传统模型,最佳达34.1%准确率
- 为低资源手语研究提供首个标准化基准,适合手语技术开发者
自动手语识别在弥合聋人与听人沟通鸿沟中至关重要,但现有数据集多聚焦于美国手语。针对罗马尼亚孤立手语识别(RoISLR),尚无大规模标准化数据集,制约了研究进展。本文提出新数据集RoCoISLR,包含超过9,000个视频样本,覆盖近6,000个标准化词汇,来源多样。我们在一致实验设置下评估七种先进视频识别模型(I3D、SlowFast、Swin Transformer、TimeSformer、Uniformer、VideoMAE、PoseConv3D),并与广泛使用的WLASL2000数据集进行对比。结果表明,基于Transformer的架构优于卷积基线模型;其中Swin Transformer达到34.1%的Top-1准确率。该基准揭示了低资源手语中长尾类别分布带来的挑战,为系统性罗氏手语识别研究奠定基础。
原文摘要 · Abstract (English)
Automatic sign language recognition plays a crucial role in bridging the communication gap between deaf communities and hearing individuals; however, most available datasets focus on American Sign Language. For Romanian Isolated Sign Language Recognition (RoISLR), no large-scale, standardized dataset exists, which limits research progress. In this work, we introduce a new corpus for RoISLR, named RoCoISLR, comprising over 9,000 video samples that span nearly 6,000 standardized glosses from multiple sources. We establish benchmark results by evaluating seven state-of-the-art video recognition models-I3D, SlowFast, Swin Transformer, TimeSformer, Uniformer, VideoMAE, and PoseConv3D-under consistent experimental setups, and compare their performance with that of the widely used WLASL2000 corpus. According to the results, transformer-based architectures outperform convolutional baselines; Swin Transformer achieved a Top-1 accuracy of 34.1%. Our benchmarks highlight the challenges associated with long-tail class distributions in low-resource sign languages, and RoCoISLR provides the initial foundation for systematic RoISLR research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。