首个大规模真实场景连续手语识别数据集,支持复杂环境下的手语理解。
Isharah: A Large-Scale Multi-Scene Dataset for Continuous Sign Language Recognition
- 用手机拍摄的多场景视频构建,覆盖真实环境中的拍摄差异。
- 包含3万段视频、18位聋人及专业手语者表演,每段有逐词标注。
- 支持跨说话人、未见句子的手语识别与翻译,适合实际应用研究。
当前手语识别(SLR)基准主要聚焦孤立手语,而连续手语识别(CSLR)数据集稀缺且多在受控环境中采集,难以支撑真实场景下的鲁棒系统。为此,我们提出Isharah,首个大规模、多场景、非受控环境下的连续手语识别数据集。该数据集通过手语者自拍手机视频收集,涵盖高度变化的拍摄距离、角度和分辨率,更贴近真实使用场景。数据集包含30,000段视频,由18位聋人及专业手语者完成,每段视频均提供逐词(gloss-level)标注,支持构建手语识别(CSLR)与手语翻译(SLT)系统。本文还引入多个基准任务,包括跨说话人、未见句子的CSLR,以及基于和无基于词义的SLT。Isharah数据集已公开:https://snalyami.github.io/Isharah_CSLR/。
原文摘要 · Abstract (English)
Current benchmarks for sign language recognition (SLR) focus mainly on isolated SLR, while there are limited datasets for continuous SLR (CSLR), which recognizes sequences of signs in a video. Additionally, existing CSLR datasets are collected in controlled settings, which restricts their effectiveness in building robust real-world CSLR systems. To address these limitations, we present Isharah, a large multi-scene dataset for CSLR. It is the first dataset of its type and size that has been collected in an unconstrained environment using signers' smartphone cameras. This setup resulted in high variations of recording settings, camera distances, angles, and resolutions. This variation helps with developing sign language understanding models capable of handling the variability and complexity of real-world scenarios. The dataset consists of 30,000 video clips performed by 18 deaf and professional signers. Additionally, the dataset is linguistically rich as it provides a gloss-level annotation for all dataset's videos, making it useful for developing CSLR and sign language translation (SLT) systems. This paper also introduces multiple sign language understanding benchmarks, including signer-independent and unseen-sentence CSLR, along with gloss-based and gloss-free SLT. The Isharah dataset is available on https://snalyami.github.io/Isharah_CSLR/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。