构建车载手语数据集,助力聋哑人出行无障碍
The In-Car Sign Language Corpus (ICSL): A Multi-Modal Resource for Constrained-Space Sign Language Recognition

- 采集真实车内多模态手语数据,含2D摄像头与3D传感器信号
- 覆盖超150万帧画面,支持深度学习模型训练与评估
- 专为狭小封闭空间手语识别设计,适合交通无障碍研究
本文针对共享出行服务中手语使用面临的挑战,提出面向巴西手语(Libras)的车载手语数据集(ICSL)。该数据集包含高精度实验室动作捕捉数据与真实车内多模态记录,由2D相机和3D飞行时间传感器同步采集。整体数据量超过150万帧,涵盖多种车内场景下的手语表现。数据提供词素标注及非词素语言元素标注,用于支持约束空间下手语识别模型的训练与评估。该资源为构建鲁棒的“真实世界”手语识别系统与领域自适应研究奠定基础。车内环境具有空间受限、遮挡多、视角非正面等特点,是手语识别的重要挑战场景。
原文摘要 · Abstract (English)
This paper addresses the challenges of using sign language within shared mobility services, such as taxis, carpools, or ride-sharing platforms. The use of sign language recognition (SLR) in real-world, confined environments, specifically vehicle interiors remains largely unexplored. To motivate research in this area, we present the In-Car Sign Language (ICSL) dataset for Brazilian Sign Language (Libras), with the long-term goal of improving public transport accessibility for the Deaf and Hard-of-Hearing community. The dataset consists of: (1) high-precision laboratory motion capture (MoCap) data to establish an idealized linguistic baseline and (2) real-world multi-modal in-car recordings captured using a 2D camera and 3D Time-of-Flight sensors. The dataset provides a basis for comparative analyses between synthesized signing avatar animations and recorded real signing interpreter videos, which enable future research into robust "in-the-wild" SLR models and domain adaptation. We describe in detail the use cases, the setup, the data collection protocol, and the metadata structure of the corpus. In total, we recorded a multimodal dataset exceeding 1.5 million frames, comprising the synchronized multimodal streams described above featuring Libras users across various in-car scenarios. The corpus is provided with gloss annotation of lexical signs and non-lexical sign language elements specially designed to support the training and evaluation of deep neural networks for constrained space recognition. In-vehicle signing offers a technically significant example of a constrained, occluded, and non-frontal environment. While recognizing the diverse communication strategies already employed by the Deaf community, identifying automotive-specific limitations provides a useful stepping stone for research into enhancing in-car accessibility and passenger quality of life.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。