首个面向澳大利亚手语的多视角多模态词级识别数据集,助力手语技术发展。
MM-WLAuslan: Multi-View Multi-Modal Word-Level Australian Sign Language Recognition Dataset
- 构建多视角多模态数据采集系统,覆盖73位表演者
- 包含3215个常见词、超28万段视频,规模领先
- 适合手语识别、跨视角建模等研究,支持无障碍交流
孤立手语识别(ISLR)关注单个手语词素的识别。由于地域差异,开发区域专属的ISLR数据集对促进交流与研究至关重要。作为澳大利亚特有手语的澳式手语(Auslan)目前尚缺乏大规模词级数据集。为此,我们构建了首个大规模多视角多模态词级澳式手语识别数据集,命名为MM-WLAuslan。相比现有公开数据集,该数据集具有三大优势:(1)数据量最大,(2)词汇最丰富,(3)多模态摄像机视角最多样。具体而言,我们在演播室环境中记录了超过28.2万段手语视频,涵盖3,215个常用澳式手语词素,由73位表演者完成。拍摄系统包含三种Kinect-V2相机与一台RealSense相机,呈半球形分布于表演者前方,同步录制四路视频。我们还基于状态最优方法,在多视角、跨相机、跨视角等多种设置下进行了基准测试,结果表明该数据集极具挑战性。我们希望此数据集能推动澳式手语技术发展,并为全球手语研究提供支持。所有数据与基准测试代码均已开源。
原文摘要 · Abstract (English)
Isolated Sign Language Recognition (ISLR) focuses on identifying individual sign language glosses. Considering the diversity of sign languages across geographical regions, developing region-specific ISLR datasets is crucial for supporting communication and research. Auslan, as a sign language specific to Australia, still lacks a dedicated large-scale word-level dataset for the ISLR task. To fill this gap, we curate \underline{\textbf{the first}} large-scale Multi-view Multi-modal Word-Level Australian Sign Language recognition dataset, dubbed MM-WLAuslan. Compared to other publicly available datasets, MM-WLAuslan exhibits three significant advantages: (1) the largest amount of data, (2) the most extensive vocabulary, and (3) the most diverse of multi-modal camera views. Specifically, we record 282K+ sign videos covering 3,215 commonly used Auslan glosses presented by 73 signers in a studio environment. Moreover, our filming system includes two different types of cameras, i.e., three Kinect-V2 cameras and a RealSense camera. We position cameras hemispherically around the front half of the model and simultaneously record videos using all four cameras. Furthermore, we benchmark results with state-of-the-art methods for various multi-modal ISLR settings on MM-WLAuslan, including multi-view, cross-camera, and cross-view. Experiment results indicate that MM-WLAuslan is a challenging ISLR dataset, and we hope this dataset will contribute to the development of Auslan and the advancement of sign languages worldwide. All datasets and benchmarks are available at MM-WLAuslan.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。