构建首个以语言学描述区分易混淆手语的日本手语数据集
JSL-DC: A Word-Level Japanese Sign Language Dataset with Linguist-Derived Descriptions for Distinguishing Confusable Signs

- 基于聋人语言学家设计的270个常用词库,采集19名使用者视频
- 在易混淆手语识别上模型准确率提升9.8%,超越现有方法
- 专为听人父母学习手语设计,适合手语教育与识别研究
有效掌握手语对聋童成长至关重要,但95%聋童出生在听力父母家庭,父母常缺乏手语能力。手语识别技术可助力亲子沟通工具开发。然而,日本手语(JSL)缺乏大规模、多说话者数据集,制约模型泛化能力。为此,我们提出JSL-DC,目前规模最大的JSL视频数据集,包含36.7万段视频,来自19名日常使用JSL的聋人。整个流程由聋人主导:270个核心词由聋人及混血聋人语言学家选定,用于促进亲子交流;所有参与者均为聋人;数据经两阶段聋人语言学家审核。此外,我们提供语言学描述以区分易混淆手语。基于这些描述设计的模型,在易混淆子集上比现有最优方法提升9.8%。该数据集及其语言学描述将按CC-BY 4.0许可开源,推动手语识别研究发展。
原文摘要 · Abstract (English)
Effective sign language (SL) acquisition is crucial for deaf children, yet 95% are born to hearing parents who often lack proficiency in SL. SL recognition can power learning tools to help parents communicate with their children. However, Japanese Sign Language (JSL) lacks large-scale, multi-signer datasets, hindering the development of models that can generalize to new users. To address this gap, we introduce JSL-DC, the largest JSL dataset by video count, comprising 36.7K videos from 19 signers. The entire process was Deaf-centric: the lexicon comprising 270 JSL words was selected by Deaf and Coda linguists to facilitate parent-child communication, all participants were Deaf individuals who use JSL daily, and the data underwent a two-stage review process involving Deaf linguists. Moreover, we provide linguist-derived descriptions for distinguishing confusable signs. We demonstrate that the proposed model inspired by the descriptions outperforms state-of-the-art recognition methods by 9.8% on the confusable subset. The dataset, along with its linguistic description that inspires new models, will be released under a CC-BY 4.0 license to accelerate research in SL recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。