arXiv:2410.18295cs.AIcs.CL2024-10被引 2

构建首个肯尼亚手语数据集,助力聋人与健听者沟通无障碍

Kenyan Sign Language (KSL) Dataset: Using Artificial Intelligence (AI) in Bridging Communication Barrier among the Deaf Learners

  • 采集48名教师与400名聋生的自发与诱导手语数据
  • 产出约1.4万句英文对应手语词表、2万段手语视频及4000个音位参数标注
  • 为非洲手语数字化提供可复用框架,适合残障科技与语言计算研究者

肯尼亚手语(KSL)是肯尼亚聋人群体的主要交流语言,广泛用于从学前到大学阶段的教学与社会互动。然而,聋人与健听者之间仍存在显著语言障碍。为此,2023-2024年开展的AI4KSL项目致力于构建开放获取的数字人工智能手语数据集。本研究通过采集48名聋教育教师和400名聋生参与的阅读与歌唱任务,收集了约14,000条英文句子对应的KSL Gloss(词表),涵盖约4,000个词汇;同时生成约20,000段手语视频(单字或句子)。第二阶段产出10,000段分割与分段的手语视频。第三阶段将4,000个手语词转录为5个发音部位参数,采用HamNoSys系统标注。该数据集为实现英-肯尼亚手语自动翻译提供了基础,推动聋人教育包容性发展。

原文摘要 · Abstract (English)

Kenyan Sign Language (KSL) is the primary language used by the deaf community in Kenya. It is the medium of instruction from Pre-primary 1 to university among deaf learners, facilitating their education and academic achievement. Kenyan Sign Language is used for social interaction, expression of needs, making requests and general communication among persons who are deaf in Kenya. However, there exists a language barrier between the deaf and the hearing people in Kenya. Thus, the innovation on AI4KSL is key in eliminating the communication barrier. Artificial intelligence for KSL is a two-year research project (2023-2024) that aims to create a digital open-access AI of spontaneous and elicited data from a representative sample of the Kenyan deaf community. The purpose of this study is to develop AI assistive technology dataset that translates English to KSL as a way of fostering inclusion and bridging language barriers among deaf learners in Kenya. Specific objectives are: Build KSL dataset for spoken English and video recorded Kenyan Sign Language and to build transcriptions of the KSL signs to a phonetic-level interface of the sign language. In this paper, the methodology for building the dataset is described. Data was collected from 48 teachers and tutors of the deaf learners and 400 learners who are Deaf. Participants engaged mainly in sign language elicitation tasks through reading and singing. Findings of the dataset consisted of about 14,000 English sentences with corresponding KSL Gloss derived from a pool of about 4000 words and about 20,000 signed KSL videos that are either signed words or sentences. The second level of data outcomes consisted of 10,000 split and segmented KSL videos. The third outcome of the dataset consists of 4,000 transcribed words into five articulatory parameters according to HamNoSys system.

手语识别无障碍技术数据集AI for Good

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。