构建首个斯瓦希里语问答数据集,推动非洲低资源语言AI发展
SwaQuAD-24: QA Benchmark Dataset in Swahili
- 基于SQuAD等基准设计高质量斯瓦希里语问答对
- 涵盖语言多样性,支持机器翻译与医疗聊天机器人应用
- 注重伦理隐私与包容性,适合非洲NLP研究者使用
本文提出构建一个斯瓦希里语问答(QA)基准数据集,以解决斯瓦希里语在自然语言处理(NLP)中代表性不足的问题。借鉴SQuAD、GLUE、KenSwQuAD和KLUE等成熟基准,该数据集将提供高质量、已标注的问答对,涵盖斯瓦希里语的语言多样性和复杂性。数据集旨在支持机器翻译、信息检索以及医疗聊天机器人等社会服务应用。开发过程中强调数据隐私、偏见缓解与包容性等伦理考量。未来计划扩展至特定领域内容、多模态融合及更广泛的众包协作。该数据集致力于促进东非地区技术革新,为低资源语言的NLP研究与应用提供关键资源。
原文摘要 · Abstract (English)
This paper proposes the creation of a Swahili Question Answering (QA) benchmark dataset, aimed at addressing the underrepresentation of Swahili in natural language processing (NLP). Drawing from established benchmarks like SQuAD, GLUE, KenSwQuAD, and KLUE, the dataset will focus on providing high-quality, annotated question-answer pairs that capture the linguistic diversity and complexity of Swahili. The dataset is designed to support a variety of applications, including machine translation, information retrieval, and social services like healthcare chatbots. Ethical considerations, such as data privacy, bias mitigation, and inclusivity, are central to the dataset development. Additionally, the paper outlines future expansion plans to include domain-specific content, multimodal integration, and broader crowdsourcing efforts. The Swahili QA dataset aims to foster technological innovation in East Africa and provide an essential resource for NLP research and applications in low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。