arXiv:2505.14311cs.CL2025-05被引 4

梳理豪萨语NLP现状,推动低资源语言研究发展

HausaNLP: Current Status, Challenges and Future Directions for Hausa Natural Language Processing

  • 构建豪萨语资源目录,整合数据集与工具
  • 指出豪萨语在大模型中存在分词与方言问题
  • 呼吁扩充数据、改进建模并加强社区协作

豪萨语自然语言处理近年关注度上升,但作为拥有超1.2亿母语者和8000万第二语言使用者的低资源语言,仍研究不足。尽管高资源语言取得显著进展,豪萨语NLP仍面临开源数据集匮乏、模型表征能力弱等挑战。本文系统梳理豪萨语在文本分类、机器翻译、命名实体识别、语音识别及问答等基础任务上的现有资源与研究进展,发布HausaNLP(https://catalog.hausanlp.org)目录,汇聚数据集、工具与研究成果以提升可及性。同时讨论豪萨语融入大语言模型时的分词不优与方言差异问题,提出未来应聚焦数据扩展、语言建模优化及社区协同,为豪萨语NLP发展提供基础,并为多语言NLP研究提供参考。

原文摘要 · Abstract (English)

Hausa Natural Language Processing (NLP) has gained increasing attention in recent years, yet remains understudied as a low-resource language despite having over 120 million first-language (L1) and 80 million second-language (L2) speakers worldwide. While significant advances have been made in high-resource languages, Hausa NLP faces persistent challenges, including limited open-source datasets and inadequate model representation. This paper presents an overview of the current state of Hausa NLP, systematically examining existing resources, research contributions, and gaps across fundamental NLP tasks: text classification, machine translation, named entity recognition, speech recognition, and question answering. We introduce HausaNLP (https://catalog.hausanlp.org), a curated catalog that aggregates datasets, tools, and research works to enhance accessibility and drive further development. Furthermore, we discuss challenges in integrating Hausa into large language models (LLMs), addressing issues of suboptimal tokenization and dialectal variation. Finally, we propose strategic research directions emphasizing dataset expansion, improved language modeling approaches, and strengthened community collaboration to advance Hausa NLP. Our work provides both a foundation for accelerating Hausa NLP progress and valuable insights for broader multilingual NLP research.

豪萨语低资源语言NLP综述多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。