梳理提格里尼亚语NLP研究现状,为低资源语言发展提供路线图。
Natural Language Processing for Tigrinya: Current State and Future Directions
- 系统分析2011-2025年超50项研究,覆盖15类下游任务
- 发现从规则系统向神经模型演进,资源建设是关键驱动力
- 提出形态感知建模与社区共建资源等未来方向
尽管有数百万人使用,提格里尼亚语在自然语言处理研究中仍严重缺失。本文对2011至2025年间超过50项提格里尼亚语NLP研究进行了全面综述,系统评估了计算资源、模型与应用在十五类下游任务中的进展,包括形态分析、词性标注、命名实体识别、机器翻译、问答、语音识别与合成等。分析显示,研究轨迹从基础的规则系统逐步转向现代神经架构,进步始终由资源创建里程碑推动。我们识别出由提格里尼亚语形态特性及资源匮乏引发的关键挑战,并提出形态感知建模、跨语言迁移与社区主导资源开发等有前景的研究方向。本工作既可作为研究参考,亦可作为推进提格里尼亚语NLP的路线图。已公开整理的调研文献与资源合集。
原文摘要 · Abstract (English)
Despite being spoken by millions of people, Tigrinya remains severely underrepresented in Natural Language Processing (NLP) research. This work presents a comprehensive survey of NLP research for Tigrinya, analyzing over 50 studies from 2011 to 2025. We systematically review the current state of computational resources, models, and applications across fifteen downstream tasks, including morphological processing, part-of-speech tagging, named entity recognition, machine translation, question-answering, speech recognition, and synthesis. Our analysis reveals a clear trajectory from foundational, rule-based systems to modern neural architectures, with progress consistently driven by milestones in resource creation. We identify key challenges rooted in Tigrinya's morphological properties and resource scarcity, and highlight promising research directions, including morphology-aware modeling, cross-lingual transfer, and community-centered resource development. This work serves both as a reference for researchers and as a roadmap for advancing Tigrinya NLP. An anthology of surveyed studies and resources is publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。