聚焦吉尔吉斯语的NLP困境,呼吁社区共建语言资源。
KyrgyzNLP: Challenges, Progress, and Future
- 以母语者标注数据为基础,评估吉尔吉斯语NLP真实水平。
- 吉尔吉斯语被列为'勉强维持'资源匮乏语言,仅少数公开数据集可用。
- 强调社区驱动建设,提出未来研究与资源发展路线图。
大型语言模型在众多基准测试中表现优异,推动了语言与非语言任务的AI应用进展,但主要惠及资源丰富语言,使低资源语言(LRLs)处于劣势。本文聚焦于低资源语言——吉尔吉斯语(kyrgyz tili)的NLP现状。母语者参与的人工评估及标注数据,是确保低资源语言可靠性能不可或缺的部分,尤其当自动评估存在局限时。在近期对突厥语族语言资源的评估中,吉尔吉斯语被评为'Scraping By'(勉强维持),尽管其使用者达数百万,且在吉尔吉斯斯坦及海外侨民中日益重要,却缺乏官方地位。我们回顾已有研究,发现多数公开资源为近年新创,除词典外极少其他数据。虽有部分论文取得进展,仍需大量工作。尽管商业与政府部门已有兴趣与支持,吉尔吉斯语资源建设仍面临严峻挑战。本文强调社区主导资源建设的重要性,以保障可持续发展,并提出当前最紧迫的挑战及未来研究与资源发展的路线图。
原文摘要 · Abstract (English)
Large language models (LLMs) have excelled in numerous benchmarks, advancing AI applications in both linguistic and non-linguistic tasks. However, this has primarily benefited well-resourced languages, leaving less-resourced ones (LRLs) at a disadvantage. In this paper, we highlight the current state of the NLP field in the specific LRL: kyrgyz tili. Human evaluation, including annotated datasets created by native speakers, remains an irreplaceable component of reliable NLP performance, especially for LRLs where automatic evaluations can fall short. In recent assessments of the resources for Turkic languages, Kyrgyz is labeled with the status 'Scraping By', a severely under-resourced language spoken by millions. This is concerning given the growing importance of the language, not only in Kyrgyzstan but also among diaspora communities where it holds no official status. We review prior efforts in the field, noting that many of the publicly available resources have only recently been developed, with few exceptions beyond dictionaries (the processed data used for the analysis is presented at https://kyrgyznlp.github.io/). While recent papers have made some headway, much more remains to be done. Despite interest and support from both business and government sectors in the Kyrgyz Republic, the situation for Kyrgyz language resources remains challenging. We stress the importance of community-driven efforts to build these resources, ensuring the future advancement sustainability. We then share our view of the most pressing challenges in Kyrgyz NLP. Finally, we propose a roadmap for future development in terms of research topics and language resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。