用大模型提升尼泊尔语命名实体识别,解决低资源语言难题
Generative AI for Named Entity Recognition in Low-Resource Language Nepali
- 采用提示工程调优大模型,适配尼泊尔语命名实体识别任务
- 实测显示主流大模型在尼泊尔语上仍具潜力但性能有限
- 为低资源语言NLP研究提供可复用的实验框架与数据参考
生成式人工智能(GenAI),特别是大型语言模型(LLMs),显著推动了自然语言处理(NLP)任务的发展,如命名实体识别(NER),即从文本中识别人物、地点、组织等实体。由于具备从少量数据中学习的能力,LLMs在低资源语言中尤为有前景。然而,针对尼泊尔语这一低资源语言的GenAI模型表现尚未得到充分评估。本文研究了前沿大模型在尼泊尔语NER中的应用,通过多种提示技术进行实验,以评估其有效性。结果揭示了在低资源环境下使用大模型进行NER所面临的挑战与机遇,为尼泊尔语等语言的NLP研究提供了重要贡献。
原文摘要 · Abstract (English)
Generative Artificial Intelligence (GenAI), particularly Large Language Models (LLMs), has significantly advanced Natural Language Processing (NLP) tasks, such as Named Entity Recognition (NER), which involves identifying entities like person, location, and organization names in text. LLMs are especially promising for low-resource languages due to their ability to learn from limited data. However, the performance of GenAI models for Nepali, a low-resource language, has not been thoroughly evaluated. This paper investigates the application of state-of-the-art LLMs for Nepali NER, conducting experiments with various prompting techniques to assess their effectiveness. Our results provide insights into the challenges and opportunities of using LLMs for NER in low-resource settings and offer valuable contributions to the advancement of NLP research in languages like Nepali.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。