用大模型提示工程提升古拉丁语命名实体识别效果
Transfer Learning for Named Entity Recognition of Classical Latin through LLM Prompting
- 通过提示工程调用商用大模型处理古拉丁语文本
- 在粗粒度和细粒度任务中均获第一名,全指标最优
- 为冷门语言提供大模型迁移学习的有效路径
随着古拉丁语文本数字化资源增加及大语言模型(LLMs)的突破,本文参与EvaLatin 2026的命名实体识别(NER)共享任务。任务分为两个子任务:11类粗粒度NER与28类细粒度NER,分别在严格和模糊评价模式下进行。通过针对gemini-2.5-pro和claude-sonnet-4-5的提示工程,证明了古拉丁语这一低资源语言可借助跨语言迁移学习,利用大模型社区的进展实现性能飞跃。报告的方法在两项子任务中均排名第一,所有评估指标与模式下的得分均优于其他提交方案。
原文摘要 · Abstract (English)
With the increase in digitized resources of Classical Latin texts and modern breakthroughs of Large Language Models (LLMs), I contribute to ancient language research by participating in EvaLatin 2026. This paper describes Team uOttawa's system description and results for the Named Entity Recognition (NER) shared task. The task is divided into two subtasks: coarse-grained NER with 11 classes and fine-grained NER with 28 classes, each evaluated under strict and fuzzy regimes. Through prompt engineering of commercial LLMs gemini-2.5-pro and claude-sonnet-4-5, I show that the underrepresented ancient Latin language can take advantage of cross-lingual transfer learning by using advancements made by the wider LLM development community. Overall, the methods discussed in this report demonstrate very strong results, placing first in both NER subtasks and achieving the best scores across all evaluation metrics and regimes among all submissions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。