arXiv:2602.15753cs.CL2026-02被引 3

用大模型零样本完成四门古语言的词形还原与词性标注

Under-resourced studies of under-resourced languages: lemmatization and POS-tagging with LLM annotators for historical Armenian, Georgian, Greek and Syriac

  • 利用GPT-4和Mistral等大模型,在少样本/零样本下处理古希腊语等四门低资源语言
  • 无需微调即在多数语言上达到媲美甚至超越传统模型的标注效果
  • 为无数据支持的古语文本标注提供可行方案,适合历史语言学研究者

低资源语言在自然语言处理中长期面临词形还原与词性标注的挑战。本文研究了近期大语言模型(包括GPT-4变体和开源的Mistral模型)在少量样本和零样本设置下,对四种在历史与语言上差异显著的低资源语言——古希腊语、古典亚美尼亚语、古格鲁吉亚语和叙利亚语——进行词形还原与词性标注的能力。通过一个包含对齐训练集与跨域测试集的新基准,我们评估了基础模型在两类任务上的表现,并与专门设计的RNN基线模型PIE进行对比。结果表明,即使未经微调,大语言模型在大多数语言的少样本设置下,词性标注与词形还原性能均达到竞争性或更优水平。尽管复杂形态与非拉丁字母仍带来显著挑战,但研究证明大语言模型在缺乏数据时可作为有效的标注辅助工具,为语言注释任务提供可信且实用的起点。

原文摘要 · Abstract (English)

Low-resource languages pose persistent challenges for Natural Language Processing tasks such as lemmatization and part-of-speech (POS) tagging. This paper investigates the capacity of recent large language models (LLMs), including GPT-4 variants and open-weight Mistral models, to address these tasks in few-shot and zero-shot settings for four historically and linguistically diverse under-resourced languages: Ancient Greek, Classical Armenian, Old Georgian, and Syriac. Using a novel benchmark comprising aligned training and out-of-domain test corpora, we evaluate the performance of foundation models across lemmatization and POS-tagging, and compare them with PIE, a task-specific RNN baseline. Our results demonstrate that LLMs, even without fine-tuning, achieve competitive or superior performance in POS-tagging and lemmatization across most languages in few-shot settings. Significant challenges persist for languages characterized by complex morphology and non-Latin scripts, but we demonstrate that LLMs are a credible and relevant option for initiating linguistic annotation tasks in the absence of data, serving as an effective aid for annotation.

大模型应用古语言处理零样本学习词性标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。