用大模型提升中世纪罗曼语词性标注,效果显著优于传统方法。
From Traditional Taggers to LLMs: A Comparative Study of POS Tagging for Medieval Romance Languages

- 对比规则、统计与大模型在零样本、少样本等场景下的表现
- 微调和多语言训练使准确率提升最明显,跨语言迁移对资源少的语言尤其有效
- 适合数字人文研究者参考,推动历史文本的智能处理
中世纪罗曼语的词性标注因拼写变异、形态复杂及标注资源有限而面临挑战。本文系统评估了大语言模型(LLMs)在三种中世纪语言——中世纪奥克语、中世纪加泰罗尼亚语和中世纪法语中的词性标注表现。在零样本提示、少样本提示、单语微调和跨语言迁移学习等多种设置下,对比传统基于规则和统计的标注器。实验基于历史语料库显示,基于大模型的方法始终优于传统方法,其中微调和多语言训练带来最大提升。特别是跨语言迁移对资源匮乏的语言有显著帮助,而针对特定目标语言的双语训练甚至优于更广泛的多语言配置。结果强调了语言亲缘关系与数据特征在历史自然语言处理迁移策略设计中的重要性。研究为现代神经方法应用于中世纪文本处理提供了实证支持,并为数字人文研究中的大模型词性标注流程部署提供实践指导。所有代码、模型与处理后的数据集均已开源,确保可复现性。
原文摘要 · Abstract (English)
Part-of-speech (POS) tagging for Medieval Romance languages remains challenging due to orthographic variation, morphological complexity, and limited annotated resources. This paper presents a systematic empirical evaluation of large language models (LLMs) for POS tagging across three medieval varieties: Medieval Occitan, Medieval Catalan, and Medieval French. We compare traditional rule-based and statistical taggers with modern open-source LLMs under zero-shot prompting, few-shot prompting, monolingual fine-tuning, and cross-lingual transfer learning settings. Experiments on historically grounded datasets show that LLM-based approaches consistently outperform traditional taggers, with fine-tuning and multilingual training yielding the largest improvements. In particular, cross-lingual transfer learning substantially benefits under-resourced varieties, while targeted bilingual training can outperform broader multilingual configurations for specific target languages. The results highlight the importance of linguistic proximity and dataset characteristics when designing transfer strategies for historical NLP. These findings provide empirical insights into the applicability of modern neural methods to medieval text processing and provide practical guidance for deploying LLM-based POS tagging pipelines in digital humanities research. All code, models, and processed datasets are released for reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。