研究中世纪罗曼语低资源场景下的词性标注关键因素
Unveiling Factors for Enhanced POS Tagging: A Study of Low-Resource Medieval Romance Languages
- 系统测试微调、提示工程等方法对词性标注的影响
- 发现大模型在处理历史拼写变异时表现有限
- 提出适合低资源历史语言的有效标注技术,适合数字人文研究者
词性标注(POS)是自然语言处理的基础环节,尤其在计算语言学与数字人文学科交叉领域对历史文本分析至关重要。尽管现代大语言模型(LLMs)在古代语言处理上取得进展,但其应用于中世纪罗曼语仍面临历时语言演变、拼写变异和标注数据稀缺等独特挑战。本研究系统考察了中世纪奥克语、西班牙语和法语文本在圣经、圣徒传记、医学及饮食等不同领域的词性标注性能影响因素。通过严谨实验,评估了微调策略、提示工程、模型架构、解码方式及跨语言迁移学习对标注准确率的影响。结果揭示大模型在处理历史语言变异和非标准化拼写方面存在明显局限,同时识别出若干能有效应对低资源历史语言挑战的专用技术。
原文摘要 · Abstract (English)
Part-of-speech (POS) tagging remains a foundational component in natural language processing pipelines, particularly critical for historical text analysis at the intersection of computational linguistics and digital humanities. Despite significant advancements in modern large language models (LLMs) for ancient languages, their application to Medieval Romance languages presents distinctive challenges stemming from diachronic linguistic evolution, spelling variations, and labeled data scarcity. This study systematically investigates the central determinants of POS tagging performance across diverse corpora of Medieval Occitan, Medieval Spanish, and Medieval French texts, spanning biblical, hagiographical, medical, and dietary domains. Through rigorous experimentation, we evaluate how fine-tuning approaches, prompt engineering, model architectures, decoding strategies, and cross-lingual transfer learning techniques affect tagging accuracy. Our results reveal both notable limitations in LLMs' ability to process historical language variations and non-standardized spelling, as well as promising specialized techniques that effectively address the unique challenges presented by low-resource historical languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。