测试开源大模型在古奥克语词性标注中的表现,发现其处理历史语言有明显局限。
Modern Models, Medieval Texts: A POS Tagging Study of Old Occitan
- 对比两个文本语料,评估大模型对古奥克语词性标注能力
- 模型在拼写和句法高度变异时表现显著下降
- 研究结果对历史语言计算处理有实用指导意义
大型语言模型(LLMs)在自然语言处理中表现出色,但在处理历史语言方面的效果仍不明确。本研究考察了开源大模型在古奥克语(一种拼写不统一、历时变化显著的历史语言)的词性标注(POS tagging)任务中的表现。通过对比两组不同语料——圣徒传记类与医学文本——评估当前模型在低资源历史语言处理中的应对能力。研究发现,当面对极端拼写与句法变异时,大模型性能存在严重局限。我们进行了详细的错误分析,并提出改进模型在历史语言处理中表现的具体建议。该研究深化了对大模型在复杂语言环境下的能力理解,同时为计算语言学与历史语言研究提供了实际参考。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable capabilities in natural language processing, yet their effectiveness in handling historical languages remains largely unexplored. This study examines the performance of open-source LLMs in part-of-speech (POS) tagging for Old Occitan, a historical language characterized by non-standardized orthography and significant diachronic variation. Through comparative analysis of two distinct corpora-hagiographical and medical texts-we evaluate how current models handle the inherent challenges of processing a low-resource historical language. Our findings demonstrate critical limitations in LLM performance when confronted with extreme orthographic and syntactic variability. We provide detailed error analysis and specific recommendations for improving model performance in historical language processing. This research advances our understanding of LLM capabilities in challenging linguistic contexts while offering practical insights for both computational linguistics and historical language studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。