arXiv:2409.13870cs.CLcs.AI2024-09被引 2

用指令微调大模型,修复古希腊碑文与纸草文献的缺失文字。

Instruct-Tuning Pretrained Causal Language Models for Ancient Greek Papyrology and Epigraphy

  • 用指令微调Llama 3.1 8B模型,直接恢复古希腊文本缺字。
  • 文本修复CER仅14.9%,10字符内准确率达73.5%。
  • 适合古籍数字化、考古文本复原等历史语言研究者。

本文实验了对预训练因果语言模型(Meta的Llama 3.1 8B Instruct)进行指令微调,以辅助修复古希腊铭文和文书纸草中的缺失或模糊字符。采用简单的指令式方法和95%/5%训练/测试划分,纸草文本修复模型在不超过10个字符的序列上实现14.9%的字符错误率(CER)、73.5%的top-1准确率和86.0%的top-20准确率。地理归属模型达到66.4% top-1准确率和79.9% top-3准确率;时间归类平均偏差21.7年,中位数为0年。对于铭文,修复模型的CER为20.5%,top-1准确率为63.7%,top-20为83.0%;地理归属准确率分别为75.0%和83.7%,时间归类平均偏差37.1年,中位数3年。在共享测试集及新修订铭文上,指令微调模型在文本修复方面优于当前最优模型Ithaca,且具备忽略空格重建的能力,符合古代文本连续书写的特征。尽管地理与时间归类性能略逊于Ithaca,但在80%/10%/10%划分下重训后仍胜出。结果表明,使用指令模板微调大型预训练因果语言模型用于古籍校勘具有潜力。

原文摘要 · Abstract (English)

This article presents an experiment in fine-tuning a pretrained causal language model (Meta's Llama 3.1 8B Instruct) to assist with restoring missing or illegible characters in ancient Greek inscriptions and documentary papyri. Utilizing a straightforward instruction-based approach and a 95%/5% train/test split, the papyrus restoration model achieved a character error rate (CER) of 14.9%, a top-1 accuracy of 73.5%, and a top-20 accuracy of 86.0% for sequences up to 10 characters. A model was also fine-tuned for geographic attribution, reaching a top-1 accuracy of 66.4% and a top-3 accuracy of 79.9%. In chronological attribution, it demonstrated an average deviation of 21.7 years from the actual terminus post/ante quem, with a median deviation of 0 years. For inscriptions, the restoration model achieved a CER of 20.5%, a top-1 accuracy of 63.7%, and a top-20 accuracy of 83.0% for sequences up to 10 characters. In geographic attribution, it attained a top-1 accuracy of 75.0% and a top-3 accuracy of 83.7%, while in dating, it had an average deviation of 37.1 years and a median deviation of 3 years from the actual date range. Benchmarked against the state-of-the-art model (Ithaca) on a shared test set and on recently edited inscriptions, the instruction-tuned models excelled in text restoration, while also offering the practical advantage of ignoring spaces during reconstruction, which aligns with the scriptio continua of ancient textual artifacts. However, their performance in geographic and chronological attribution was lower than Ithaca's. To evaluate the approach in a more even setup, the instruction model was retrained with an 80%/10%/10% train-validation-test split, and still outperformed Ithaca in text restoration. The results suggest that fine-tuning larger pretrained causal language models using instruction templates for emendations and conjectures to ancient texts holds promise.

古希腊文文本修复大模型应用数字人文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。