arXiv:2505.20113cs.CLcs.AI2025-05被引 4

针对19世纪意大利文献,构建了首个中文命名实体识别数据集。

Named Entity Recognition in Historical Italian: The Case of Giacomo Leopardi's Zibaldone

  • 基于莱奥帕尔迪《札记》构建2899条实体标注数据集
  • 微调的BERT模型在文献引用识别上表现更优
  • 大模型指令调优后仍难处理历史人文文本

全球文本遗产的数字化带来了重大挑战,尤其在计算机科学与人文学科交叉领域。历史文本常存在拼写变异、结构残缺和数字化错误等问题,亟需适应性强的计算方法。大型语言模型(LLMs)虽已革新自然语言处理,但尚无针对意大利语历史文本的系统评估。本研究填补空白,基于19世纪学术笔记——贾科莫·莱奥帕尔迪的《札记》(Zibaldone, 1898),构建了一个新数据集,包含2,899个人物、地点与文学作品的实体引用。该数据集用于可复现实验,对比了领域专用BERT模型与前沿大模型如LLaMa3.1的表现。结果表明,经指令微调的大模型在处理历史人文文本时面临多重困难,而微调的命名实体识别模型在如文献引用等复杂实体类型上仍具更强鲁棒性。

原文摘要 · Abstract (English)

The increased digitization of world's textual heritage poses significant challenges for both computer science and literary studies. Overall, there is an urgent need of computational techniques able to adapt to the challenges of historical texts, such as orthographic and spelling variations, fragmentary structure and digitization errors. The rise of large language models (LLMs) has revolutionized natural language processing, suggesting promising applications for Named Entity Recognition (NER) on historical documents. In spite of this, no thorough evaluation has been proposed for Italian texts. This research tries to fill the gap by proposing a new challenging dataset for entity extraction based on a corpus of 19th century scholarly notes, i.e. Giacomo Leopardi's Zibaldone (1898), containing 2,899 references to people, locations and literary works. This dataset was used to carry out reproducible experiments with both domain-specific BERT-based models and state-of-the-art LLMs such as LLaMa3.1. Results show that instruction-tuned models encounter multiple difficulties handling historical humanistic texts, while fine-tuned NER models offer more robust performance even with challenging entity types such as bibliographic references.

命名实体识别历史文本意大利语BERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。