用大模型生成历史文本标注数据,提升法语和中文古籍的自然语言处理效果。
Ground Truth Generation for Multilingual Historical NLP using LLMs
- 用大模型生成16-20世纪法语和1900-1950年中文文本的标注数据。
- 基于合成数据微调spaCy,使词性标注、分词和命名实体识别性能显著提升。
- 证明少量合成数据可有效改善低资源历史文本的分析工具,适合数字人文研究者。
历史与低资源自然语言处理因标注数据稀缺及与现代网络语料存在领域差异而面临挑战。本文提出利用大语言模型(LLMs)为16至20世纪法语文本和1900至1950年中文文本生成真实标注数据。通过在部分语料上使用LLM生成的标注数据进行微调,我们成功提升了spaCy在特定时期语料上的词性标注(POS)、词形还原和命名实体识别(NER)性能。结果表明,领域专用模型至关重要,且即使少量合成数据也能显著改进计算人文学科中对低资源语料的处理工具。
原文摘要 · Abstract (English)
Historical and low-resource NLP remains challenging due to limited annotated data and domain mismatches with modern, web-sourced corpora. This paper outlines our work in using large language models (LLMs) to create ground-truth annotations for historical French (16th-20th centuries) and Chinese (1900-1950) texts. By leveraging LLM-generated ground truth on a subset of our corpus, we were able to fine-tune spaCy to achieve significant gains on period-specific tests for part-of-speech (POS) annotations, lemmatization, and named entity recognition (NER). Our results underscore the importance of domain-specific models and demonstrate that even relatively limited amounts of synthetic data can improve NLP tools for under-resourced corpora in computational humanities research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。