arXiv:2509.22345cs.CL2025-09被引 1

构建16世纪英国宗教攻击性语言语料库,助力历史文本的计算分析

The InviTE Corpus: Annotating Invectives in Tudor English Texts for Computational Modeling

  • 通过专家标注构建近2000句早期现代英语攻击性语言语料
  • 基于历史预训练模型微调后在攻击性检测上表现更优
  • 适合历史语言学与数字人文研究者使用

本文旨在将自然语言处理技术应用于历史研究,聚焦都铎英格兰宗教改革时期的宗教攻击性言论。我们设计了一套从原始数据到迭代标注的工作流程,最终构建了名为InviTE的语料库,包含近2000个16世纪英格兰的早期现代英语句子,均经专家标注攻击性语言特征。随后,我们评估并比较了微调的BERT模型与零样本提示指令微调的大语言模型(LLMs)的表现,结果表明在历史数据上预训练并微调的模型在攻击性检测任务中更具优势。

原文摘要 · Abstract (English)

In this paper, we aim at the application of Natural Language Processing (NLP) techniques to historical research endeavors, particularly addressing the study of religious invectives in the context of the Protestant Reformation in Tudor England. We outline a workflow spanning from raw data, through pre-processing and data selection, to an iterative annotation process. As a result, we introduce the InviTE corpus -- a corpus of almost 2000 Early Modern English (EModE) sentences, which are enriched with expert annotations regarding invective language throughout 16th-century England. Subsequently, we assess and compare the performance of fine-tuned BERT-based models and zero-shot prompted instruction-tuned large language models (LLMs), which highlights the superiority of models pre-trained on historical data and fine-tuned to invective detection.

历史语言攻击性语言数字人文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。