将古法语和拉丁语手稿的识别结果按编辑规范预归一化,提升可读性与下游工具兼容性。
Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin
- 提出预编辑归一化任务,保留原始手稿特征同时生成易用文本。
- 构建466万样本训练集和1800样本黄金标准评估集,模型CER达6.7%。
- 适合历史文献数字化、数字人文研究者及古籍NLP工具开发者。
自动文本识别(ATR)的进步提升了对历史档案的访问能力,但古文字转录与标准化数字版之间仍存在方法论鸿沟。基于如CATMuS等更注重古文字学数据集训练的ATR模型具有更强泛化能力,但其原始输出与读者及下游NLP工具兼容性差;而专为生成标准化输出训练的模型则在新领域适应性差,易过度归一化或产生幻觉。本文提出预编辑归一化(PEN)任务,即根据编辑惯例对ATR生成的图形字符输出进行归一化,兼顾古文字保真度与实用性。我们基于CoMMA语料库构建新数据集,并通过passim对齐数字化的古法语和拉丁语版本,同时创建人工校正的黄金标准评估集。采用ByT5序列到序列模型在归一化和预标注任务上进行基准测试。贡献包括正式定义PEN任务、构建466万样本银质训练集、1800样本黄金评估集,以及实现6.7% CER的归一化模型,显著优于现有方法。
原文摘要 · Abstract (English)
Recent advances in Automatic Text Recognition (ATR) have improved access to historical archives, yet a methodological divide persists between palaeographic transcriptions and normalized digital editions. While ATR models trained on more palaeographically-oriented datasets such as CATMuS have shown greater generalizability, their raw outputs remain poorly compatible with most readers and downstream NLP tools, thus creating a usability gap. On the other hand, ATR models trained to produce normalized outputs have been shown to struggle to adapt to new domains and tend to over-normalize and hallucinate. We introduce the task of Pre-Editorial Normalization (PEN), which consists in normalizing graphemic ATR output according to editorial conventions, which has the advantage of keeping an intermediate step with palaeographic fidelity while providing a normalized version for practical usability. We present a new dataset derived from the CoMMA corpus and aligned with digitized Old French and Latin editions using passim. We also produce a manually corrected gold-standard evaluation set. We benchmark this resource using ByT5-based sequence-to-sequence models on normalization and pre-annotation tasks. Our contributions include the formal definition of PEN, a 4.66M-sample silver training corpus, a 1.8k-sample gold evaluation set, and a normalization model achieving a 6.7% CER, substantially outperforming previous models for this task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。