用四阶段流程自动精准转录中世纪拉丁文法律文书。
An HTR-LLM Workflow for High-Accuracy Transcription and Analysis of Abbreviated Latin Court Hand
- 先用LLM优化训练数据训练专用手写识别模型。
- 多模态LLM纠错后,再扩展缩写词为完整拉丁文。
- 最后通过命名实体校正提升术语准确性,适合历史研究者。
本文提出并验证了一种四阶段高精度转录与分析中世纪拉丁文法律文献的流程。第一阶段使用基于新型“纯净真实标注”数据清洗方法训练的专用手写文本识别(HTR)模型生成基础转录结果;第二阶段将基础转录与原始图像输入多模态大语言模型(LLM),实现上下文驱动的后处理纠错;第三阶段利用提示引导的LLM将缩略文本展开为完整的学术拉丁文;第四阶段再次使用LLM进行命名实体校正(NEC),统一专有名词并生成模糊读音的合理替代项。通过案例研究验证,该流程在与学术标准对照下达到2%-7%的词错误率(WER)。结果表明,这种混合多阶段方法能有效自动化最耗时的转录环节,生成高质量、可分析的输出,是当前技术背景下极具实用价值的解决方案。
原文摘要 · Abstract (English)
This article presents and validates an ideal, four-stage workflow for the high-accuracy transcription and analysis of challenging medieval legal documents. The process begins with a specialized Handwritten Text Recognition (HTR) model, itself created using a novel "Clean Ground Truth" curation method where a Large Language Model (LLM) refines the training data. This HTR model provides a robust baseline transcription (Stage 1). In Stage 2, this baseline is fed, along with the original document image, to an LLM for multimodal post-correction, grounding the LLM's analysis and improving accuracy. The corrected, abbreviated text is then expanded into full, scholarly Latin using a prompt-guided LLM (Stage 3). A final LLM pass performs Named-Entity Correction (NEC), regularizing proper nouns and generating plausible alternatives for ambiguous readings (Stage 4). We validate this workflow through detailed case studies, achieving Word Error Rates (WER) in the range of 2-7% against scholarly ground truths. The results demonstrate that this hybrid, multi-stage approach effectively automates the most laborious aspects of transcription while producing a high-quality, analyzable output, representing a powerful and practical solution for the current technological landscape.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。