将古籍数字化从抄写升级为结构化分析,提升准确性与效率。
Quid est VERITAS? A Modular Framework for Archival Document Analysis
- 分四阶段模块化处理,支持按需定义提取目标
- 相比商用OCR,错词率降低67.6%,处理提速三倍
- 适合历史文献研究者,支持语义检索与智能问答
历史文献的数字化长期局限于字符级转录,生成的文本缺乏结构与语义信息,难以支持深度计算分析。我们提出VERITAS(Vision-Enhanced Reading, Interpretation, and Transcription of Archival Sources),一个模块化、模型无关的框架,将数字化重构为包含转录、版面分析与语义增强的集成流程。该流程分为预处理、提取、精炼与增强四个阶段,采用基于模式的架构,使研究者可声明式地指定提取目标。我们在超过1600页的文艺复兴时期《米兰史》权威版本上评估VERITAS,结果表明,相比商用OCR基线,其词错误率相对降低67.6%,在考虑人工校正后,端到端处理时间减少三倍。我们进一步通过检索增强生成系统查询转录语料,验证了输出在支持历史研究方面的下游价值。
原文摘要 · Abstract (English)
The digitisation of historical documents has traditionally been conceived as a process limited to character-level transcription, producing flat text that lacks the structural and semantic information necessary for substantive computational analysis. We present VERITAS (Vision-Enhanced Reading, Interpretation, and Transcription of Archival Sources), a modular, model-agnostic framework that reconceptualises digitisation as an integrated workflow encompassing transcription, layout analysis, and semantic enrichment. The pipeline is organised into four stages - Preprocessing, Extraction, Refinement, and Enrichment - and employs a schema-driven architecture that allows researchers to declaratively specify their extraction objectives. We evaluate VERITAS on the critical edition of Bernardino Corio's Storia di Milano, a Renaissance chronicle of over 1,600 pages. Results demonstrate that the pipeline achieves a 67.6% relative reduction in word error rate compared to a commercial OCR baseline, with a threefold reduction in end-to-end processing time when accounting for manual correction. We further illustrate the downstream utility of the pipeline's output by querying the transcribed corpus through a retrieval-augmented generation system, demonstrating its capacity to support historical inquiry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。