用大模型重构亚美尼亚历史报纸阅读顺序,误差降七成六。
Semantic-Guided Reading Order Reconstruction in Historical Armenian Newspapers with LLMs

- 结合语义分区与生成式大模型的混合方法
- 误差比最强几何基线降低76%,多页场景更稳定
- 适合资源匮乏的历史文献快速标注
本文针对亚美尼亚历史报纸中复杂版式与语言资源稀缺的问题,提出阅读顺序重建方法。构建了包含66页的新标注数据集,对比几何启发式、基于YOLO的版面解析、端到端文档模型ECLAIR,以及融合语义区域检测与生成式大模型的混合方法。实验表明,该混合方法误差最低,在多页场景和噪声OCR下仍具鲁棒性,相比最强几何基线减少76%排序错误。方法旨在作为数据增强策略,支持高资源匮乏场景下的快速标注。同时发布专用于历史亚美尼亚印刷体的Tesseract OCR模型。
原文摘要 · Abstract (English)
This paper addresses reading order reconstruction in historical Armenian newspapers, which combine complex layouts with limited language resources. We introduce a new annotated dataset of 66 pages and compare geometric heuristics, YOLO-based layout parsing, an end-to-end document model ECLAIR, and a hybrid method combining semantic zone detection with a generative LLM. Our hybrid method achieves the lowest error rates of all evaluated approaches, reducing ordering errors by up to 76% over the strongest geometric baseline, and remains robust in multi-page settings and under noisy OCR. Rather than targeting production the method is designed as a data bootstrapping strategy enabling rapid annotation in highly under-resourced scenarios. Alongside the dataset, we release a specialized Tesseract OCR model for historical Armenian print.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。