arXiv:2607.00596cs.CV2026-07被引 1

用大模型重构亚美尼亚历史报纸阅读顺序,误差降七成六。

Semantic-Guided Reading Order Reconstruction in Historical Armenian Newspapers with LLMs

论文配图:Semantic-Guided Reading Order Reconstruction in Historical Armenian Newspapers with LLMs
图 1 · 摘自论文原文
  • 结合语义分区与生成式大模型的混合方法
  • 误差比最强几何基线降低76%,多页场景更稳定
  • 适合资源匮乏的历史文献快速标注

本文针对亚美尼亚历史报纸中复杂版式与语言资源稀缺的问题,提出阅读顺序重建方法。构建了包含66页的新标注数据集,对比几何启发式、基于YOLO的版面解析、端到端文档模型ECLAIR,以及融合语义区域检测与生成式大模型的混合方法。实验表明,该混合方法误差最低,在多页场景和噪声OCR下仍具鲁棒性,相比最强几何基线减少76%排序错误。方法旨在作为数据增强策略,支持高资源匮乏场景下的快速标注。同时发布专用于历史亚美尼亚印刷体的Tesseract OCR模型。

原文摘要 · Abstract (English)

This paper addresses reading order reconstruction in historical Armenian newspapers, which combine complex layouts with limited language resources. We introduce a new annotated dataset of 66 pages and compare geometric heuristics, YOLO-based layout parsing, an end-to-end document model ECLAIR, and a hybrid method combining semantic zone detection with a generative LLM. Our hybrid method achieves the lowest error rates of all evaluated approaches, reducing ordering errors by up to 76% over the strongest geometric baseline, and remains robust in multi-page settings and under noisy OCR. Rather than targeting production the method is designed as a data bootstrapping strategy enabling rapid annotation in highly under-resourced scenarios. Alongside the dataset, we release a specialized Tesseract OCR model for historical Armenian print.

历史文献阅读顺序大模型低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。