针对黑人报纸档案设计布局感知OCR,提升低资源环境下的文字识别质量。
Layout-Aware OCR for Black Digital Archives with Unsupervised Evaluation
- 融合合成布局生成与YOLO检测器,构建适应黑人报纸布局的OCR流水线。
- 在400页数据上,布局感知方法减少区域冗余,提升结构多样性。
- 提出无监督评估框架,适合标注稀缺的历史档案场景。
尽管具有重要的文化和历史价值,黑人数字档案在人工智能研究和基础设施中仍严重缺位。尤其在数字化黑人报纸时,不一致的排版、视觉退化及有限的标注布局数据阻碍了准确转录,尽管已有多种系统声称能良好处理光学字符识别(OCR)。本文提出一种专为黑人报纸档案设计的布局感知OCR流程,并引入适用于低资源档案场景的无监督评估框架。方法结合合成布局生成、增强数据上的模型预训练,以及最先进的You Only Look Once(YOLO)检测器融合。采用三种无标注评估指标:语义连贯性得分(SCS)、区域熵(RE)和文本冗余度得分(TRS),分别衡量语言流畅性、信息多样性与区域间冗余程度。在包含十种黑人报纸标题共400页的数据集上,布局感知OCR相比全图基线显著提升了结构多样性并降低了冗余,仅略有牺牲连贯性。结果凸显在AI文档理解中尊重文化布局逻辑的重要性,为未来社区驱动、伦理导向的档案智能系统奠定基础。
原文摘要 · Abstract (English)
Despite their cultural and historical significance, Black digital archives continue to be a structurally underrepresented area in AI research and infrastructure. This is especially evident in efforts to digitize historical Black newspapers, where inconsistent typography, visual degradation, and limited annotated layout data hinder accurate transcription, despite the availability of various systems that claim to handle optical character recognition (OCR) well. In this short paper, we present a layout-aware OCR pipeline tailored for Black newspaper archives and introduce an unsupervised evaluation framework suited to low-resource archival contexts. Our approach integrates synthetic layout generation, model pretraining on augmented data, and a fusion of state-of-the-art You Only Look Once (YOLO) detectors. We used three annotation-free evaluation metrics, the Semantic Coherence Score (SCS), Region Entropy (RE), and Textual Redundancy Score (TRS), which quantify linguistic fluency, informational diversity, and redundancy across OCR regions. Our evaluation on a 400-page dataset from ten Black newspaper titles demonstrates that layout-aware OCR improves structural diversity and reduces redundancy compared to full-page baselines, with modest trade-offs in coherence. Our results highlight the importance of respecting cultural layout logic in AI-driven document understanding and lay the foundation for future community-driven and ethically grounded archival AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。