现有OCR评估忽视历史文献,导致黑人报纸等边缘档案被系统性遗漏。
A Survey of OCR Evaluation Methods and Metrics and the Invisibility of Historical Documents
- 基于PRISMA框架审查2006–2025年论文与数据集,分析评估方法
- 黑人报纸等历史文献极少出现在训练和评测数据中,结构错误常被忽略
- 评估标准偏重现代文档的字符准确率,忽视布局崩塌等真实问题
光学字符识别(OCR)与文档理解系统日益依赖大规模视觉与视觉-语言模型,但评估仍集中于现代、西方及机构化文档。这一倾向掩盖了系统在历史与边缘档案中的表现,其版式、字体与材料退化严重影响识别效果。本研究采用PRISMA框架,回顾2006至2025年间关于OCR与文档理解的论文及基准数据集,分析训练数据、基准设计与评估指标,尤其关注黑人历史报纸。结果显示,黑人报纸及其他社区生成的历史文献极少出现在训练或评测数据中;多数评估聚焦现代排版下的字符准确率与任务成功率,未能捕捉历史报纸常见的列塌陷、排版错误与文本幻觉等结构性失败。结合已有实证研究与重要黑人报刊档案统计,我们指出评估缺口导致结构性隐形与代表性伤害,根源在于组织与制度层面的行为与结构,受基准激励与数据治理决策影响。
原文摘要 · Abstract (English)
Optical character recognition (OCR) and document understanding systems increasingly rely on large vision and vision-language models, yet evaluation remains centered on modern, Western, and institutional documents. This emphasis masks system behavior in historical and marginalized archives, where layout, typography, and material degradation shape interpretation. This study examines how OCR and document understanding systems are evaluated, with particular attention to Black historical newspapers. We review OCR and document understanding papers, as well as benchmark datasets, which are published between 2006 and 2025 using the PRISMA framework. We look into how the studies report training data, benchmark design, and evaluation metrics for vision transformer and multimodal OCR systems. During the review, we found that Black newspapers and other community-produced historical documents rarely appear in reported training data or evaluation benchmarks. Most evaluations emphasize character accuracy and task success on modern layouts. They rarely capture structural failures common in historical newspapers, including column collapse, typographic errors, and hallucinated text. To put these findings into perspective, we use previous empirical studies and archival statistics from significant Black press collections to show how evaluation gaps lead to structural invisibility and representational harm. We propose that these gaps occur due to organizational (meso) and institutional (macro) behaviors and structure, shaped by benchmark incentives and data governance decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。