用大模型修复古籍文字识别错误,提升数字遗产可读性。
ICDAR 2026 HIPE-OCRepair Competition on LLM-Assisted OCR Post-Correction for Historical Documents
- 基于大模型对历史文献的文本纠错,按段落级处理无图像输入。
- 多语言效果差异明显,低噪声数据易过度修正,需避免误改。
- 提供公开数据集与评估框架,支持后续研究复现与对比。
我们报告了 ICDAR 2026 HIPE-OCRepair-2026 竞赛的结果,该竞赛聚焦于大语言模型(LLM)辅助的历史文献光学字符识别(OCR)后处理。当前数字遗产面临长期挑战:大量已数字化文档受旧版 OCR 错误影响,而大规模重新扫描不切实际。尽管大模型为这一问题带来新机遇,其在多语言、多文档类型及不同噪声条件下的有效性,以及幻觉风险仍缺乏充分理解。本竞赛旨在:(i) 评估现代 OCR 后处理系统的能力;(ii) 提供基于 HIPE-OCRepair-2026 数据集的可复现评估框架,该数据集整合了现有与新收集的历史资料,涵盖 17 至 20 世纪的英、法、德文报纸与印刷品。参赛者需对噪声文本进行纠错,以段落或文章为单位,无法访问原始图像。评估采用检索导向而非校勘式评分,贴近实际数字馆藏的搜索与访问需求。四支队伍提交系统,涵盖零样本提示到持续预训练与微调策略,揭示不同适配方法的优劣。结果表明,现代 LLM 辅助系统可显著提升 OCR 质量,但性能在数据集、语言和噪声水平间存在差异。低噪声输入下过度修正成为共性问题,凸显仅依赖字符错误率评估的局限性。数据集、评分器与评估流程均已公开发布,以支持未来研究。
原文摘要 · Abstract (English)
We present the results of HIPE-OCRepair-2026, an ICDAR competition on LLM-assisted OCR post-correction of historical documents. OCR post-correction remains a long-standing challenge in digital heritage: large-scale collections of digitized documents are affected by legacy OCR errors, while re-digitization at scale remains impractical. Large language models (LLMs) offers a major opportunity to revisit this challenge, yet their effectiveness across languages, document types, and noise conditions - and their tendency to hallucinate - remains insufficiently understood. HIPE-OCRepair-2026 pursues two objectives: (i) to evaluate the capabilities of modern OCR post-correction systems, and (ii) to provide a reproducible evaluation framework anchored in the HIPE-OCRepair-2026 dataset, a harmonized multilingual resource consolidating existing and newly curated historical datasets. Participants were tasked with correcting noisy OCR transcripts from historical newspapers and printed works in English, French, and German (17th-20th century), working at the level of coherent transcription units (paragraphs or articles) without access to source images. The evaluation adopts a retrieval-oriented rather than diplomatic scoring approach, reflecting the practical use case of search and access over digitized collections. Four teams submitted systems ranging from zero-shot prompting to continued pre-training and fine-tuning, offering insights into the merits of different adaptation strategies. Results show that modern LLM-assisted systems can significantly improve OCR quality, but performance varies across datasets, languages, and noise levels. Over-correction on low-noise inputs emerges as a recurring challenge, highlighting the importance of evaluation beyond character error reduction. The dataset, scorer, and evaluation pipeline are publicly released to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。