让OCR错误诊断与修复可执行,提升复杂文档识别准确率
OCR-EDR: Rendering-Aware Diagnosis and Repair for Closed-Loop OCR Improvement

- 基于渲染一致性检测,逐像素诊断并定位OCR错误
- 修复86.23%错误输入,公式F1提升超30个百分点
- 适合需要高精度文档解析的科研与工业场景
尽管文档OCR系统在常规文档上表现良好,但复杂公式、结构化文本和长尾格式仍易出错。现有评估方法多为聚合指标,难以支持细粒度错误分析与性能改进。本文提出OCR-EDR(OCR错误诊断与修复)框架,从细粒度诊断到迭代修复,结合源图像、可编辑的OCR预测及其渲染结果,联合判断预测与渲染是否与源一致,保留有效预测(包括渲染等价项),同时定位真实错误。随后执行可执行修正,并可请求新渲染以迭代评估。我们构建了OCRErrBench数据集,涵盖文本、公式、精确匹配与渲染等价正例及真实错误。开发的DocEDR模型在该数据集上诊断准确率达94.78%。它修复86.23%错误输入至视觉一致,相较DOCR-Inspector-7B在DOCRcaseBench上公式Case-F1提升30.99个百分点,在UniMER-Test的四个错误子集上公式CDM最高提升4.62个百分点。结果表明,OCR-EDR将细粒度分析转化为可验证修正与性能提升。
原文摘要 · Abstract (English)
Although document OCR systems perform increasingly well on routine documents, complex formulas, structured text, and long-tail formats remain error-prone. OCR predictions may omit fine-grained content or hallucinate unsupported outputs, while equivalent encodings of the same visible content must be accommodated. Existing OCR evaluation methods mostly report aggregate metrics, offering limited support for analyzing case-level errors and improving OCR performance. We propose OCR-EDR (OCR Error Diagnosis and Repair), a rendering-aware framework that advances from fine-grained diagnosis to iterative repair. Given a source image, an editable OCR prediction, and its rendered image, OCR-EDR first jointly assesses whether the prediction and its rendering are consistent with the source, preserving valid predictions, including rendering-equivalent ones, while diagnosing and localizing genuine errors. It then applies executable edits and may request an updated rendering for iterative reassessment. We construct OCRErrBench from diverse real OCR predictions, covering text and formulas, exact and rendering-equivalent positives, and genuine errors, and develop the DocEDR model to execute the diagnosis--repair loop. On OCRErrBench, DocEDR achieves 94.78% diagnostic accuracy. It repairs 86.23% of erroneous inputs to visual consistency, raises formula Case-F1 by 30.99 percentage points over DOCR-Inspector-7B on DOCRcaseBench, and improves formula CDM by up to 4.62 percentage points on the identified Bad subsets of four OCR systems on UniMER-Test. These results show that OCR-EDR turns fine-grained OCR analysis into verified corrections and performance gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。