arXiv:2607.16203cs.LGcs.AI2026-07

无需标注数据,自动评估并选择最优OCR工具

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth

论文配图:DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
图 1 · 摘自论文原文
  • 通过三阶段纠错与排序,无须真实标签评估OCR性能
  • 多MLLM集成结果显著逼近人工标注排名
  • 适合缺乏标注数据的现实文档处理场景

文档解析是视觉问答和关键信息提取等任务的基础,将扫描图像转换为结构化文本、视觉与布局信息。尽管已有多种OCR引擎和多模态大模型(MLLMs)用于此目的,但在标签稀缺情况下为特定文档集选择合适解析方案仍具挑战。本文系统评估了多种OCR引擎与前沿MLLM在跨领域、跨语言扫描文档基准上的文本识别表现。针对多数OCR引擎上下文推理能力有限及人工标注成本高的问题,提出DocOCR-Eval——一种无需标注的自动化OCR评估与选型框架。该框架采用三阶段纠错与排序策略,逼近基于标注的工具排序。实验表明,融合多个MLLM的结果可逐步提升与人工标注排名的一致性。大量实验证明,在真实、标注受限场景下仍可实现可靠工具选择,为多样化实际文档集部署提供实用指导。

原文摘要 · Abstract (English)

Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting textual, visual, and layout information. While numerous Optical Character Recognition (OCR) engines and multimodal large language models (MLLMs) have been developed for this purpose, selecting an appropriate document parsing solution for a given document collection remains challenging, particularly in label-scarce settings. In this work, we conduct a systematic evaluation of text recognition performance across a diverse set of OCR engines and state-of-the-art MLLMs on multiple scanned document benchmarks spanning different domains and languages. Motivated by the limited contextual reasoning capabilities of many OCR engines and the high cost of manual annotations, we propose DocOCR-Eval, an annotation-free evaluation framework for automatic OCR assessment and selection. DocOCR-Eval employs a three-staged correction and ranking strategy to approximate annotation-based tool ordering without ground-truth labels. We show that aggregating across multiple MLLMs progressively improves alignment with annotation-based rankings. Extensive experiments further demonstrate that reliable OCR tool selection can be achieved in realistic, label-limited settings, providing practical guidance for deploying document parsing systems across diverse real-world document collections.

OCR评估无标注文档解析MLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。