arXiv:2412.02210cs.CV2024-12ICCV被引 93

构建首个全面的OCR评测基准,评估大模型识字能力。

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

  • 设计四大文本任务赛道,覆盖多场景、多语言与文档解析
  • 含7058张真实应用图像,41%来自实际使用场景
  • 揭示大模型在文字定位、多角度识别与重复幻觉上的缺陷

大型多模态模型(LMMs)在自然语言指令下识别文档图像方面表现优异,但其在处理结构复杂、视觉细节丰富的文本时的识字能力仍不明确。当前缺乏全面的评测基准来有效衡量LMMs的识字能力,现有基准多局限于特定场景或任务。为此,我们提出CC-OCR,一个涵盖多样化场景、任务与挑战的综合性基准。该基准包含四个以OCR为核心的赛道:多场景文本阅读、多语言文本阅读、文档解析与关键信息提取,共41个子集,7,058张全标注图像,其中41%源自真实应用场景,首次公开发布。我们评估了九个主流LMMs,揭示其在文本定位、多方向识别及重复内容幻觉方面的优劣。CC-OCR旨在全面评估LMMs在以OCR为中心的任务中的表现,推动该领域持续进步。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) have demonstrated impressive performance in recognizing document images with natural language instructions. However, it remains unclear to what extent capabilities in literacy with rich structure and fine-grained visual challenges. The current landscape lacks a comprehensive benchmark to effectively measure the literate capabilities of LMMs. Existing benchmarks are often limited by narrow scenarios and specified tasks. To this end, we introduce CC-OCR, a comprehensive benchmark that possesses a diverse range of scenarios, tasks, and challenges. CC-OCR comprises four OCR-centric tracks: multi-scene text reading, multilingual text reading, document parsing, and key information extraction. It includes 39 subsets with 7,058 full annotated images, of which 41% are sourced from real applications, and released for the first time. We evaluate nine prominent LMMs and reveal both the strengths and weaknesses of these models, particularly in text grounding, multi-orientation, and hallucination of repetition. CC-OCR aims to comprehensively evaluate the capabilities of LMMs on OCR-centered tasks, facilitating continued progress in this crucial area.

OCR评测多模态模型文档理解真实数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。