arXiv:2605.03903cs.CL2026-05

剖析大模型在真实文档场景下的失败原因,揭示性能与实际部署间的差距。

CC-OCR V2: Fine-Grained Attribution of LMM Failures in Real-World Visual Document Understanding

论文配图:CC-OCR V2: Fine-Grained Attribution of LMM Failures in Real-World Visual Document Understanding
图 1 · 摘自论文原文
  • 构建涵盖16个子任务、7093样本的细粒度评估框架
  • 发现模型在不同文档类型和采集条件下表现差异显著
  • 提供10类细粒度文档因素标注,助力故障归因分析

近期大型多模态模型(LMMs)在以光学字符识别(OCR)为核心的文档理解与处理任务中取得了显著进展。现有基准主要评估LMMs在多样化任务中的表现,或分析文档特征对性能的影响,但对真实文档采集条件下的可靠性缺乏深入洞察。为此,我们提出CC-OCR v2,一个面向真实文档处理中LMM失败归因的综合性基准。该基准涵盖五个核心文档处理能力,包含16个子任务、74种应用场景和7,093个样本,覆盖五项评估赛道、十类文档类别及32种语言。每个样本均标注了十类细粒度文档因素,支持跨文档类型与采集条件的系统性故障归因。对17个代表性LMM的实验表明,其性能在任务、文档类别及真实条件间存在显著差异;即便整体准确率相近,模型在特定文档因素下的失败模式仍截然不同,凸显基准性能与实际可靠部署之间的巨大鸿沟。数据集与评估工具已开源:https://github.com/eioss/CC-OCR-V2。

原文摘要 · Abstract (English)

Recent Large Multimodal Models (LMMs) have achieved remarkable progress on OCR-centric document understanding and processing tasks. Existing benchmarks primarily evaluate LMMs across diverse tasks to reflect practical document-processing workflows or analyze how document characteristics influence model performance. However, they provide limited insight into the reliability of LMMs under real-world document acquisition conditions, where factors such as lighting, screen displays, imaging quality, and capture methods can substantially affect performance. To bridge this gap, we present CC-OCR v2, a comprehensive benchmark for attributing LMM failures in real-world document processing. CC-OCR v2 provides a unified evaluation framework covering five core document-processing capabilities through 16 subtasks, 74 application scenarios, and 7,093 samples. The benchmark spans five evaluation tracks, ten document categories, and 32 languages in the recognition suite. Beyond task-level evaluation, each sample is annotated with ten fine-grained document factors, enabling systematic attribution of model failures across document types and acquisition conditions. Extensive experiments on 17 representative LMMs reveal substantial performance variation across tasks, document categories, and real-world conditions. Moreover, models with comparable overall accuracy often exhibit fundamentally different failure patterns under specific document factors, highlighting a significant gap between benchmark-level performance and reliable deployment in practical applications. The dataset and evaluation toolkit are publicly available at https://github.com/eioss/CC-OCR-V2.

多模态文档理解故障归因评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。