arXiv:2505.17163cs.LGcs.AI2025-05被引 27

首个系统评估多模态大模型在图文文本推理能力的基准,发现顶尖模型准确率仍不足50%。

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning

  • 构建1069个标注样本,覆盖6类核心推理能力与18项实际任务
  • 首次提供分步推理过程标注,可同时评估答案与思维链
  • 揭示当前最强模型在复杂图文推理上普遍表现不佳,亟待突破

近年来,多模态慢思考系统在多种视觉推理任务中表现出色。然而,由于缺乏专门且系统的基准,其在富文本图像推理任务中的能力仍鲜有研究。为此,我们提出OCR-Reasoning,一个全新的基准,旨在系统评估多模态大语言模型在富文本图像推理任务中的表现。该基准包含1,069个由人工标注的样本,涵盖6种核心推理能力与18项实际应用场景。与现有基准仅提供最终答案不同,本基准还提供了详细的分步推理过程,支持对模型输出的答案及推理路径进行双重评估,实现对富文本推理能力的全面衡量。基于此基准,我们对最新MLLMs进行了全面评估,结果表明即使最先进的模型在该基准上的准确率也均未超过50%,说明富文本图像推理仍是亟待解决的重大挑战。基准与评估脚本已公开于https://github.com/SCUT-DLVCLab/OCR-Reasoning。

原文摘要 · Abstract (English)

Recent advancements in multimodal slow-thinking systems have demonstrated remarkable performance across various visual reasoning tasks. However, their capabilities in text-rich image reasoning tasks remain understudied due to the absence of a dedicated and systematic benchmark. To address this gap, we propose OCR-Reasoning, a novel benchmark designed to systematically assess Multimodal Large Language Models on text-rich image reasoning tasks. Specifically, OCR-Reasoning comprises 1,069 human-annotated examples spanning 6 core reasoning abilities and 18 practical reasoning tasks in text-rich visual scenarios. Unlike existing text-rich image understanding benchmarks that only provide a final answer, this benchmark additionally provides a detailed step-by-step reasoning process. This dual annotation enables the evaluation of both the models' final answers and their reasoning processes, thereby offering a holistic assessment of text-rich reasoning capabilities. By leveraging this benchmark, we conducted a comprehensive evaluation of the latest MLLMs. Our results demonstrate that even the most advanced MLLMs exhibit substantial difficulties in text-rich image reasoning tasks, with none achieving an accuracy above 50\% on our benchmark, indicating that the challenges of text-rich image reasoning are an urgent issue to be addressed. The benchmark and evaluation scripts are available at https://github.com/SCUT-DLVCLab/OCR-Reasoning.

多模态图文推理评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。