arXiv:2505.12766cs.CV2025-05被引 6

构建新基准,测试大模型从图文中推理复杂逻辑问题的能力

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?

  • 设计六类视觉场景,150道题覆盖多种逻辑推理挑战
  • 现有大模型在复杂推理任务上表现仍不理想,需提升
  • 适合研究多模态推理、OCR应用的学者与工程师参考

大型多模态模型(LMMs)已展现出强大的光学字符识别(OCR)相关能力。现有评测基准主要聚焦于简单的视觉问答或图文解析任务,但对模型基于OCR信息解决复杂逻辑推理问题的能力研究不足。为此,我们提出Reasoning-OCR基准,要求模型在丰富的视觉-文本线索基础上完成复杂推理任务。该基准涵盖六类视觉场景,包含150道精心设计的问题,分为六类推理挑战,且尽量避免依赖领域专有知识的影响。评估结果揭示了主流私有与开源LMM在不同推理挑战中的表现差异,凸显提升其推理能力的紧迫性。我们希望Reasoning-OCR能推动未来基于OCR线索的复杂推理研究。数据集已公开于 https://github.com/Hxyz-123/ReasoningOCR。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) have become increasingly versatile, accompanied by impressive Optical Character Recognition (OCR) related capabilities. Existing OCR-related benchmarks emphasize evaluating LMMs' abilities of relatively simple visual question answering, visual-text parsing, etc. However, the extent to which LMMs can deal with complex logical reasoning problems based on OCR cues is relatively unexplored. To this end, we introduce the Reasoning-OCR benchmark, which challenges LMMs to solve complex reasoning problems based on the cues that can be extracted from rich visual-text. Reasoning-OCR covers six visual scenarios and encompasses 150 meticulously designed questions categorized into six reasoning challenges. Additionally, Reasoning-OCR minimizes the impact of field-specialized knowledge. Our evaluation offers some insights for proprietary and open-source LMMs in different reasoning challenges, underscoring the urgent to improve the reasoning performance. We hope Reasoning-OCR can inspire and facilitate future research on enhancing complex reasoning ability based on OCR cues. Reasoning-OCR is publicly available at https://github.com/Hxyz-123/ReasoningOCR.

多模态推理OCR逻辑推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。