arXiv:2505.02018cs.CV2025-05ICML被引 37

构建多学科跨语言的研究生级推理评测集,评估大模型复杂推理能力。

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

  • 设计涵盖108个学科的双语推理题库,覆盖语言与多模态模型
  • 顶尖模型如OpenAI o1在多模态任务中仅达53.2%准确率
  • 适合评估大模型在真实复杂场景下的综合推理能力

推理是智能的核心,能够整合已有知识解决复杂问题。尽管取得显著进展,现有推理评测基准常无法严格评估真实世界复杂问题所需的细微推理能力,尤其在多学科和多模态情境下。本文提出一个研究生水平、多学科、中英双语的推理评测基准R-Bench,用于评估语言模型与多模态模型的推理能力。R-Bench包含1,094道语言模型题目(覆盖108个学科)和665道多模态模型题目(覆盖83个学科),均经过精心设计以确保难度校准、学科平衡和跨语言对齐,可作为奥林匹克级多学科评测基准。我们评估了包括OpenAI o1、GPT-4o、DeepSeek-R1等主流模型。实验表明,先进模型在复杂推理任务中表现不佳,尤其在多模态推理方面。即使最先进模型OpenAI o1在多模态评测中也仅达到53.2%准确率。数据与代码已公开。

原文摘要 · Abstract (English)

Reasoning stands as a cornerstone of intelligence, enabling the synthesis of existing knowledge to solve complex problems. Despite remarkable progress, existing reasoning benchmarks often fail to rigorously evaluate the nuanced reasoning capabilities required for complex, real-world problemsolving, particularly in multi-disciplinary and multimodal contexts. In this paper, we introduce a graduate-level, multi-disciplinary, EnglishChinese benchmark, dubbed as Reasoning Bench (R-Bench), for assessing the reasoning capability of both language and multimodal models. RBench spans 1,094 questions across 108 subjects for language model evaluation and 665 questions across 83 subjects for multimodal model testing in both English and Chinese. These questions are meticulously curated to ensure rigorous difficulty calibration, subject balance, and crosslinguistic alignment, enabling the assessment to be an Olympiad-level multi-disciplinary benchmark. We evaluate widely used models, including OpenAI o1, GPT-4o, DeepSeek-R1, etc. Experimental results indicate that advanced models perform poorly on complex reasoning, especially multimodal reasoning. Even the top-performing model OpenAI o1 achieves only 53.2% accuracy on our multimodal evaluation. Data and code are made publicly available at here.

推理评测多模态大模型跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。