arXiv:2507.03483cs.CLcs.AI2025-07NeurIPS被引 7

构建首个中英双语多学科推理数据集,评估大模型跨领域思维能力。

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

  • 设计双语多模态数据收集流程,含人工校验与可扩展框架
  • 覆盖300个学科、11万道题,提供高质量推理路径标注
  • 推出专用验证器,精准评估模型跨学科推理表现

本文提出BMMR,一个大规模双语、多模态、多学科推理数据集,用于推动和评估大视觉语言模型(LMMs)的发展。BMMR包含11万道大学水平题目,覆盖300个联合国教科文组织定义的学科,涵盖选择题、填空题和开放式问答,数据源包括书籍、考试与测验等印刷与数字媒体。所有数据通过人机协作的可扩展框架进行筛选与标注,每条样本均配有高质量推理路径。数据集分为两部分:BMMR-Eval含20,458个高质量样本,用于在中英文环境下全面评估模型知识与跨学科推理能力;BMMR-Train含88,991个样本,支持研究与训练。此外,我们提出基于过程的多学科验证器(BMMR-Verifier),实现对推理路径的精准细粒度评估。24个模型的实验证明:(i)即使是当前最优模型(如o3和Gemini-2.5-Pro)在BMMR-Eval上仍有显著提升空间;(ii)推理模型存在学科偏见,仅在特定领域优于通用大模型;(iii)开源模型仍落后于闭源模型;(iv)在BMMR-Train上微调可缩小差距。我们还利用BMMR-Verifier开展推理链分析与深度研究,揭示了当前大模型在多学科推理中的核心挑战。数据将公开发布,期望为社区提供重要参考。

原文摘要 · Abstract (English)

In this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 110k college-level questions spanning 300 UNESCO-defined subjects, spanning diverse formats-multiple-choice, fill-in-the-blank, and open-ended QA-and sourced from both print and digital media such as books, exams, and quizzes. All data are curated and filtered via a human-in-the-loop and scalable framework, and each instance is paired with a high-quality reasoning path. The dataset is organized into two parts: BMMR-Eval that comprises 20,458 high-quality instances to comprehensively assess LMMs' knowledge and reasoning across multiple disciplines in both Chinese and English; and BMMR-Train that contains 88,991 instances to support further research and development, extending the current focus on mathematical reasoning to diverse disciplines and domains. In addition, we propose the process-based multi-discipline verifier (i.e., BMMR-Verifier) for accurate and fine-grained evaluation of reasoning paths. Extensive experiments on 24 models reveal that (i) even SOTA models (e.g., o3 and Gemini-2.5-Pro) leave substantial headroom on BMMR-Eval; (ii) reasoning models exhibit discipline bias and outperform LMMs only on specific subjects; (iii) open-source models still trail their proprietary counterparts; and (iv) fine-tuning on BMMR-Train narrows this gap. Additionally, we conduct reasoning-chain analyses using BMMR-Verifier and other in-depth studies, uncovering the challenges LMMs currently face in multidisciplinary reasoning. We will release the data, and we hope our work can offer insights and contributions to the community.

多学科推理双语数据集大模型评估推理链分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。