arXiv:2504.05782cs.CVcs.AI2025-04被引 16

构建涵盖12年制多学科的图文推理评测集,评估大模型真实解题能力。

MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models

  • 基于K-12真实考题构建多学科图文推理数据集
  • 含14万例题目,覆盖6大学科与多级难度
  • 引入动态评测框架,防止数据泄露问题

多模态推理将语言与视觉信息融合以解决问题和决策,是人类智能的核心,也是迈向通用人工智能的关键。然而当前对多模态大模型(MLLMs)多模态推理能力的评估仍不充分。现有评测基准普遍存在数据量小、领域窄、知识分布无序等问题。为此,我们提出MDK12-Bench,一个基于真实K-12考试的跨学科评测基准,涵盖数学、物理、化学、生物、地理和信息科学六大学科,包含14万条推理实例,覆盖从小学到12年级的多种难度层级。该数据集包含6,827个细粒度知识点标注,配有详细答案解析、难度标签及跨年份划分,为全面评估提供坚实基础。此外,我们设计了一种新颖的动态评估框架,通过动态变换题型、问题类型和图像风格来缓解数据污染问题。在该基准上的大量实验揭示了当前MLLMs在多模态推理中的显著局限性,为下一代模型的发展提供了重要启示。数据与代码已开源:https://github.com/LanceZPF/MDK12。

原文摘要 · Abstract (English)

Multimodal reasoning, which integrates language and visual cues into problem solving and decision making, is a fundamental aspect of human intelligence and a crucial step toward artificial general intelligence. However, the evaluation of multimodal reasoning capabilities in Multimodal Large Language Models (MLLMs) remains inadequate. Most existing reasoning benchmarks are constrained by limited data size, narrow domain coverage, and unstructured knowledge distribution. To close these gaps, we introduce MDK12-Bench, a multi-disciplinary benchmark assessing the reasoning capabilities of MLLMs via real-world K-12 examinations. Spanning six disciplines (math, physics, chemistry, biology, geography, and information science), our benchmark comprises 140K reasoning instances across diverse difficulty levels from primary school to 12th grade. It features 6,827 instance-level knowledge point annotations based on a well-organized knowledge structure, detailed answer explanations, difficulty labels and cross-year partitions, providing a robust platform for comprehensive evaluation. Additionally, we present a novel dynamic evaluation framework to mitigate data contamination issues by bootstrapping question forms, question types, and image styles during evaluation. Extensive experiment on MDK12-Bench reveals the significant limitation of current MLLMs in multimodal reasoning. The findings on our benchmark provide insights into the development of the next-generation models. Our data and codes are available at https://github.com/LanceZPF/MDK12.

多模态推理评测基准教育AI大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。