首个面向AI数学辅导的多模态评测基准,评估模型解题与引导能力。
MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
- 构建685道含关键步骤的数学题,支持分维度精细评估。
- 开源模型表现显著低于闭源模型,距人类导师仍有差距。
- 适合研究教育AI、多模态推理与自动评分的开发者使用。
有效的数学辅导不仅需要解题,还需诊断学生困难并逐步引导。尽管多模态大语言模型(MLLMs)展现潜力,现有基准大多忽略这些教学技能。我们提出MMTutorBench,首个针对AI数学辅导的基准,包含685道围绕教学关键步骤设计的问题。每道题配有特定评分标准,支持六个维度的细粒度评估,并划分为三类任务:洞察发现、操作构型、操作执行。我们评估了12个主流MLLMs,发现闭源系统表现优于开源系统,与人类导师相比仍有较大提升空间;同时在不同输入变体下表现出一致趋势:OCR流程降低辅导质量,少样本提示增益有限,基于评分标准的LLM作为裁判具有高度可靠性。这些结果凸显了MMTutorBench在推动AI辅导发展中的挑战性与诊断价值。
原文摘要 · Abstract (English)
Effective math tutoring requires not only solving problems but also diagnosing students' difficulties and guiding them step by step. While multimodal large language models (MLLMs) show promise, existing benchmarks largely overlook these tutoring skills. We introduce MMTutorBench, the first benchmark for AI math tutoring, consisting of 685 problems built around pedagogically significant key-steps. Each problem is paired with problem-specific rubrics that enable fine-grained evaluation across six dimensions, and structured into three tasks-Insight Discovery, Operation Formulation, and Operation Execution. We evaluate 12 leading MLLMs and find clear performance gaps between proprietary and open-source systems, substantial room compared to human tutors, and consistent trends across input variants: OCR pipelines degrade tutoring quality, few-shot prompting yields limited gains, and our rubric-based LLM-as-a-Judge proves highly reliable. These results highlight both the difficulty and diagnostic value of MMTutorBench for advancing AI tutoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。