arXiv:2505.17050cs.CLcs.AI2025-05被引 9

构建PBLBench基准,用专家评估提升教育AI评测可靠性

Towards Robust Evaluation of STEM Education: Leveraging MLLMs in Project-Based Learning

  • 用层次分析法构建专家驱动的评分体系,确保评估科学性
  • 15个主流模型在复杂任务中最高仅达59%排名准确率
  • 专为项目式学习设计,适合教育AI与教师辅助系统研究者

项目式学习(PBL)涉及多种高度相关的多模态数据,是STEM教育中的重要方法。随着多模态大语言模型(MLLMs)的发展,研究者开始探索其在信息检索、知识理解与数据生成等教育任务中的潜力。然而,现有基准缺乏自由输出结构和严谨的人类专家验证流程,难以真实评估教育场景下的表现。同时,因模型幻觉与不稳定性,自动化教师辅助流程进展缓慢。为此,我们提出PBLBench,一个面向领域知识与长上下文理解的新型基准,旨在挑战模型在贴近人类专家任务中的复杂推理能力。为建立可靠真值,采用层次分析法(AHP),通过专家主导的成对比较构建结构化权重评价体系。我们在15个领先的MLLM/LLM上评估该基准,发现最先进模型最高仅达59%的排名准确率,凸显其挑战性。我们相信PBLBench将推动更强大AI代理的发展,最终减轻教师负担,提升教育效率。

原文摘要 · Abstract (English)

Project-Based Learning (PBL) involves a variety of highly correlated multimodal data, making it a vital educational approach within STEM disciplines. With the rapid development of multimodal large language models (MLLMs), researchers have begun exploring their potential to enhance tasks such as information retrieval, knowledge comprehension, and data generation in educational settings. However, existing benchmarks fall short in providing both a free-form output structure and a rigorous human expert validation process, limiting their effectiveness in evaluating real-world educational tasks. Additionally, few methods have developed automated pipelines to assist with the complex responsibilities of teachers leveraging MLLMs, largely due to model hallucination and instability, which lead to unreliable implementation. To address this gap, we introduce PBLBench, a novel benchmark designed to evaluate complex reasoning grounded in domain-specific knowledge and long-context understanding, thereby challenging models with tasks that closely resemble those handled by human experts. To establish reliable ground truth, we adopt the Analytic Hierarchy Process (AHP), utilizing expert-driven pairwise comparisons to derive structured and weighted evaluation criteria. We assess the performance of 15 leading MLLMs/LLMs using PBLBench and demonstrate that even the most advanced models achieve only 59% rank accuracy, underscoring the significant challenges presented by this benchmark. We believe PBLBench will serve as a catalyst for the development of more capable AI agents, ultimately aiming to alleviate teacher workload and enhance educational productivity.

教育AI多模态模型项目式学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。