用教育评价体系测试大模型是否是好科学老师。
Is your multimodal large language model a good science tutor?
- 设计教学评价框架与模拟学生模型,评估大模型的辅导能力。
- 发现解题能力强不等于教学好,优化后模型辅导效果显著提升。
- 适合教育AI研发者和对模型教学能力有要求的研究者。
多模态大语言模型(MLLM)在科学推理任务(如ScienceQA)中表现优异,但现有评测主要关注最终答案的准确性,忽视了教学价值。本文提出一个综合教育评价框架,结合模拟学生模型,评估MLLM作为科学导师的教学表现。基于ScienceQA训练集,构建强弱导师输出的成对比较数据集,采用多种偏好优化方法微调低效模型(Qwen2-VL-2B),使其教学能力提升。结果表明,强解题能力不等于高质量教学,以教学效果为导向的优化能生成更符合教育目标的导师模型。该方法为打造真正助教型多模态大模型提供了新路径。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) demonstrate impressive performance on scientific reasoning tasks (e.g., ScienceQA). However, most existing benchmarks focus narrowly on the accuracy of the final answer while ignoring other metrics. In particular, when applying MLLMs to educational contexts, the goal is not only correctness but also the ability to teach. In this paper, we propose a framework that evaluates MLLMs as science tutors using a comprehensive educational rubric and a simulated student model that judges the teaching performance of the tutors. Given a list of candidate MLLM science tutors, we use rubric-based student judgments to produce a range of tutor performance scores, identifying both strong and weak tutors. Using the training section of the ScienceQA dataset, we then construct a data set of pairwise comparisons between the outputs of strong and weak tutors. This enables us to apply multiple preference optimization methods to fine-tune an underperforming tutor model (Qwen2-VL-2B) into more effective ones. Our results also show that strong problem-solving skills do not guarantee high-quality tutoring and that performance optimization-guided refinements can yield more educationally aligned tutor models. This approach opens avenues for building MLLMs that serve not only as problem solvers, but as genuinely helpful educational assistants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。