arXiv:2412.19755cs.AI2024-12被引 1

评测大模型能否像人一样批改图文混合的简答题并给出合理反馈。

Can MLLMs generate human-like feedback in grading multimodal short answers?

  • 用大模型幻觉生成学生常见错误,构建2197条图文答案数据集。
  • 最高62.5%准确率判断答案对错,80.36%准确率判断图像相关性。
  • 首次用评分标准评估反馈质量,适合教育AI研究者参考。

在教育领域,传统自动简答评分(ASAG)主要针对纯文本回答。然而真实评估中常包含图文混合的回答。为此,我们提出多模态简答评分与反馈(MMSAF)问题,需联合评估文本与图表内容,并提供解释性反馈。由于数据规模与采集难度,构建代表性数据集困难。为此,我们开发自动化数据生成框架,利用大模型幻觉模拟学生常见错误,构建包含2197个实例的数据集。我们在3个STEM学科中评估4个多模态大语言模型(MLLMs),结果显示其在判断答案正确性(正确/部分正确/错误)上的准确率最高达62.5%,图像相关性评估准确率最高达80.36%。此外,通过9名标注员进行人类评估,涵盖5项指标,采用基于评分标准的方法评估反馈语义质量,而非依赖重叠度。研究揭示了各MLLM的适用性及现存缺陷。

原文摘要 · Abstract (English)

In education, the traditional Automatic Short Answer Grading (ASAG) with feedback problem has focused primarily on evaluating text-only responses. However, real-world assessments often include multimodal responses containing both diagrams and text. To address this limitation, we introduce the Multimodal Short Answer Grading with Feedback (MMSAF) problem, which requires jointly evaluating textual and diagrammatic content while also providing explanatory feedback. Collecting data representative of such multimodal responses is challenging due to both scale and logistical constraints. To mitigate this, we develop an automated data generation framework that leverages LLM hallucinations to mimic common student errors, thereby constructing a dataset of 2,197 instances. We evaluate 4 Multimodal Large Language Models (MLLMs) across 3 STEM subjects, showing that MLLMs achieve accuracies of up to 62.5% in predicting answer correctness (correct/partially correct/incorrect) and up to 80.36% in assessing image relevance. This also includes a human evaluation with 9 annotators across 5 parameters, including a rubric-based approach. The rubrics also serve as a way to evaluate the feedback quality semantically rather than using overlap-based approaches. Our findings highlight which MLLMs are better suited for such tasks while also pointing out to drawbacks of the remaining MLLMs.

多模态教育AI反馈生成大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。