用多智能体框架评估学生科学草图,提升评估准确性与教学价值。
SketchMind: A Multi-Agent Cognitive Framework for Assessing Student-Drawn Scientific Sketches
- 分模块智能体协同解析评分标准、理解草图、对齐认知水平并生成修改建议。
- 在3575幅草图上平均准确率达77.1%,较基线提升21.4个百分点。
- 反馈质量获专家认可,适合用于支持学生概念性思维发展的智能教育系统。
科学草图(如模型图)是洞察学生概念理解的有力工具,但人工智能对这类自由形式、视觉多样性的作品进行自动化评估仍面临重大挑战。现有方法多将草图评价视为图像分类或单一视觉语言模型任务,缺乏可解释性、教学契合度及跨认知层级的适应性。为此,我们提出SketchMind——一种基于认知理论的多智能体框架,用于评估与优化学生绘制的科学草图。该框架包含负责评分标准解析、草图感知、认知对齐及迭代反馈与草图修改的模块化智能体,实现个性化、透明化的评估。我们在涵盖六个科学测评项(最高达布卢姆认知层次第6级)的3,575幅学生草图数据集上进行评估。相比无SRG的GPT-4o基线(平均准确率55.6%),集成SRG后达到77.1%平均准确率(+21.4%绝对提升)。多智能体协同结合SRG进一步提升性能,例如GPT-4.1在草图预测准确率上平均提升8.9%,优于所有单智能体流水线。人类评估者对使用GPT-4.1的SketchMind生成的反馈与共创草图评分达4.1/5,显著高于基线模型(如GPT-4o为2.3)。专家认为该系统可通过引导修订有效促进概念发展。代码与(待审批)数据集将公开以支持可复现性与未来研究。
原文摘要 · Abstract (English)
Scientific sketches (e.g., models) offer a powerful lens into students' conceptual understanding, yet AI-powered automated assessment of such free-form, visually diverse artifacts remains a critical challenge. Existing solutions often treat sketch evaluation as either an image classification task or monolithic vision-language models, which lack interpretability, pedagogical alignment, and adaptability across cognitive levels. To address these limitations, we present SketchMind, a cognitively grounded, multi-agent framework for evaluating and improving student-drawn scientific sketches. SketchMind comprises modular agents responsible for rubric parsing, sketch perception, cognitive alignment, and iterative feedback with sketch modification, enabling personalized and transparent evaluation. We evaluate SketchMind on a curated dataset of 3,575 student-generated sketches across six science assessment items with different highest order of Bloom's level that require students to draw models to explain phenomena. Compared to baseline GPT-4o performance without SRG (average accuracy: 55.6%), and with SRG integration achieves 77.1% average accuracy (+21.4% average absolute gain). We also demonstrate that multi-agent orchestration with SRG enhances SketchMind performance, for example, GPT-4.1 gains an average 8.9% increase in sketch prediction accuracy, outperforming single-agent pipelines across all items. Human evaluators rated the feedback and co-created sketches generated by \textsc{SketchMind} with GPT-4.1, which achieved an average of 4.1 out of 5, significantly higher than those of baseline models (e.g., 2.3 for GPT-4o). Experts noted the system's potential to meaningfully support conceptual growth through guided revision. Our code and (pending approval) dataset will be released to support reproducibility and future research in AI-driven education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。