评测并提升大模型生成教育可视化的能力,让讲解更符合学习规律。
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
- 设计多领域多层级测评基准,评估模型的视觉推理能力
- 提出协作式多智能体框架,实现从理解到可视化输出的全流程优化
- 适合教育科技研发者、智能教学系统设计者使用
尽管扩散模型和大视觉语言模型等基础模型已在教育场景广泛应用,但其生成具有教学有效性的视觉解释能力仍有限。现有方法多聚焦文本推理,忽视结构化可解释可视化对概念理解的关键作用。为此,我们提出EduVisBench——一个跨学科、多层级的基准测试,包含需视觉支撑求解的多样化STEM问题,并基于教育理论设计细粒度评估标准。实证分析表明,现有模型在分解复杂推理并转化为符合人类认知过程的视觉表达方面普遍表现不佳。为此,我们提出EduVisAgent,一种多智能体协同框架,通过专用智能体分工完成教学规划、推理分解、元认知提示与可视化设计。实验结果表明,EduVisAgent显著优于所有基线,性能提升40.2%,生成的可视化更具教育一致性。EduVisBench与EduVisAgent已开源:https://github.com/aiming-lab/EduVisBench 及 https://github.com/aiming-lab/EduVisAgent。
原文摘要 · Abstract (English)
While foundation models (FMs), such as diffusion models and large vision-language models (LVLMs), have been widely applied in educational contexts, their ability to generate pedagogically effective visual explanations remains limited. Most existing approaches focus primarily on textual reasoning, overlooking the critical role of structured and interpretable visualizations in supporting conceptual understanding. To better assess the visual reasoning capabilities of FMs in educational settings, we introduce EduVisBench, a multi-domain, multi-level benchmark. EduVisBench features diverse STEM problem sets requiring visually grounded solutions, along with a fine-grained evaluation rubric informed by pedagogical theory. Our empirical analysis reveals that existing models frequently struggle with the inherent challenge of decomposing complex reasoning and translating it into visual representations aligned with human cognitive processes. To address these limitations, we propose EduVisAgent, a multi-agent collaborative framework that coordinates specialized agents for instructional planning, reasoning decomposition, metacognitive prompting, and visualization design. Experimental results show that EduVisAgent substantially outperforms all baselines, achieving a 40.2% improvement and delivering more educationally aligned visualizations. EduVisBench and EduVisAgent are available at https://github.com/aiming-lab/EduVisBench and https://github.com/aiming-lab/EduVisAgent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。