arXiv:2604.05005cs.CYcs.AI2026-04被引 1

评测大模型生成带图解的K12 STEM教育内容能力,兼顾图文一致性与教学效果。

EduIllustrate: Towards Scalable Automated Generation Of Multimodal Educational Content

  • 构建多模态教育内容生成基准,含230道跨学科跨年级题目
  • Gemini 3.0 Pro Preview表现最佳,准确率达87.8%,成本最低仅$0.12/题
  • 顺序锚定机制提升视觉一致性13%,且大幅降低生成成本

大语言模型在教育辅助中广泛应用,但现有评估仍集中于问答与辅导任务。在多媒体教学内容生成方面存在关键空白:即能否生成图文一致、几何准确、步骤清晰的解释。我们提出EduIllustrate,一个针对K-12 STEM问题的文本-图表交替生成评估基准。该基准包含230道涵盖五个学科和三个年级水平的问题,采用标准化生成流程,通过顺序锚定机制确保跨图表视觉一致性,并基于多媒体学习理论设计八维评价体系,覆盖文本与视觉质量。对十款LLM的评估显示性能差异显著:Gemini 3.0 Pro Preview以87.8%的准确率领先,Kimi-K2.5则实现最佳成本效益(80.8%准确率,每题仅$0.12)。工作流消融实验表明,顺序锚定使视觉一致性提升13%,同时成本降低94%。20位专家的人工评估验证了大模型作为评分者在客观维度上的可靠性(ρ≥0.83),但在主观视觉评估上仍有局限。

原文摘要 · Abstract (English)

Large language models are increasingly used as educational assistants, yet evaluation of their educational capabilities remains concentrated on question-answering and tutoring tasks. A critical gap exists for multimedia instructional content generation -- the ability to produce coherent, diagram-rich explanations that combine geometrically accurate visuals with step-by-step reasoning. We present EduIllustrate, a benchmark for evaluating LLMs on interleaved text-diagram explanation generation for K-12 STEM problems. The benchmark comprises 230 problems spanning five subjects and three grade levels, a standardized generation protocol with sequential anchoring to enforce cross-diagram visual consistency, and an 8-dimension evaluation rubric grounded in multimedia learning theory covering both text and visual quality. Evaluation of ten LLMs reveals a wide performance spread: Gemini 3.0 Pro Preview leads at 87.8\%, while Kimi-K2.5 achieves the best cost-efficiency (80.8\% at \\$0.12/problem). Workflow ablation confirms sequential anchoring improves Visual Consistency by 13\% at 94\% lower cost. Human evaluation with 20 expert raters validates LLM-as-judge reliability for objective dimensions ($ρ\geq 0.83$) while revealing limitations on subjective visual assessment.

教育AI多模态生成大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。