arXiv:2602.13243cs.CYcs.AI2026-02被引 3

验证大模型对中小学科学教材的评价是否靠谱,为智能教学设计提供依据。

Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K--12 Science Instructional Materials

  • 用专家评审对比三款大模型对12套教材的评分与理由。
  • 发现模型在标准匹配上准确率高,但对教学情境理解有明显偏差。
  • 成果将用于训练更懂教育的生成式AI助手,适合教育科技研发者参考。

为中小学科学课程设计高质量、符合标准的教学材料耗时且需专业知识。本研究考察人类专家在评审AI生成的评价时关注什么,旨在将这些洞见转化为未来基于生成式AI的教学材料设计代理的设计原则。研究从OpenSciEd和基于项目学习的多重素养等可信项目中选取12个涵盖生命、物理与地球科学的优质课程单元,使用包含9项评估指标的EQuIP量规,分别提示GPT-4o、Claude与Gemini对每个单元进行数值评分与书面论证,共生成648条评价输出。两位科学教育专家独立评审全部结果,标记评分与理由的一致性(1)或不一致(0),并提供对AI推理的定性反思。该过程揭示了大模型判断与专家视角在哪些方面一致或偏离,暴露了其推理优势、盲区及情境理解的细微差异。这些发现将直接指导领域专用生成式AI代理的开发,以支持中小学科学教育中高质量教学材料的设计。

原文摘要 · Abstract (English)

Designing high-quality, standards-aligned instructional materials for K--12 science is time-consuming and expertise-intensive. This study examines what human experts notice when reviewing AI-generated evaluations of such materials, aiming to translate their insights into design principles for a future GenAI-based instructional material design agent. We intentionally selected 12 high-quality curriculum units across life, physical, and earth sciences from validated programs such as OpenSciEd and Multiple Literacies in Project-based Learning. Using the EQuIP rubric with 9 evaluation items, we prompted GPT-4o, Claude, and Gemini to produce numerical ratings and written rationales for each unit, generating 648 evaluation outputs. Two science education experts independently reviewed all outputs, marking agreement (1) or disagreement (0) for both scores and rationales, and offering qualitative reflections on AI reasoning. This process surfaces patterns in where LLM judgments align with or diverge from expert perspectives, revealing reasoning strengths, gaps, and contextual nuances. These insights will directly inform the development of a domain-specific GenAI agent to support the design of high-quality instructional materials in K--12 science education.

生成式AI教育评估大模型评测K12教育

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。