arXiv:2502.10410cs.CYcs.AI2025-02被引 7

用AI自动评估教学资源质量,提升AI生成课件的准确性与安全性。

Auto-Evaluation: A Critical Measure in Driving Improvements in Quality and Safety of AI-Generated Lesson Resources

  • 构建AI评估代理,自动检测课件中的题目难度
  • 通过与专家评价对比,准确率持续提升
  • 适合教育科技开发者和政策制定者参考

英国公立教育机构奥克国家学院拥有约13,000个由专业教师设计并质量保证的开放教育资料(OER),覆盖国家课程所有科目。基于此内容库,我们开发了免费可用的AI助教工具Aila,支持教师进行课程规划。同时,我们依据实证驱动的课程原则,对教学设计的每个环节进行了编码与示范。为在大规模上评估Aila生成课件的质量,我们构建了一个AI自动评估代理,用于推动输出质量的持续改进。通过对比人工评价与自动评估结果,不断优化该代理,其与专家人类评估者的吻合度得以提升。本文以多选题难度这一质量基准为例,展示这一迭代评估流程,并探讨其对类似项目及整个教育领域的潜在贡献。

原文摘要 · Abstract (English)

As a publicly funded body in the UK, Oak National Academy is in a unique position to innovate within this field as we have a comprehensive curriculum of approximately 13,000 open education resources (OER) for all National Curriculum subjects, designed and quality-assured by expert, human teachers. This has provided the corpus of content needed for building a high-quality AI-powered lesson planning tool, Aila, that is free to use and, therefore, accessible to all teachers across the country. Furthermore, using our evidence-informed curriculum principles, we have codified and exemplified each component of lesson design. To assess the quality of lessons produced by Aila at scale, we have developed an AI-powered auto-evaluation agent,facilitating informed improvements to enhance output quality. Through comparisons between human and auto-evaluations, we have begun to refine this agent further to increase its accuracy, measured by its alignment with an expert human evaluator. In this paper we present this iterative evaluation process through an illustrative case study focused on one quality benchmark - the level of challenge within multiple-choice quizzes. We also explore the contribution that this may make to similar projects and the wider sector.

AI教育自动评估教学设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。