arXiv:2507.22947cs.CYcs.CL2025-07被引 2

为教育场景中的大模型评估提供自动化框架,解决评测标准缺失问题。

ELMES: An Automated Framework for Evaluating Large Language Models in Educational Scenarios

  • 通过配置文件构建多智能体对话,灵活设计教育场景。
  • 采用大模型判官机制量化教学能力,实现主观评价客观化。
  • 覆盖四大教育场景,帮助研究者发现模型的上下文优势与短板。

大语言模型(LLM)为教育带来变革性机遇,但评估标准在不同教育场景间差异显著,许多新兴场景缺乏合适评测指标。现有基准主要衡量通用智能,而非教学能力。为此,我们提出ELMES——一个开源自动化评估框架,专为教育场景中的LLM设计。该框架采用模块化架构,通过简单配置文件即可生成动态多智能体对话,无需复杂编程。其混合评估引擎利用大模型作为评判者,客观量化传统上主观的教学指标。我们在四大关键教育场景中系统评估了前沿大模型:知识点解释、引导式解题教学、跨学科教案生成和情境化问题生成,采用教育专家合作开发的细粒度指标。结果揭示各模型在不同场景下的能力分布特征,展现特定上下文中的优势与局限。ELMES为教育工作者和研究人员提供了低门槛的评估工具,显著降低多样化教育应用的适配成本,推动大模型在教学实践中的落地。框架已公开于https://github.com/sii-research/elmes.git。

原文摘要 · Abstract (English)

The emergence of Large Language Models (LLMs) presents transformative opportunities for education, generating numerous novel application scenarios. However, significant challenges remain: evaluation metrics vary substantially across different educational scenarios, while many emerging scenarios lack appropriate assessment metrics. Current benchmarks predominantly measure general intelligence rather than pedagogical capabilities. To address this gap, we introduce ELMES, an open-source automated evaluation framework specifically designed for assessing LLMs in educational settings. ELMES features a modular architecture that enables researchers to create dynamic, multi-agent dialogues through simple configuration files, facilitating flexible scenario design without requiring extensive programming expertise. The framework incorporates a hybrid evaluation engine that objectively quantifies traditionally subjective pedagogical metrics using an LLM-as-a-Judge methodology. We conduct systematic benchmarking of state-of-the-art LLMs across four critical educational scenarios: Knowledge Point Explanation, Guided Problem-Solving Teaching, Interdisciplinary Lesson Plan Generation, and Contextualized Question Generation, employing fine-grained metrics developed in collaboration with education specialists. Our results demonstrate distinct capability distributions among models, revealing context-specific strengths and limitations. ELMES provides educators and researchers with an accessible evaluation framework that significantly reduces adaptation barriers for diverse educational applications while advancing the practical implementation of LLMs in pedagogy. The framework is publicly available at \emph{https://github.com/sii-research/elmes.git}.

大模型评估教育AI自动化评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。