用自动化框架构建教育场景细粒度评估标准,解决长尾教学评价难题。
Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios

- 通过多智能体互动与自进化模块,自动构建场景化评估规则
- 覆盖330个教育场景,含1000+二级指标,验证顶尖模型在创造力与价值观上的差异
- 适合教育AI研发者、评测人员及关注教学能力评估的研究者
评估大语言模型在教育中的表现,需衡量其教学能力而非仅知识掌握。现有基准侧重通用正确性或依赖人工设计的评分标准,难以扩展至长尾教育场景。本文提出Elmes*,一个端到端框架,用于构建、优化和应用细粒度的场景特定评估标准。Elmes*结合声明式多智能体引擎(教师-学生-评审)与SceneGen模块,从专家定义的教学维度协同优化评估指标与测试数据。基于Elmes*,我们构建了Edu-330数据集,涵盖11个学科、3个年级段、10种任务类型,共330个场景,超过1000个二级指标。在Edu-330及四个专家制定的黄金标准场景上实验表明:教育能力具有多维性——顶级模型主要差异体现在创造力与价值观融合;知识强的模型可能在苏格拉底式引导中失败;教育专用模型InnoSpark获得最高的人工平均评分。使用大模型作为评审者可保持与人类相当的排名一致性,且评分方差更小,但存在如自我偏好等评委特异性偏差。消融实验显示,专家评分少样本锚定提升人机对齐效果,而推理强化与贪婪解码策略则依赖模型。Elmes*为基于教学原理的大语言模型评估提供可扩展的诊断基础设施。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know. Existing benchmarks emphasize domain-general correctness or depend on manually designed rubrics that scale poorly to long-tail pedagogical scenarios. We introduce Elmes*, an end-to-end framework for constructing, refining, and applying fine-grained scenario-specific rubrics. Elmes* combines a declarative multi-agent engine for teacher--student--judge interactions with SceneGen, a self-evolving module that co-optimizes evaluation criteria and test data from expert-defined pedagogical dimensions. Using Elmes*, we build Edu-330, covering 330 scenarios across 11 subjects, 3 grade bands, and 10 task types, with over 1{,}000 second-level indicators. Experiments on Edu-330 and four expert-authored gold-standard scenarios show that educational capability is multidimensional: top-tier LLMs differ mainly in creativity and values integration, knowledge-strong models may fail at Socratic scaffolding, and the education-specialized InnoSpark achieves the best human-evaluated average score. LLM judges preserve human-comparable rankings with much lower scoring variance, but exhibit judge-specific biases such as self-preference. Ablations show that expert-scored few-shot anchoring improves human--LLM alignment, while reasoning enforcement and greedy decoding are model-dependent. Elmes* thus provides scalable diagnostic infrastructure for pedagogically grounded LLM evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。