arXiv:2608.09548cs.CLcs.AI2026-08

首个综合评估教育类大模型能力的基准,涵盖准确性、安全性、教学性与育人目标。

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

  • 构建四维评测体系,统一标准评估模型在教育场景下的综合表现。
  • 中文模型在安全性能上领先,尤其在本土规范内容上优势明显。
  • 专用教育模型未显著提升教学效果,高阶育人能力普遍存在盲区。

大语言模型正被广泛应用于教育领域,作为导师、助教和内容生成工具。这些角色对模型提出的要求远超普通问答:需兼具准确性、敏感指令下的安全性、教学实用性及与教育目标的一致性。现有基准多孤立评估各项指标,无法整体衡量教育适用性。本文提出ELBench,首个统一协议下评估四项核心能力(通用能力、安全可信、基础教育、高阶培育)的多维度基准,融合公开数据与新构建的安全与培育数据。评估九个模型(七个主流通用模型和两个教育专用模型),得出三点发现:第一,模块级表现比总分更具信息量——前六名模型总体得分无显著差异,但各模块优劣不一,且安全性能与教学实用性呈强负相关(r = -0.83);第二,中国开发模型在安全模块领先,尤其在地区性规范内容上优势最大,对普适性危害内容仍有优势但减弱;第三,两个教育专用模型未在任一教育模块领先,高阶培育模块所有模型均呈现系统性盲区:在结构化判断任务中趋同于同一非参考选项,偏好教学风格而非目标契合度,导致该模块评分普遍偏低且无法区分模型。这提示领域后训练是否跟得上前沿系统仍存疑问。

原文摘要 · Abstract (English)

Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed to be accurate, safe under sensitive prompts, instructionally useful, and aligned with pedagogical goals at the same time. Existing benchmarks evaluate these requirements largely in isolation, so none assesses education-facing suitability as an integrated profile. We introduce ELBench, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data. We evaluate nine models, seven frontier general-purpose systems and two education-specialized variants, and report three findings. First, module-level profiles are more informative than a single aggregate: the top six models are statistically indistinguishable on overall score, yet their module leaders differ substantially, and safety is anti-correlated with practical teaching (r = -0.83). Second, the Chinese-developed models lead the safety module, the most discriminative in the suite; this advantage is largest on region-specific normative content and narrows, but does not vanish, on universal-harm content. Third, the two education-specialized models lead neither education module, and on High-Level Cultivation all models share a systematic blind spot: on the structured judgment task they converge on the same non-reference option, favoring pedagogical style over fit to the stated goal, so the module scores uniformly low and does not separate models. This raises, but does not resolve, whether domain post-training keeps pace with frontier systems on education tasks.

大模型评测教育AI安全对齐基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。