评测教育大模型的知识、技能与态度,更真实反映其教学能力。
OpenLearnLM Benchmark: A Unified Framework for Evaluating Knowledge, Skill, and Attitude in Educational Large Language Models
- 从教育评估理论出发,分知识、技能、态度三维度构建评测框架。
- 覆盖12.4万+题目,涵盖多学科和布卢姆分类法难度层级。
- 发现无模型全优,凸显多维评测必要性,适合教育AI研究者使用。
大型语言模型在教育领域应用日益广泛,但现有评测基准侧重单一技能,缺乏学习科学基础。我们提出OpenLearnLM基准,一个基于教育评估理论的统一框架,从知识(课程对齐内容与教学理解)、技能(基于四层中心-角色-场景-子场景结构的情境化能力)和态度(一致性与欺骗抵抗性)三个维度评估大模型。该基准包含超过12.4万个题目,覆盖多个学科、教育角色和布卢姆分类法中的不同难度层级。知识维度采用权威基准的真实测评题,态度维度借鉴Anthropic的对齐伪装方法,在不同监控条件下检测行为不一致。对七款前沿模型的评估显示:Claude-Opus-4.5虽知识较弱,但在实践技能上表现优异;Grok-4.1-fast知识领先,但存在对齐隐患。值得注意的是,没有任何模型在所有维度上全面领先,验证了多维度评测的必要性。OpenLearnLM提供了一个开源、全面的框架,助力提升大模型在真实教育场景中的准备度。
原文摘要 · Abstract (English)
Large Language Models are increasingly deployed as educational tools, yet existing benchmarks focus on narrow skills and lack grounding in learning sciences. We introduce OpenLearnLM Benchmark, a theory-grounded framework evaluating LLMs across three dimensions derived from educational assessment theory: Knowledge (curriculum-aligned content and pedagogical understanding), Skills (scenario-based competencies organized through a four-level center-role-scenario-subscenario hierarchy), and Attitude (alignment consistency and deception resistance). Our benchmark comprises 124K+ items spanning multiple subjects, educational roles, and difficulty levels based on Bloom's taxonomy. The Knowledge domain prioritizes authentic assessment items from established benchmarks, while the Attitude domain adapts Anthropic's Alignment Faking methodology to detect behavioral inconsistency under varying monitoring conditions. Evaluation of seven frontier models reveals distinct capability profiles: Claude-Opus-4.5 excels in practical skills despite lower content knowledge, while Grok-4.1-fast leads in knowledge but shows alignment concerns. Notably, no single model dominates all dimensions, validating the necessity of multi-axis evaluation. OpenLearnLM provides an open, comprehensive framework for advancing LLM readiness in authentic educational contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。