arXiv:2512.00290cs.CLcs.AI2025-12被引 1

构建中文教育领域分级认知评估体系,精准测试大模型教学能力

EduEval: A Hierarchical Cognitive Benchmark for Evaluating Large Language Models in Chinese Education

  • 基于布鲁姆与韦伯框架设计六维认知分类体系
  • 涵盖11000+真实题目的多模态教育任务,覆盖小学到高中
  • 发现开源模型在复杂推理上优于闭源模型,需分维度优化

大型语言模型在教育应用中潜力巨大,但未经检验的部署可能影响教育质量,亟需严格评估。本文提出EduEval,一个面向中文K-12教育的分层评估基准。该基准有三大贡献:(1) 认知框架:提出EduAbility Taxonomy,融合布鲁姆分类学与韦伯深度认知模型,涵盖记忆、理解、应用、推理、创造与伦理六大认知维度;(2) 真实性:整合真实考试题、课堂对话、学生作文及专家设计任务,反映真实教育挑战;(3) 规模:包含24类任务,超11,000个题目,覆盖小学至高中。我们评估了14个主流LLMs在零样本与少样本设置下的表现,发现模型在事实性任务上表现良好,但在课堂对话分类和创造性内容生成上表现不一致。有趣的是,部分开源模型在复杂教育推理任务上超越闭源系统。少样本提示效果在不同认知维度间差异显著,表明教育目标应采用差异化策略。这些发现为定制化优化中文教育场景下的LLMs提供了靶向评估指标。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate significant potential for educational applications. However, their unscrutinized deployment poses risks to educational standards, underscoring the need for rigorous evaluation. We introduce EduEval, a comprehensive hierarchical benchmark for evaluating LLMs in Chinese K-12 education. This benchmark makes three key contributions: (1) Cognitive Framework: We propose the EduAbility Taxonomy, which unifies Bloom's Taxonomy and Webb's Depth of Knowledge to organize tasks across six cognitive dimensions including Memorization, Understanding, Application, Reasoning, Creativity, and Ethics. (2) Authenticity: Our benchmark integrates real exam questions, classroom conversation, student essays, and expert-designed prompts to reflect genuine educational challenges; (3) Scale: EduEval comprises 24 distinct task types with over 11,000 questions spanning primary to high school levels. We evaluate 14 leading LLMs under both zero-shot and few-shot settings, revealing that while models perform well on factual tasks, they struggle with classroom dialogue classification and exhibit inconsistent results in creative content generation. Interestingly, several open source models outperform proprietary systems on complex educational reasoning. Few-shot prompting shows varying effectiveness across cognitive dimensions, suggesting that different educational objectives require tailored approaches. These findings provide targeted benchmarking metrics for developing LLMs specifically optimized for diverse Chinese educational tasks.

教育AI认知评估大模型评测中文教育

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。