arXiv:2409.12739cs.CL2024-09被引 7

首个中文教育价值观评估基准,测试大模型的教育理念与文化理解。

Edu-Values: Towards Evaluating the Chinese Education Values of Large Language Models

  • 构建涵盖7大核心价值的1418道题评测体系,含多模态与传统文化题。
  • 中文大模型在教育价值观上优于英文模型,通义千问2得分81.37最高。
  • 发现模型在师德与教育哲学上表现薄弱,用其建知识库可提升对齐效果。

本文提出Edu-Values,首个针对中文教育价值观的评估基准,涵盖专业理念、教师职业道德、教育法律法规、文化素养、教育知识与技能、基础能力及学科知识共七项核心价值。精心设计1418道题目,包含选择题、多模态问答、主观分析、对抗性提示及中华传统文化(简答)题型。基于人工反馈对21个顶尖大语言模型进行自动评估,得出三大发现:(1) 由于教育文化差异,中文大模型整体优于英文模型,其中通义千问2以81.37分排名第一;(2) 模型在教师职业道德与专业理念方面普遍表现不佳;(3) 利用Edu-Values构建外部知识库用于RAG,显著提升模型对齐效果。该基准有效验证了其应用价值。

原文摘要 · Abstract (English)

In this paper, we present Edu-Values, the first Chinese education values evaluation benchmark that includes seven core values: professional philosophy, teachers' professional ethics, education laws and regulations, cultural literacy, educational knowledge and skills, basic competencies and subject knowledge. We meticulously design 1,418 questions, covering multiple-choice, multi-modal question answering, subjective analysis, adversarial prompts, and Chinese traditional culture (short answer) questions. We conduct human feedback based automatic evaluation over 21 state-of-the-art (SoTA) LLMs, and highlight three main findings: (1) due to differences in educational culture, Chinese LLMs outperform English LLMs, with Qwen 2 ranking the first with a score of 81.37; (2) LLMs often struggle with teachers' professional ethics and professional philosophy; (3) leveraging Edu-Values to build an external knowledge repository for RAG significantly improves LLMs' alignment. This demonstrates the effectiveness of the proposed benchmark.

教育评估大模型评测价值观对齐中文LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。