arXiv:2508.10005cs.CLcs.AI2025-08被引 5

构建中文教育题生成评估基准,检验大模型出题能力。

From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation

  • 设计五维评估框架,覆盖知识点、难度、题型等维度。
  • 基于900个样本测试46个主流模型,发现出题能力普遍不足。
  • 适用于教育AI研究者与智能教学系统开发者。

大型语言模型在数学解题方面表现出色,但将答案转化为高质量教育题目仍面临挑战且研究不足。为推进教育题生成(EQG)并评估大模型生成具有教学价值与教育效果题目的能力,我们提出EQGBench,一个专用于中文教育题生成的综合性评估基准。该基准建立在涵盖数学、物理、化学三门初中基础学科的900个评估样本数据集上,包含不同知识点、难度梯度和题型要求的用户查询,模拟真实教学场景。对46个主流大模型进行系统评估后发现,当前模型在生成体现教育价值、促进学生综合能力发展的题目方面仍有显著提升空间。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities in mathematical problem-solving. However, the transition from providing answers to generating high-quality educational questions presents significant challenges that remain underexplored. To advance Educational Question Generation (EQG) and facilitate LLMs in generating pedagogically valuable and educationally effective questions, we introduce EQGBench, a comprehensive benchmark specifically designed for evaluating LLMs' performance in Chinese EQG. EQGBench establishes a five-dimensional evaluation framework supported by a dataset of 900 evaluation samples spanning three fundamental middle school disciplines: mathematics, physics, and chemistry. The dataset incorporates user queries with varying knowledge points, difficulty gradients, and question type specifications to simulate realistic educational scenarios. Through systematic evaluation of 46 mainstream large models, we reveal significant room for development in generating questions that reflect educational value and foster students' comprehensive abilities.

教育AI题生成大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。