多智能体协作生成个性化数学题,兼顾多样性与目标对齐
EduAgentQG: Multi-Agent Personalized Mathematics Question Generation with Explicit Diversity and Objective-Aware Evaluation
- 设计闭环流程:规划-写作-评估-优化-检查,分步控制生成质量
- 在10273道题上验证,多样性与目标一致性均优于现有方法
- 适合教育科技研发者、个性化学习系统设计者使用
智能教育中,个性化数学题生成旨在满足教学要求并支持自适应测评与学习。现有基于大模型的单/多智能体方法虽提升生成灵活性,但仍依赖聚合反馈或模型随机性,难以同时保证维度级目标对齐与可控多样性。为此,我们提出EduAgentQG,一个具备显式多样性和目标感知评估的多智能体协同框架。该框架将题目生成组织为规划、写作、评估、优化、检查的闭环流程:结构化生成计划与多方向引导候选生成,细粒度评估验证逻辑正确性、可解性及知识概念、难度、年级水平、核心素养等维度的目标一致性。我们首次构建涵盖1-9年级共10,273道题的数学题生成基准,覆盖634个知识点、16项核心素养与三种难度等级;评估分为两个子集:MathChoice(489个选择题目标)和MathBlank(500个填空题目标)。实验表明,EduAgentQG在多样性、目标一致性与胜率上持续优于COT、COT$_N$、ReAct和EQPR。
原文摘要 · Abstract (English)
In intelligent education, personalized mathematics question generation aims to produce mathematics questions that satisfy educational requirements while supporting adaptive assessment and learning. Existing LLM-based single-agent and multi-agent methods improve generation flexibility, but they still tend to rely on aggregated feedback or model randomness, making it difficult to jointly ensure dimension-wise objective alignment and controllable diversity. To address these challenges, we propose EduAgentQG, a multi-agent collaborative framework for personalized mathematics question generation with explicit diversity and objective-aware evaluation. EduAgentQG organizes question generation as a closed-loop process of planning, writing, evaluation, refinement, and checking: structured generation plans and multiple generation directions guide candidate generation, while fine-grained evaluation verifies logical correctness, solvability, and objective alignment in knowledge concepts, difficulty, grade level, and core competencies. We first construct a mathematics question generation benchmark containing 10,273 questions across Grades 1-9, covering 634 knowledge concepts, 16 core competencies, and three difficulty levels; for evaluation, it is organized into two subsets: MathChoice, with 489 educational objectives for multiple-choice question generation, and MathBlank, with 500 educational objectives for fill-in-the-blank question generation. Experiments show that EduAgentQG consistently outperforms COT, COT$_N$, ReAct, and EQPR in diversity, Objective Consistency, and Win Rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。