arXiv:2601.11792cs.AIcs.CL2026-01被引 2

提出自进化多角色协作框架,生成创新且高正确的数学题

A self-evolving multi-role collaborative framework with fine-grained difficulty guidance for innovative mathematical problem generation

  • 多角色协同迭代优化,结合自评与反馈提升题目质量
  • 引入细粒度难度模型,生成题目创新性显著提升
  • 适合教育AI研究者与智能出题系统开发者

数学问题生成(MPG)是智能教育的重要方向。尽管大语言模型(LLMs)已实现高正确率,但缺乏创新性且区分度差。本文提出创新数学问题生成(IMPG)任务,构建自进化、多角色协作框架,包含采样器、生成器、评估器、状态机与记忆模块,通过迭代优化确保题目正确性。提出改进的难度模型实现细粒度引导,并采用数据驱动的关联引导路径采样(DAPS)算法增强语义合理性。构建高质量的高中数学题数据集HSM3K-CN,采用持续预训练(CPT)、监督微调(SFT)和组相对策略优化(GRPO)的多阶段训练流程,提升模型生成与评估能力。通过知识蒸馏将专家模型评估能力迁移至学生模型,实现系统自进化。实验表明,相比基线模型,该方法在保持高正确率的同时,显著提升生成题目的创新性。

原文摘要 · Abstract (English)

Mathematical problem generation (MPG) is a significant research direction in the field of intelligent education. In recent years, the rapid development of large language models (LLMs) has enabled new technological approaches to problem-generation tasks. Although existing LLMs can achieve high correctness rates, they generally lack innovation and exhibit poor discrimination. In this paper, we propose the task of innovative math problem generation (IMPG). To solve the IMPG task, this paper proposes a self-evolving, multi-role collaborative framework with fine-grained difficulty guidance. First, a multi-role collaborative mechanism comprising a sampler, generator, evaluator, state machine, and memory is constructed, ensuring the correctness of generated problems through iterative optimization informed by self-assessment and external feedback. Second, we introduce an improved difficulty model to quantify difficulty and provide fine-grained guidance. We adopt the data-driven association-guided path sampling (DAPS) algorithm to enhance the semantic rationality of sampled encodings. Third, we construct the HSM3K-CN dataset, which comprises high-quality high school math problems. A multi-stage training pipeline is adopted, incorporating continual pre-training (CPT), supervised fine-tuning (SFT), and group relative policy optimization (GRPO), to enhance the generation and evaluation capabilities of the base model. Finally, system self-evolution is achieved by transferring evaluation capabilities from the expert model to the apprentice model via distillation. Experiments show that, compared to baseline models, our proposed method significantly improves the innovation of the generated problems while maintaining a high correctness rate.

数学生成多角色协作自进化难度控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。