用双向动态生成题目,让大模型更省数据地学会解数学题。
Bidirectional Curriculum Generation: A Multi-Agent Framework for Data-Efficient Mathematical Reasoning
- 构建多智能体系统,根据模型错误自动调难或降难题目。
- 仅用少量样本就达到比基线更好的数学推理性能。
- 适合想高效训练数学能力的AI研究者和开发者。
提升大语言模型的数学推理能力通常需要海量数据,但数据效率仍是关键瓶颈。传统单向课程学习(由易到难)会盲目提升难度,即使基础漏洞未补全,导致大量计算浪费在无法解决的问题上。为最大化每条训练样本的教学价值,我们提出一种新型双向课程生成框架。该多智能体系统模拟自适应教学,建立闭环反馈机制:不仅能通过复杂化问题挑战模型,更关键的是能简化问题以修复特定推理错误。这一机制确保模型始终只接收当前阶段最有效的数据。基于最优教学节奏定理,该方法优化了学习路径,在显著少于基线所需指令样本的情况下,实现了更优的推理表现。
原文摘要 · Abstract (English)
Enhancing mathematical reasoning in Large Language Models typically demands massive datasets, yet data efficiency remains a critical bottleneck. While Curriculum Learning attempts to structure this process, standard unidirectional approaches (simple-to-complex) suffer from inefficient sample utilization: they blindly escalate complexity even when foundational gaps persist, leading to wasted computation on unsolvable problems. To maximize the instructional value of every training sample, we introduce a novel Bidirectional Curriculum Generation framework. Unlike rigid trajectories, our multi-agent ecosystem mimics adaptive pedagogy to establish a closed feedback loop. It dynamically generates data by either complicating problems to challenge the model or, crucially, simplying them to repair specific reasoning failures. This mechanism ensures that the model consumes only the most effective data at any given stage. Grounded in the Optimal Pacing Theorem, our approach optimizes the learning trajectory, significantly outperforming baselines while achieving superior reasoning performance with substantially fewer instruction samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。