用智能代理合成高质量推理数据,让小模型也能达到顶尖水平。
Agentic Proposing: Enhancing Large Language Model Reasoning via Compositional Skill Synthesis
- 设计智能代理动态组合推理技能,自动生成结构严谨的题目。
- 仅用1.1万条合成数据,30B模型在AIME25上达91.6%准确率。
- 适合想低成本训练强推理模型的研究者和开发者。
提升大语言模型复杂推理能力依赖高质量、可验证的数据集,但人工标注成本高昂且难以扩展。现有合成方法常面临权衡:保持结构有效性会限制问题难度,放宽约束又易生成不一致或无解实例。为此,我们提出Agentic Proposing框架,将问题合成建模为目标驱动的序列决策过程,由专用代理动态选择并组合模块化推理技能。通过内部反思与工具使用迭代优化,我们基于多粒度策略优化(MGPO)训练出Agentic-Proposer-4B,生成数学、编程与科学领域的高精度、可验证训练轨迹。实证结果表明,基于该合成数据训练的下游求解器显著优于主流基线,并具备强跨领域泛化能力。特别地,仅使用11,000条合成轨迹训练的30B模型,在AIME25上达到91.6%准确率,媲美前沿的闭源模型如GPT-5,证明少量高质量合成信号可有效替代海量人工标注数据。
原文摘要 · Abstract (English)
Advancing complex reasoning in large language models relies on high-quality, verifiable datasets, yet human annotation remains cost-prohibitive and difficult to scale. Current synthesis paradigms often face a recurring trade-off: maintaining structural validity typically restricts problem complexity, while relaxing constraints to increase difficulty frequently leads to inconsistent or unsolvable instances. To address this, we propose Agentic Proposing, a framework that models problem synthesis as a goal-driven sequential decision process where a specialized agent dynamically selects and composes modular reasoning skills. Through an iterative workflow of internal reflection and tool-use, we develop the Agentic-Proposer-4B using Multi-Granularity Policy Optimization (MGPO) to generate high-precision, verifiable training trajectories across mathematics, coding, and science. Empirical results demonstrate that downstream solvers trained on agent-synthesized data significantly outperform leading baselines and exhibit robust cross-domain generalization. Notably, a 30B solver trained on only 11,000 synthesized trajectories achieves a state-of-the-art 91.6% accuracy on AIME25, rivaling frontier-scale proprietary models such as GPT-5 and proving that a small volume of high-quality synthetic signals can effectively substitute for massive human-curated datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。