用近端发展区理论生成智能体训练数据,提升大模型推理能力。
AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis
- 基于近端发展区理论生成模型能学但无法独立解决的任务数据。
- 训练出的300亿参数模型在人类终极考试等基准上领先现有水平。
- 适合研究大模型智能体、强化学习与自动数据构建的开发者。
在大语言模型能力边界任务上进行训练是激发高级推理能力的关键。本文提出一种受近端发展区(ZPD)教育理论启发的数据合成方法,将能力边界定义为模型无法独立完成但可通过指导掌握的任务。我们构建了AgentFrontier引擎,一个自动化数据生成管道,可生成高质量、多学科且精准位于模型ZPD内的数据,支持知识密集型持续预训练和复杂推理任务的定向后训练。基于该框架,我们设计了动态自动评估基准ZPD Exam,用于测评模型在前沿任务上的表现。使用合成数据训练的AgentFrontier-30B-A3B模型在Humanity's Last Exam等严苛基准上达到当前最优性能,甚至超越部分领先闭源智能体。本工作表明,基于ZPD引导的数据合成是一种可扩展、高效构建更强大大模型智能体的路径。
原文摘要 · Abstract (English)
Training large language model agents on tasks at the frontier of their capabilities is key to unlocking advanced reasoning. We introduce a data synthesis approach inspired by the educational theory of the Zone of Proximal Development (ZPD), which defines this frontier as tasks an LLM cannot solve alone but can master with guidance. To operationalize this, we present the AgentFrontier Engine, an automated pipeline that synthesizes high-quality, multidisciplinary data situated precisely within the LLM's ZPD. This engine supports both continued pre-training with knowledge-intensive data and targeted post-training on complex reasoning tasks. From the same framework, we derive the ZPD Exam, a dynamic and automated benchmark designed to evaluate agent capabilities on these frontier tasks. We train AgentFrontier-30B-A3B model on our synthesized data, which achieves state-of-the-art results on demanding benchmarks like Humanity's Last Exam, even surpassing some leading proprietary agents. Our work demonstrates that a ZPD-guided approach to data synthesis offers a scalable and effective path toward building more capable LLM agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。