40亿参数模型实现长周期探索,性能超越更大模型。
AgentCPM-Explore: Realizing Long-Horizon Deep Exploration for Edge-Scale Agents
- 通过参数空间融合与奖励去噪提升小模型推理稳定性。
- 在4个基准上达40亿级模型最优,5个任务超80亿模型。
- 适合资源受限场景下需要强探索能力的智能体应用。
尽管基于大语言模型的智能体在解决复杂任务方面展现出巨大潜力,但现有系统仍严重依赖大规模模型,导致边缘端规模模型的能力未被充分挖掘。本文首次系统研究了40亿参数级别智能体模型的训练方法。我们识别出三大瓶颈:监督微调中的灾难性遗忘、强化学习中对奖励信号噪声的敏感性,以及长上下文场景下的冗余信息导致的推理退化。为此,我们提出AgentCPM-Explore——一个具有高知识密度和强探索能力的紧凑型40亿参数智能体模型。其采用整体式训练框架,包含参数空间模型融合、奖励信号去噪和上下文信息精炼。通过深度探索,AgentCPM-Explore在40亿级模型中达到当前最优性能,在4个基准上匹配或超越80亿级先进模型,甚至在5个基准上优于Claude-4.5-Sonnet和DeepSeek-v3.2等更大模型。值得注意的是,其在GAIA文本任务上以pass@64获得97.09%准确率。结果表明,边缘模型的瓶颈并非能力上限,而是推理稳定性。基于此稳定训练框架,AgentCPM-Explore有效释放了边缘模型长期被低估的潜力。
原文摘要 · Abstract (English)
While Large Language Model (LLM)-based agents have shown remarkable potential for solving complex tasks, existing systems remain heavily reliant on large-scale models, leaving the capabilities of edge-scale models largely underexplored. In this paper, we present the first systematic study on training agentic models at the 4B-parameter scale. We identify three primary bottlenecks hindering the performance of edge-scale models: catastrophic forgetting during Supervised Fine-Tuning (SFT), sensitivity to reward signal noise during Reinforcement Learning (RL), and reasoning degradation caused by redundant information in long-context scenarios. To address the issues, we propose AgentCPM-Explore, a compact 4B agent model with high knowledge density and strong exploration capability. We introduce a holistic training framework featuring parameter-space model fusion, reward signal denoising, and contextual information refinement. Through deep exploration, AgentCPM-Explore achieves state-of-the-art (SOTA) performance among 4B-class models, matches or surpasses 8B-class SOTA models on four benchmarks, and even outperforms larger-scale models such as Claude-4.5-Sonnet or DeepSeek-v3.2 in five benchmarks. Notably, AgentCPM-Explore achieves 97.09% accuracy on GAIA text-based tasks under pass@64. These results provide compelling evidence that the bottleneck for edge-scale models is not their inherent capability ceiling, but rather their inference stability. Based on our well-established training framework, AgentCPM-Explore effectively unlocks the significant, yet previously underestimated, potential of edge-scale models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。