arXiv:2511.05951cs.AI2025-11被引 5

开源打造高效智能体,80亿参数模型性能媲美更大模型

Klear-AgentForge: Forging Agentic Intelligence through Posttraining Scaling

  • 用合成数据微调加多轮强化学习,提升模型工具使用与编程能力
  • 在多个智能体评测中表现领先,80亿参数模型超越同类大模型
  • 全流程开源,适合研究者复现和开发自主智能体系统

尽管强大的智能体模型不断涌现,但关键的训练细节缺失阻碍了开源社区发展高性能模型。本研究提出一个完整且完全开源的训练流程,基于 Qwen3-8B 基座模型构建高精度智能体模型 Klear-Qwen3-AgentForge。通过合成数据监督微调(SFT)结合多轮强化学习(RL),充分释放模型在多种智能体任务中的潜力。我们在工具使用与编程领域的多个智能体基准上进行专项实验。Klear-Qwen3-AgentForge-8B 在同规模语言模型中达到顶尖水平,且在性能上仍可与显著更大的模型竞争。

原文摘要 · Abstract (English)

Despite the proliferation of powerful agentic models, the lack of critical post-training details hinders the development of strong counterparts in the open-source community. In this study, we present a comprehensive and fully open-source pipeline for training a high-performance agentic model for interacting with external tools and environments, named Klear-Qwen3-AgentForge, starting from the Qwen3-8B base model. We design effective supervised fine-tuning (SFT) with synthetic data followed by multi-turn reinforcement learning (RL) to unlock the potential for multiple diverse agentic tasks. We perform exclusive experiments on various agentic benchmarks in both tool use and coding domains. Klear-Qwen3-AgentForge-8B achieves state-of-the-art performance among LLMs of similar size and remains competitive with significantly larger models.

智能体强化学习开源模型Qwen

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。