arXiv:2607.22083cs.AIcs.CL2026-07被引 1

30亿参数小模型实现强智能代理能力,性能超越更大模型。

Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model

  • 用循环变压器结构在不增加参数下提升容量,从头预训练28万亿词。
  • 在代码、办公、工具调用等任务上超越Qwen3.5-9B和Gemma4-12B。
  • 适合本地部署的个人智能助手,支持长流程任务与高效推理。

我们提出 Nanbeige4.2-3B,一个仅含30亿非嵌入参数的紧凑通用智能体模型。该模型在代码代理、办公代理及复杂工具使用任务中表现强劲,同时在数学、编程和科学推理方面保持高竞争力。模型在28万亿个标记上从头预训练,采用循环变压器(Looped Transformer)复用层堆栈以提升容量而不增加参数。为增强SFT数据和轨迹构建,通过真实部署与大规模合成扩展可执行环境、任务资源和智能体框架多样性。强化学习流水线结合混合模式的思维与非思维响应的RLHF,提升整体质量并减少失败案例;通过长度可控的推理强化学习平衡准确率与推理效率;采用基于结果与过程奖励的智能体强化学习,稳定长周期训练。大量评估显示,Nanbeige4.2-3B在多种智能体基准测试中优于更大模型,包括Qwen3.5-9B和Gemma4-12B,同时在推理与对齐任务上仍具竞争力。OpenClaw性能进一步验证其作为小型本地个人助手的适用性。

原文摘要 · Abstract (English)

We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabilities in mathematics, coding, and science. Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters. For SFT data and trajectory construction, we expand the diversity of executable environments, task assets, and agentic scaffolds through real-world deployment and large-scale synthesis. Our RL pipeline applies mixed-mode RLHF over Think and Non-Think responses to improve overall model quality and reduce failure cases, length-controlled reasoning RL to balance accuracy and reasoning efficiency, and agentic RL with outcome and process rewards to stabilize long-horizon training. Extensive evaluations show that Nanbeige4.2-3B outperforms larger models, including Qwen3.5-9B and Gemma4-12B, across diverse agentic benchmarks while remaining competitive on reasoning and alignment tasks. Performance with OpenClaw further supports its use as a compact local personal assistant.

智能代理小模型强化学习本地部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。