30亿参数小模型实现推理、对齐与工具调用,性能媲美大模型。
Nanbeige4.1-3B: A Small General Model that Reasons, Aligns, and Acts
- 融合点对点与成对奖励,提升推理与人类偏好对齐
- 复杂度感知强化学习,优化代码生成的正确性与效率
- 支持600步工具调用,适合复杂长程任务的智能体应用
我们提出Nanbeige4.1-3B,一个仅含30亿参数的统一通用语言模型,可同时实现强大的代理行为、代码生成与通用推理。据我们所知,它是首个在单一模型中达成此多能力平衡的开源小型语言模型(SLM)。为提升推理与偏好对齐,我们结合点对点与成对奖励建模,确保输出高质量且符合人类偏好。针对代码生成,设计复杂度感知奖励,在强化学习中优化正确性与效率。在深度搜索中,通过复杂数据合成并引入回合级监督训练,实现稳定长时序工具交互,使模型可靠执行最多600步工具调用以解决复杂问题。大量实验表明,Nanbeige4.1-3B显著优于同规模先前模型(如Nanbeige4-3B-2511和Qwen3-4B),甚至超越更大模型(如Qwen3-30B-A3B)。结果表明,小模型可同时具备广度能力与深度专长,重新定义了30亿参数模型的潜力。
原文摘要 · Abstract (English)
We present Nanbeige4.1-3B, a unified generalist language model that simultaneously achieves strong agentic behavior, code generation, and general reasoning with only 3B parameters. To the best of our knowledge, it is the first open-source small language model (SLM) to achieve such versatility in a single model. To improve reasoning and preference alignment, we combine point-wise and pair-wise reward modeling, ensuring high-quality, human-aligned responses. For code generation, we design complexity-aware rewards in Reinforcement Learning, optimizing both correctness and efficiency. In deep search, we perform complex data synthesis and incorporate turn-level supervision during training. This enables stable long-horizon tool interactions, allowing Nanbeige4.1-3B to reliably execute up to 600 tool-call turns for complex problem-solving. Extensive experimental results show that Nanbeige4.1-3B significantly outperforms prior models of similar scale, such as Nanbeige4-3B-2511 and Qwen3-4B, even achieving superior performance compared to much larger models, such as Qwen3-30B-A3B. Our results demonstrate that small models can achieve both broad competence and strong specialization simultaneously, redefining the potential of 3B parameter models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。