30亿参数小模型实现顶尖可验证推理能力,媲美大模型。
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models

- 基于分阶段微调与强化学习,提升小模型推理能力
- 在AIME26达94.3分(优化后97.1),LiveCodeBench达80.2%
- 兼具强泛化性与指令可控性,适合高效推理场景
本文介绍VibeThinker-3B,一个仅30亿参数的紧凑密集模型,旨在探索小模型在可验证推理任务中的极限。基于Spectrum-to-Signal后训练范式,通过课程式监督微调、多领域强化学习和离线自蒸馏构建优化流程。实验表明,该模型在高挑战性任务上达到前沿水平:AIME26得分为94.3(使用命题级测试时缩放后达97.1),LiveCodeBench v6 Pass@1为80.2%,在近期未见的LeetCode竞赛中接受率达96.1%。性能接近或超越深求问V3.2、GLM-5、Gemini 3 Pro等大模型。IFEval得分为93.4,证实其指令控制严格性未受损。研究支持参数压缩-覆盖假说,认为可验证推理可压缩为紧凑核心,而开放域知识需广泛参数覆盖。这表明小模型不仅是部署替代品,更是通向高性能的重要路径。
原文摘要 · Abstract (English)
This technical report introduces VibeThinker-3B, a compact dense model with 3B parameters developed to investigate how far verifiable reasoning can be pushed within a strictly small-model regime. Building upon the Spectrum-to-Signal post-training paradigm, we systematically enhance the model through an optimized pipeline that includes curriculum-based supervised fine-tuning, multi-domain reinforcement learning, and offline self-distillation. Experimental evaluations demonstrate that VibeThinker-3B achieves frontier-level performance on highly demanding verifiable tasks. Specifically, it attains a score of 94.3 on AIME26 (improving to 97.1 with claim-level test-time scaling), an 80.2 Pass@1 on LiveCodeBench v6, and exhibits strong out-of-distribution generalization with a 96.1\% acceptance rate on recent unseen LeetCode contests. This effectively places it in the performance band of first-tier reasoning systems, matching or exceeding flagship models that are orders of magnitude larger, such as DeepSeek V3.2, GLM-5, and Gemini 3 Pro. Furthermore, a score of 93.4 on IFEval confirms that this extreme reasoning enhancement does not compromise strict instruction controllability. Extending our previous 1.5B work, these findings motivate the Parametric Compression-Coverage Hypothesis, which views verifiable reasoning as compressible into compact reasoning cores, while open-domain knowledge and general-purpose competence require broad parameter coverage over facts, concepts, and long-tail scenarios. This perspective suggests that compact models are not merely deployment-efficient substitutes, but a complementary path toward frontier-level performance in parameter-dense capability regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。