15亿参数小模型通过多样性优化,实现媲美大模型的推理能力。
Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B
- 用两阶段多样性蒸馏+最大熵策略优化,从海量解中提炼正确逻辑
- 在数学和编程基准上超越400倍大的深求模型,训练成本仅7800美元
- 适合追求低成本高效率推理的开发者与研究者
挑战小模型无法具备强推理能力的固有认知,本文提出基于谱-信号原理(SSP)的1.5B参数稠密模型VibeThinker-1.5B。该方法先通过两阶段多样性探索蒸馏(SFT)生成多样化解,再以最大熵引导策略优化(RL)强化正确信号。总训练成本仅7800美元,其推理能力优于闭源模型Magistral Medium和Claude Opus 4,与开源模型GPT OSS-20B Medium相当。尤其在三个数学基准上显著超越400倍大的DeepSeek R1:AIME24(80.3 vs. 79.8)、AIME25(74.4 vs. 70.0)、HMMT25(50.4 vs. 41.7),相较其基线模型提升显著(6.7→80.3、4.3→74.4、0.6→50.4)。在LiveCodeBench V6上得分为51.1,优于Magistral Medium的50.3和基线的0.0。结果表明,小模型可通过优化机制实现类大模型推理能力,大幅降低训练与推理成本,推动先进AI研究普惠化。
原文摘要 · Abstract (English)
Challenging the prevailing consensus that small models inherently lack robust reasoning, this report introduces VibeThinker-1.5B, a 1.5B-parameter dense model developed via our Spectrum-to-Signal Principle (SSP). This challenges the prevailing approach of scaling model parameters to enhance capabilities, as seen in models like DeepSeek R1 (671B) and Kimi k2 (>1T). The SSP framework first employs a Two-Stage Diversity-Exploring Distillation (SFT) to generate a broad spectrum of solutions, followed by MaxEnt-Guided Policy Optimization (RL) to amplify the correct signal. With a total training cost of only $7,800, VibeThinker-1.5B demonstrates superior reasoning capabilities compared to closed-source models like Magistral Medium and Claude Opus 4, and performs on par with open-source models like GPT OSS-20B Medium. Remarkably, it surpasses the 400x larger DeepSeek R1 on three math benchmarks: AIME24 (80.3 vs. 79.8), AIME25 (74.4 vs. 70.0), and HMMT25 (50.4 vs. 41.7). This is a substantial improvement over its base model (6.7, 4.3, and 0.6, respectively). On LiveCodeBench V6, it scores 51.1, outperforming Magistral Medium's 50.3 and its base model's 0.0. These findings demonstrate that small models can achieve reasoning capabilities comparable to large models, drastically reducing training and inference costs and thereby democratizing advanced AI research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。