32B参数模型通过高效训练与推理,达到顶尖数学推理水平。
K2-Think: A Parameter-Efficient Reasoning System
- 结合长思维链微调与测试时扩展技术提升小模型性能
- 在数学推理上超越更大模型,推理速度超2000词/秒
- 开源免费,适合追求高效推理的开发者与研究者
K2-Think 是一个基于 Qwen2.5 的推理系统,仅用 32B 参数即达到顶尖性能,表现媲美甚至超越 GPT-OSS 120B 和 DeepSeek v3.1 等大型模型。该系统通过六大核心技术:长思维链监督微调、可验证奖励的强化学习(RLVR)、代理式规划、测试时扩展、推测解码以及推理优化硬件,结合公开开源数据集实现。在数学推理方面,其在开源模型中取得领先成绩,同时在代码与科学任务中也表现优异。结果表明,通过集成后训练策略与推理时增强,小模型也能实现高性能,显著提升开源推理系统的可及性与性价比。K2-Think 已开源,可通过 Cerebras Wafer-Scale Engine 实现每请求超过 2,000 词/秒的推理速度。
原文摘要 · Abstract (English)
K2-Think is a reasoning system that achieves state-of-the-art performance with a 32B parameter model, matching or surpassing much larger models like GPT-OSS 120B and DeepSeek v3.1. Built on the Qwen2.5 base model, our system shows that smaller models can compete at the highest levels by combining advanced post-training and test-time computation techniques. The approach is based on six key technical pillars: Long Chain-of-thought Supervised Finetuning, Reinforcement Learning with Verifiable Rewards (RLVR), Agentic planning prior to reasoning, Test-time Scaling, Speculative Decoding, and Inference-optimized Hardware, all using publicly available open-source datasets. K2-Think excels in mathematical reasoning, achieving state-of-the-art scores on public benchmarks for open-source models, while also performing strongly in other areas such as Code and Science. Our results confirm that a more parameter-efficient model like K2-Think 32B can compete with state-of-the-art systems through an integrated post-training recipe that includes long chain-of-thought training and strategic inference-time enhancements, making open-source reasoning systems more accessible and affordable. K2-Think is freely available at k2think.ai, offering best-in-class inference speeds of over 2,000 tokens per second per request via the Cerebras Wafer-Scale Engine.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。