32B开源模型在数学与编程推理上达到顶尖水平。
AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale
- 基于Qwen2.5-32B微调,结合监督与强化学习提升推理能力。
- 在AIME 2024、AIME 2025和LiveCodeBench分别获85.3、74.4、70.3分。
- 适合关注可部署、可复现的中等规模模型的研究者与开发者。
我们提出AM-Thinking-v1,一个32B参数的稠密语言模型,推动了推理能力的边界,体现了开源协作的精神。该模型超越DeepSeek-R1,媲美顶级Mixture-of-Experts(MoE)模型如Qwen3-235B-A22B和Seed1.5-Thinking。在AIME 2024、AIME 2025和LiveCodeBench上分别取得85.3、74.4和70.3的高分,展现出同规模开源模型中的顶尖数学与编码能力。模型完全基于开源的Qwen2.5-32B基础模型及公开查询数据,通过精心设计的后训练流程——结合监督微调与强化学习——实现卓越推理性能。本工作表明,开源社区可在32B这一兼具性能与实用性的规模上实现高水平表现。我们希望以此激发更多协作,推动中等规模模型在推理能力上的突破,同时保持创新的可及性。模型已开源至Hugging Face。
原文摘要 · Abstract (English)
We present AM-Thinking-v1, a 32B dense language model that advances the frontier of reasoning, embodying the collaborative spirit of open-source innovation. Outperforming DeepSeek-R1 and rivaling leading Mixture-of-Experts (MoE) models like Qwen3-235B-A22B and Seed1.5-Thinking, AM-Thinking-v1 achieves impressive scores of 85.3 on AIME 2024, 74.4 on AIME 2025, and 70.3 on LiveCodeBench, showcasing state-of-the-art mathematical and coding capabilities among open-source models of similar scale. Built entirely from the open-source Qwen2.5-32B base model and publicly available queries, AM-Thinking-v1 leverages a meticulously crafted post-training pipeline - combining supervised fine-tuning and reinforcement learning - to deliver exceptional reasoning capabilities. This work demonstrates that the open-source community can achieve high performance at the 32B scale, a practical sweet spot for deployment and fine-tuning. By striking a balance between top-tier performance and real-world usability, we hope AM-Thinking-v1 inspires further collaborative efforts to harness mid-scale models, pushing reasoning boundaries while keeping accessibility at the core of innovation. We have open-sourced our model on \href{https://huggingface.co/a-m-team/AM-Thinking-v1}{Hugging Face}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。