150亿参数模型通过训练设计实现顶尖多模态推理,无需强化学习或超大规模算力。
Apriel-1.5-15b-Thinker
- 分三阶段渐进训练:扩深推理能力、合成数据提升视觉理解、高质量指令微调
- 在人工分析智能指数上得52分,媲美大模型但资源消耗少得多
- 适合算力有限的研究者,开源全部训练代码与模型供复用
我们提出 Apriel-1.5-15B-Thinker,一个拥有150亿参数的开源多模态推理模型,其前沿性能源于训练设计而非单纯规模。基于 Pixtral-12B,采用三阶段渐进方法:(1)深度扩展以增强推理能力,无需从头预训练;(2)分阶段持续预训练,先建立基础文本与视觉理解,再通过针对性合成数据生成提升空间结构、组合理解与细粒度感知能力;(3)在精选指令-响应对上进行高质量纯文本监督微调,涵盖数学、编程、科学及工具使用等领域的显式推理链。值得注意的是,模型未使用强化学习或偏好优化,凸显了数据驱动持续预训练的价值。在 Artificial Analysis Intelligence Index 上得分52,与 DeepSeek-R1-0528 相当,但计算资源需求显著更低。在十项图像基准上,平均性能距 Gemini-2.5-Flash 与 Claude Sonnet-3.7 仅差五分,是单卡部署条件下关键突破。结果表明,精心设计的中期训练可弥补显著能力差距,使前沿多模态推理对资源受限组织更易获得。模型检查点、训练配方与评估协议已按 MIT 许可开源,以推动开放研究。
原文摘要 · Abstract (English)
We present Apriel-1.5-15B-Thinker, a 15-billion parameter open-weights multimodal reasoning model that achieves frontier-level performance through training design rather than sheer scale. Starting from Pixtral-12B, we apply a progressive three-stage methodology: (1) depth upscaling to expand reasoning capacity without pretraining from scratch, (2) staged continual pre-training that first develops foundational text and vision understanding, then enhances visual reasoning through targeted synthetic data generation addressing spatial structure, compositional understanding, and fine-grained perception, and (3) high-quality text-only supervised fine-tuning on curated instruction-response pairs with explicit reasoning traces spanning mathematics, coding, science, and tool use. Notably, our model achieves competitive results without reinforcement learning or preference optimization, isolating the contribution of our data-centric continual pre-training approach. On the Artificial Analysis Intelligence Index, Apriel-1.5-15B-Thinker attains a score of 52, matching DeepSeek-R1-0528 despite requiring significantly fewer computational resources. Across ten image benchmarks, its performance is on average within five points of Gemini-2.5-Flash and Claude Sonnet-3.7, a key achievement for a model operating within single-GPU deployment constraints. Our results demonstrate that thoughtful mid-training 2 design can close substantial capability gaps without massive scale, making frontier-level multimodal reasoning accessible to organizations with limited infrastructure. We release the model checkpoint, all training recipes, and evaluation protocols under the MIT license to to advance open-source research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。