用纯文本强化学习训练出更强推理模型,保留原模型能力并提升多模态理解。
Magistral
- 从零构建强化学习流水线,仅依赖自研模型与基础设施。
- 纯文本强化学习保持并提升原模型的指令遵循与函数调用能力。
- 开源小规模模型Magistral Small,适合研究与部署推理任务。
我们提出Magistral,Mistral首个推理模型及自研可扩展强化学习(RL)流水线。不同于依赖已有实现或先前模型的蒸馏轨迹,我们采用从头构建的方法,仅使用自身模型与基础设施。显著成果包括:探索纯强化学习训练大语言模型的极限、提出一种强制模型生成推理语言的简单方法,并验证仅在文本数据上进行强化学习可保留初始检查点的大部分能力。实验表明,该方法在多模态理解、指令遵循和函数调用方面表现维持或优于原始模型。本文推出基于Mistral Medium 3通过纯强化学习训练的Magistral Medium,并开源Magistral Small(Apache 2.0),其包含来自Magistral Medium的冷启动数据。
原文摘要 · Abstract (English)
We introduce Magistral, Mistral's first reasoning model and our own scalable reinforcement learning (RL) pipeline. Instead of relying on existing implementations and RL traces distilled from prior models, we follow a ground up approach, relying solely on our own models and infrastructure. Notably, we demonstrate a stack that enabled us to explore the limits of pure RL training of LLMs, present a simple method to force the reasoning language of the model, and show that RL on text data alone maintains most of the initial checkpoint's capabilities. We find that RL on text maintains or improves multimodal understanding, instruction following and function calling. We present Magistral Medium, trained for reasoning on top of Mistral Medium 3 with RL alone, and we open-source Magistral Small (Apache 2.0) which further includes cold-start data from Magistral Medium.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。