arXiv:2605.26494cs.AIcs.CL2026-05被引 35

迷你激活实现最大真实智能,模型仅用98亿参数每词激活却达顶尖性能。

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

论文配图:The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
图 1 · 摘自论文原文
  • 采用极小激活机制,每词仅激活98亿参数,大幅降低计算开销。
  • 在编程、搜索、办公任务等基准上达到前沿水平,效率与性能兼备。
  • 支持自主调试与自我演进,适合需要高效智能代理的场景。

我们推出MiniMax-M2系列,一类基于‘极小激活释放最大真实智能’理念的专家混合语言模型。旗舰模型M2总参数量为2299亿,每词仅激活98亿参数。该系列从零开始设计用于智能体部署,包含三大组件:(i) 由智能体驱动的数据流水线,生成大规模可验证轨迹,涵盖智能体编程与协作任务,基于可执行工作区与对齐奖励;(ii) Forge,一个可扩展的原生智能体强化学习系统,适配长周期智能体轨迹,结合窗口化FIFO调度、前缀树合并、推理优化及训练-推理-智能体解耦设计,支持白盒与黑盒智能体;(iii) 最新M2.7检查点已初步实现自演化——能自主调试训练过程并修改自身架构。从M2到M2.7,该组合将极小激活优势转化为在智能体编程、深度搜索、办公任务和推理基准上的顶级表现。

原文摘要 · Abstract (English)

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale, verifiable trajectories across agentic coding and agentic cowork, each grounded in an executable workspace and an artifact-aligned reward; (ii) Forge, a scalable agent-native RL system that adapts to long-horizon agent trajectories, paired with windowed-FIFO scheduling, prefix-tree merging, inference optimization, and a clean training-inference-agent decoupling that supports both white-box and black-box agents; (iii) the latest M2.7 checkpoint takes an early step toward self-evolution -- autonomously debugging training runs and modifying its own scaffold. Across M2 through M2.7, this combination translates a mini-activation footprint into frontier-tier performance on agentic coding, deep search, office-task, and reasoning benchmarks.

语言模型智能体极小激活自演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。