35B参数模型通过优化实现百亿级性能,推理成本更低。
Mach-Mind-4-Flash Technical Report

- 采用多专家路由与动态教师调度,训练提速17%
- 混合奖励下避免性能退化,推理链压缩46%且精度损失小于0.7%
- 适合追求高性价比智能体的开发者与研究者
我们提出Mach-Mind-4-Flash,一个350亿参数的专家混合(MoE)智能体模型,激活参数仅30亿。仅通过后训练优化,无需扩大预训练计算量,其性能达到或超越1000亿参数级别模型。通过引入可扩展的智能体交互环境用于大规模强化学习,模型在真实应用任务中获得显著提升。整个流程包含三个阶段:(1) 统一的强化学习/操作策略训练框架,结合动态多教师调度与操作层级加速,实现端到端训练速度提升17%;(2) 在推理、通用和智能体三类任务轨道上并行训练多个领域专用强化学习专家,再通过多教师在线策略蒸馏(MOPD)融合为单一通用模型——一种基于路由反KL目标的方法,有效消除混合奖励强化学习中的性能波动;(3) 混合中位长度策略优化(HMPO),单阶段高效压缩推理链,压缩率19%–46%,精度损失不超过0.7个百分点。该模型在AIME'26上得分为92.70,在IFBench上为82.82,在Behavioral-SafetyBench上为80.74,在BFCL-v4上为75.80,在BrowseComp-zh上为72.31,在ClawBench上为84.20,表现领先或持平于激活参数规模大10–30倍的模型,同时推理成本仅为极小部分。
原文摘要 · Abstract (English)
We present Mach-Mind-4-Flash, a 35B-parameter Mixture-of-Experts (MoE) agentic model with 3B activated parameters. Through post-training optimization alone without scaling pre-training compute, the model achieves performance on par with or surpassing that of 100B-parameter-class models. By introducing scalable agentic interaction environments for large-scale reinforcement learning, the model attains significant performance gains on real-world application tasks. Our pipeline comprises three stages: (1) a unified RL/OPD training infrastructure with dynamic multi-teacher scheduling and operator-level acceleration, delivering 17\% end-to-end training speedup; (2) multiple domain-specific RL experts trained in parallel across Reasoning, General, and Agent tracks, then fused into a single generalist via Multi-Teacher On-Policy Distillation (MOPD) -- a routed reverse-KL objective that eliminates the see-saw degradation of mixed-reward RL; (3) Hybrid Median-length Policy Optimization (HMPO), a single-stage token-efficiency method that compresses reasoning chains by 19--46\% with $\le$0.7 percentage-point accuracy loss. Mach-Mind-4-Flash scores 92.70 on AIME'26, 82.82 on IFBench, 80.74 on Behavioral-SafetyBench, 75.80 on BFCL-v4, 72.31 on BrowseComp-zh, and 84.20 on ClawBench -- leading or matching models with 10--30$\times$ its activated size at a fraction of the inference cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。