11B激活参数的稀疏模型,实现高效智能体推理与自提升。
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
- 采用稀疏MoE架构,仅激活110亿参数实现高效推理。
- 在数学、编码等任务上表现媲美顶级模型,如IMO达85.4%。
- 支持多轮交互优化,适合工业级智能体部署场景。
我们提出Step 3.5 Flash,一种稀疏的专家混合(MoE)模型,旨在连接前沿智能体能力与计算效率。该模型以1960亿参数为基础,仅激活110亿参数进行推理,通过交错的3:1滑动窗口/全注意力机制和多标记预测(MTP-3)降低多轮智能体交互的延迟与成本。为实现前沿智能,设计可扩展的强化学习框架,结合可验证信号与偏好反馈,在大规模离策略训练下保持稳定,实现数学、代码与工具使用任务的持续自提升。在多项评测中表现优异:IMO-AnswerBench达85.4%,LiveCodeBench-v6(2024.08–2025.05)达86.4%,tau2-Bench达88.2%,BrowseComp(含上下文管理)达69.0%,Terminal-Bench 2.0达51.0%,性能接近GPT-5.2 xHigh与Gemini 3.0 Pro。通过重新定义效率边界,为真实工业环境中的复杂智能体部署提供高密度基础。
原文摘要 · Abstract (English)
We introduce Step 3.5 Flash, a sparse Mixture-of-Experts (MoE) model that bridges frontier-level agentic intelligence and computational efficiency. We focus on what matters most when building agents: sharp reasoning and fast, reliable execution. Step 3.5 Flash pairs a 196B-parameter foundation with 11B active parameters for efficient inference. It is optimized with interleaved 3:1 sliding-window/full attention and Multi-Token Prediction (MTP-3) to reduce the latency and cost of multi-round agentic interactions. To reach frontier-level intelligence, we design a scalable reinforcement learning framework that combines verifiable signals with preference feedback, while remaining stable under large-scale off-policy training, enabling consistent self-improvement across mathematics, code, and tool use. Step 3.5 Flash demonstrates strong performance across agent, coding, and math tasks, achieving 85.4% on IMO-AnswerBench, 86.4% on LiveCodeBench-v6 (2024.08-2025.05), 88.2% on tau2-Bench, 69.0% on BrowseComp (with context management), and 51.0% on Terminal-Bench 2.0, comparable to frontier models such as GPT-5.2 xHigh and Gemini 3.0 Pro. By redefining the efficiency frontier, Step 3.5 Flash provides a high-density foundation for deploying sophisticated agents in real-world industrial environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。