5600亿参数开源模型,专攻复杂任务规划与工具使用。
LongCat-Flash-Thinking-2601 Technical Report
- 采用专家混合架构,融合多领域训练与环境协同设计。
- 在多种代理任务中表现领先,支持超10000个环境的稳定训练。
- 新增深度并行推理模式,提升复杂问题解决能力。
我们提出 LongCat-Flash-Thinking-2601,一个拥有5600亿参数的开源混合专家(MoE)推理模型,具备卓越的代理式推理能力。该模型在广泛代理基准测试中达到开源模型的最先进水平,涵盖代理搜索、工具调用及工具集成推理。除了基准性能外,模型在复杂工具交互中表现出强泛化能力,并在噪声丰富的现实环境中保持稳健行为。其优势源于统一的训练框架,结合领域并行专家训练与后续融合,以及从预训练到后训练全链路的数据构建、环境设计、算法与基础设施协同优化。特别是,通过深入探索环境规模扩展与严谨的任务构造,显著提升了复杂工具使用下的泛化能力。为优化长尾分布、偏斜生成与多轮代理交互,我们系统扩展了异步强化学习框架DORA,实现了跨越超过20个领域、10000多个环境的大规模高效稳定训练。针对现实任务固有的噪声特性,我们系统分析并分解噪声模式,设计针对性训练流程,将不完美因素显式融入训练,从而增强实际应用中的鲁棒性。为进一步提升复杂推理任务表现,引入重思考模式,通过在推理时联合扩展深度与宽度,实现高效的并行思维放大。
原文摘要 · Abstract (English)
We introduce LongCat-Flash-Thinking-2601, a 560-billion-parameter open-source Mixture-of-Experts (MoE) reasoning model with superior agentic reasoning capability. LongCat-Flash-Thinking-2601 achieves state-of-the-art performance among open-source models on a wide range of agentic benchmarks, including agentic search, agentic tool use, and tool-integrated reasoning. Beyond benchmark performance, the model demonstrates strong generalization to complex tool interactions and robust behavior under noisy real-world environments. Its advanced capability stems from a unified training framework that combines domain-parallel expert training with subsequent fusion, together with an end-to-end co-design of data construction, environments, algorithms, and infrastructure spanning from pre-training to post-training. In particular, the model's strong generalization capability in complex tool-use are driven by our in-depth exploration of environment scaling and principled task construction. To optimize long-tailed, skewed generation and multi-turn agentic interactions, and to enable stable training across over 10,000 environments spanning more than 20 domains, we systematically extend our asynchronous reinforcement learning framework, DORA, for stable and efficient large-scale multi-environment training. Furthermore, recognizing that real-world tasks are inherently noisy, we conduct a systematic analysis and decomposition of real-world noise patterns, and design targeted training procedures to explicitly incorporate such imperfections into the training process, resulting in improved robustness for real-world applications. To further enhance performance on complex reasoning tasks, we introduce a Heavy Thinking mode that enables effective test-time scaling by jointly expanding reasoning depth and width through intensive parallel thinking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。