5600亿参数开源推理模型,显著提升思维效率与多领域能力。
Introducing LongCat-Flash-Thinking: A Technical Report
- 通过长链式思维冷启动+分域并行训练,实现高效专家模型融合。
- 在AIME-25任务中令牌消耗减少64.5%,从19,653降至6,965。
- 适合关注高效推理、智能体系统与开源大模型研究者使用。
我们提出LongCat-Flash-Thinking,一个高效的5600亿参数开源混合专家(MoE)推理模型。其先进能力源于精心设计的训练流程:从长链式思维(CoT)数据冷启动开始,最终通过大规模强化学习(RL)完成优化。首先采用优化的冷启动策略,显著提升推理潜力,并赋予模型在形式化与智能体推理方面的专长。核心创新在于分域并行训练方案,将不同领域(如STEM、代码、智能体)的优化解耦,再融合为单一近似帕累托最优模型。整个过程由动态异步回滚调度系统(DORA)驱动,相比同步方法在数万加速器上实现三倍以上训练速度提升。结果表明,LongCat-Flash-Thinking在复杂推理任务中达到开源模型领先水平。在AIME-25上,智能体推理平均令牌消耗降低64.5%(从19,653降至6,965),且不降低任务准确率。模型已开源,以推动推理系统与智能体人工智能研究进展。
原文摘要 · Abstract (English)
We present LongCat-Flash-Thinking, an efficient 560-billion-parameter open-source Mixture-of-Experts (MoE) reasoning model. Its advanced capabilities are cultivated through a meticulously crafted training process, beginning with long Chain-of-Thought (CoT) data cold-start and culminating in large-scale Reinforcement Learning (RL). We first employ a well-designed cold-start training strategy, which significantly enhances the reasoning potential and equips the model with specialized skills in both formal and agentic reasoning. Then, a core innovation is our domain-parallel training scheme, which decouples optimization across distinct domains (e.g., STEM, Code, Agentic) and subsequently fuses the resulting expert models into a single, nearly Pareto-optimal model. This entire process is powered by our Dynamic ORchestration for Asynchronous rollout (DORA) system, a large-scale RL framework that delivers a greater than threefold training speedup over synchronous methods on tens of thousands of accelerators. As a result, LongCat-Flash-Thinking achieves state-of-the-art performance among open-source models on a suite of complex reasoning tasks. The model exhibits exceptional efficiency in agentic reasoning, reducing average token consumption by 64.5% (from 19, 653 to 6, 965) on AIME-25, without degrading task accuracy. We release LongCat-Flash-Thinking to promote further advances in reasoning systems and agentic AI research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。