arXiv:2509.01322cs.CLcs.AI2025-09被引 55

5600亿参数模型实现高效推理与智能代理能力。

LongCat-Flash Technical Report

  • 零计算专家动态分配算力,每令牌平均激活270亿参数。
  • 推理速度超100 tokens/秒,每百万输出仅0.7美元成本。
  • 专为智能体任务优化,适合研究高效大模型与自动化应用。

我们提出LongCat-Flash,一个5600亿参数的专家混合(MoE)语言模型,兼顾计算效率与先进智能体能力。为实现可扩展效率,该模型引入两项创新:(a) 零计算专家,实现动态算力分配,根据上下文需求激活186亿至313亿参数(平均270亿/令牌),优化资源使用;(b) 快捷连接式MoE,扩大计算-通信重叠窗口,在同等规模下显著提升推理效率与吞吐量。我们构建了完整的大型模型扩展框架,融合超参数迁移、模型增长初始化、多维度稳定性方案及确定性计算,实现稳定可复现训练。借助架构与基础设施协同,我们在30天内完成超过20万亿令牌的训练,并实现推理速度超100 tokens/秒,每百万输出成本仅为0.7美元。为培养智能体能力,我们在优化混合数据上进行大规模预训练,并通过推理、代码与指令的中后段微调,辅以合成数据与工具使用任务增强。全面评估显示,作为非思考型基础模型,LongCat-Flash在主流模型中表现极具竞争力,尤其在智能体任务中优势突出。模型检查点已开源,促进社区研究。长猫聊天:https://longcat.ai Hugging Face:https://huggingface.co/meituan-longcat GitHub:https://github.com/meituan-longcat

原文摘要 · Abstract (English)

We introduce LongCat-Flash, a 560-billion-parameter Mixture-of-Experts (MoE) language model designed for both computational efficiency and advanced agentic capabilities. Stemming from the need for scalable efficiency, LongCat-Flash adopts two novel designs: (a) Zero-computation Experts, which enables dynamic computational budget allocation and activates 18.6B-31.3B (27B on average) per token depending on contextual demands, optimizing resource usage. (b) Shortcut-connected MoE, which enlarges the computation-communication overlap window, demonstrating notable gains in inference efficiency and throughput compared to models of a comparable scale. We develop a comprehensive scaling framework for large models that combines hyperparameter transfer, model-growth initialization, a multi-pronged stability suite, and deterministic computation to achieve stable and reproducible training. Notably, leveraging the synergy among scalable architectural design and infrastructure efforts, we complete model training on more than 20 trillion tokens within 30 days, while achieving over 100 tokens per second (TPS) for inference at a cost of \$0.70 per million output tokens. To cultivate LongCat-Flash towards agentic intelligence, we conduct a large-scale pre-training on optimized mixtures, followed by targeted mid- and post-training on reasoning, code, and instructions, with further augmentation from synthetic data and tool use tasks. Comprehensive evaluations demonstrate that, as a non-thinking foundation model, LongCat-Flash delivers highly competitive performance among other leading models, with exceptional strengths in agentic tasks. The model checkpoint of LongCat-Flash is open-sourced to foster community research. LongCat Chat: https://longcat.ai Hugging Face: https://huggingface.co/meituan-longcat GitHub: https://github.com/meituan-longcat

大模型MoE智能体高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。