5600亿参数多模态模型,支持低延迟音视频实时交互。
LongCat-Flash-Omni Technical Report
- 渐进式训练策略,从简单到复杂逐步提升多模态建模能力。
- 仅激活270亿参数,实现90%以上文本训练吞吐效率。
- 开源多模态模型,适用于音视频理解与生成任务。
我们提出LongCat-Flash-Omni,一个拥有5600亿参数的开源多模态模型,具备实时音视频交互能力。通过受课程启发的渐进式训练策略,模型在逐步增加模态序列建模复杂度的过程中,同时保持强大的单模态性能。该模型基于具备零计算专家的快捷连接混合专家(MoE)架构的LongCat-Flash,集成高效的多模态感知与语音重建模块。尽管参数量达5600亿(激活270亿),仍可实现低延迟实时交互。针对训练基础设施,我们设计了专门处理多模态数据与模型异构性的模态解耦并行方案,其效率超过文本训练的90%。大量评估表明,该模型在开源多模态基准上达到顶尖水平,并在文本、图像、视频理解及音频生成等任务中表现优异。本文详述模型架构、训练流程与数据策略,并开源模型以推动社区发展。
原文摘要 · Abstract (English)
We introduce LongCat-Flash-Omni, a state-of-the-art open-source omni-modal model with 560 billion parameters, excelling at real-time audio-visual interaction. By adopting a curriculum-inspired progressive training strategy that transitions from simpler to increasingly complex modality sequence modeling tasks, LongCat-Flash-Omni attains comprehensive multimodal capabilities while maintaining strong unimodal capability. Building upon LongCat-Flash, which adopts a high-performance Shortcut-connected Mixture-of-Experts (MoE) architecture with zero-computation experts, LongCat-Flash-Omni integrates efficient multimodal perception and speech reconstruction modules. Despite its immense size of 560B parameters (with 27B activated), LongCat-Flash-Omni achieves low-latency real-time audio-visual interaction. For training infrastructure, we developed a modality-decoupled parallelism scheme specifically designed to manage the data and model heterogeneity inherent in large-scale multimodal training. This innovative approach demonstrates exceptional efficiency by sustaining over 90% of the throughput achieved by text-only training. Extensive evaluations show that LongCat-Flash-Omni achieves state-of-the-art performance on omni-modal benchmarks among open-source models. Furthermore, it delivers highly competitive results across a wide range of modality-specific tasks, including text, image, and video understanding, as well as audio understanding and generation. We provide a comprehensive overview of the model architecture design, training procedures, and data strategies, and open-source the model to foster future research and development in the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。