开源多模态大模型Uni-MoE-2.0,支持文本、图像、语音生成与理解,性能超越多个顶尖模型。
Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- 采用动态容量专家混合架构,提升跨模态处理效率与能力。
- 在76个基准中超越Qwen2.5-Omni模型,视频理解与视听推理分别提升7%和4%。
- 适合需要多模态生成与复杂推理的开发者及研究者使用。
我们提出来自Lychee系列的Uni-MoE 2.0,一个全开源的多模态大模型(OLM),显著提升了语言主导的多模态理解、推理与生成能力。基于密集型语言模型,通过三项核心贡献:动态容量专家混合(MoE)设计、融合迭代强化学习的渐进式训练策略,以及精心构建的多模态数据匹配技术,从零开始构建Uni-MoE-2.0-Omni。该模型可实现跨模态理解,并生成图像、文本与语音。架构上,新MoE框架通过共享、路由与空专家机制,在10种跨模态输入下平衡计算效率与能力;3D RoPE确保自注意力层中时空跨模态对齐。训练方面,经跨模态预训练后,采用渐进式监督微调,激活模态特定专家,并结合均衡数据构成与迭代GSPO-DPO方法稳定强化学习训练,提升推理能力。数据上,基础模型在约750亿词的开源多模态数据上训练,配备专用语音与图像生成标记,使其能基于语言提示学习生成任务。在85个基准上的评估显示,该模型在76个基准中超过50个领先模型,性能优于训练数据达1.2万亿词的Qwen2.5-Omni,关键优势包括视频理解(平均+7%,共8项)、多模态理解(平均+7%,共4项)和音视频推理(+4%)。此外,在长时语音处理中降低4.2%的错误率,且在低级图像处理与可控生成5项指标上领先。
原文摘要 · Abstract (English)
We present Uni-MoE 2.0 from the Lychee family. As a fully open-source omnimodal large model (OLM), it substantially advances Lychee's Uni-MoE series in language-centric multimodal understanding, reasoning, and generating. Based on the dense LLM, we build Uni-MoE-2.0-Omni from scratch through three core contributions: dynamic-capacity Mixture-of-Experts (MoE) design, a progressive training strategy enhanced with an iterative reinforcement strategy, and a carefully curated multimodal data matching technique. It is capable of omnimodal understanding, as well as generating images, text, and speech. Architecturally, our new MoE framework balances computational efficiency and capability for 10 cross-modal inputs using shared, routed, and null experts, while our Omni-Modality 3D RoPE ensures spatio-temporal cross-modality alignment in the self-attention layer. For training, following cross-modal pretraining, we use a progressive supervised fine-tuning strategy that activates modality-specific experts and is enhanced by balanced data composition and an iterative GSPO-DPO method to stabilise RL training and improve reasoning. Data-wise, the base model, trained on approximately 75B tokens of open-source multimodal data, is equipped with special speech and image generation tokens, allowing it to learn these generative tasks by conditioning its outputs on linguistic cues. Extensive evaluation across 85 benchmarks demonstrates that our model achieves SOTA or highly competitive performance against leading OLMs, surpassing Qwen2.5-Omni (trained with 1.2T tokens) on over 50 of 76 benchmarks. Key strengths include video understanding (+7% avg. of 8), omnimodallity understanding (+7% avg. of 4), and audiovisual reasoning (+4%). It also advances long-form speech processing (reducing WER by 4.2%) and leads in low-level image processing and controllable generation across 5 metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。