一个稀疏统一模型,用61亿活跃参数实现多模态感知与生成的高效突破
Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
- 基于稀疏专家混合架构,仅6.1亿参数激活,提升计算效率与模型容量
- 视觉语言理解达Gemini 2.5 Pro水平,支持多轮无缝切换任务
- 可同时生成语音、音乐与音效,图像生成支持语义分割与高保真文字渲染
我们提出Ming-Flash-Omni,是Ming-Omni的升级版,基于稀疏的Ling-Flash-2.0 MoE结构,总参数量1000亿,每令牌仅激活61亿参数。该架构显著提升计算效率并扩展模型容量,推动跨视觉、语音、语言的统一多模态智能发展,迈向通用人工智能(AGI)的关键一步。相比前代,其在多模态理解与生成上均有显著提升:视觉语言理解基准得分接近Gemini 2.5 Pro;支持多轮交互中多模态任务无缝切换。语音方面,实现上下文与方言感知的强性能自动语音识别(ASR),并支持语音、声音与音乐的联合连续生成。视觉方面,引入生成式语义分割,达到独立竞争力,增强空间控制与编辑一致性,显著改善身份保留能力,并实现高保真图像内文字渲染。这些能力表明,单一统一模型可作为通用多模态智能的实用基础。
原文摘要 · Abstract (English)
We propose Ming-Flash-Omni, an upgraded version of Ming-Omni, built upon a sparser Mixture-of-Experts (MoE) variant of Ling-Flash-2.0 with 100 billion total parameters, of which only 6.1 billion are active per token. This architecture enables highly efficient scaling (dramatically improving computational efficiency while significantly expanding model capacity) and empowers stronger unified multimodal intelligence across vision, speech, and language, representing a key step toward Artificial General Intelligence (AGI). Compared to its predecessor, the upgraded version exhibits substantial improvements across multimodal understanding and generation. Notably, it achieves strong performance on vision-language understanding benchmarks, with overall scores on par with Gemini 2.5 Pro, and enables seamless switching among multimodal tasks in multi-turn interactions. In speech, it achieves strong performance in contextual and dialect-aware ASR while enabling joint, continuous-generation of speech, sound, and music. In vision, it introduces generative semantic segmentation that achieves competitive standalone performance and enhances spatial control and editing consistency, alongside marked improvements in identity preservation, and high-fidelity in-image text rendering. Together, these capabilities demonstrate that a single unified model can serve as a practical foundation for general-purpose multimodal intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。