arXiv:2505.02471cs.CV2025-05被引 15

开源统一多模态框架,支持图文生成与编辑。

Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction

  • 设计统一视觉生成器与多模态自回归模型
  • 实现文本到图像生成与指令式图像编辑
  • 适合研究多模态交互与AGI的开发者

我们提出Ming-Lite-Uni,一个开源的多模态框架,包含全新设计的统一视觉生成器和原生多模态自回归模型,用于融合视觉与语言。该框架开源实现了MetaQueries与M2-omni集成架构,并引入多尺度可学习令牌与多尺度表征对齐策略。通过固定多模态大模型(MLLM)与可学习扩散模型的结合,Ming-Lite-Uni使原生多模态自回归模型具备文本到图像生成与基于指令的图像编辑能力,突破纯视觉理解限制。实验表明其性能优异,交互过程流畅自然。所有代码与模型权重已开源,以推动社区探索。该工作与2025年3月25日发布的ChatGPT-4o等同期多模态里程碑相呼应,凸显统一模型在迈向通用人工智能(AGI)路径上的重要性。目前Ming-Lite-Uni处于α阶段,将陆续优化。

原文摘要 · Abstract (English)

We introduce Ming-Lite-Uni, an open-source multimodal framework featuring a newly designed unified visual generator and a native multimodal autoregressive model tailored for unifying vision and language. Specifically, this project provides an open-source implementation of the integrated MetaQueries and M2-omni framework, while introducing the novel multi-scale learnable tokens and multi-scale representation alignment strategy. By leveraging a fixed MLLM and a learnable diffusion model, Ming-Lite-Uni enables native multimodal AR models to perform both text-to-image generation and instruction based image editing tasks, expanding their capabilities beyond pure visual understanding. Our experimental results demonstrate the strong performance of Ming-Lite-Uni and illustrate the impressive fluid nature of its interactive process. All code and model weights are open-sourced to foster further exploration within the community. Notably, this work aligns with concurrent multimodal AI milestones - such as ChatGPT-4o with native image generation updated in March 25, 2025 - underscoring the broader significance of unified models like Ming-Lite-Uni on the path toward AGI. Ming-Lite-Uni is in alpha stage and will soon be further refined.

多模态图像生成自回归模型开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。