arXiv:2505.23661cs.CV2025-05被引 49

开源轻量模型统一多模态理解与生成,仅用1.1B参数达顶尖性能

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

  • 用可学习查询和轻量连接器打通LLM与扩散模型
  • 1.1B/3.1B参数下在多个基准上表现卓越
  • 适合追求高效开源方案的研究者和开发者

本文提出OpenUni,一个简单、轻量且完全开源的统一多模态理解与生成基线。受统一模型学习范式启发,我们采用高效训练策略,通过一组可学习查询和轻量级Transformer连接器,将现成的多模态大语言模型(LLMs)与扩散模型无缝衔接,显著降低训练复杂度与开销。在极简架构下,OpenUni可生成高质量、指令对齐的图像,并在GenEval、DPG-Bench和WISE等标准基准上实现优异表现,仅需1.1B和3.1B激活参数。为推动开放研究与社区发展,我们已公开全部模型权重、训练代码及自建数据集(含2300万张图像-文本对),地址见https://github.com/wusize/OpenUni。

原文摘要 · Abstract (English)

In this report, we present OpenUni, a simple, lightweight, and fully open-source baseline for unifying multimodal understanding and generation. Inspired by prevailing practices in unified model learning, we adopt an efficient training strategy that minimizes the training complexity and overhead by bridging the off-the-shelf multimodal large language models (LLMs) and diffusion models through a set of learnable queries and a light-weight transformer-based connector. With a minimalist choice of architecture, we demonstrate that OpenUni can: 1) generate high-quality and instruction-aligned images, and 2) achieve exceptional performance on standard benchmarks such as GenEval, DPG- Bench, and WISE, with only 1.1B and 3.1B activated parameters. To support open research and community advancement, we release all model weights, training code, and our curated training datasets (including 23M image-text pairs) at https://github.com/wusize/OpenUni.

多模态生成模型开源轻量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。