提出全新统一多模态模型架构,兼顾图像理解与生成质量。
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
- 用扩散Transformer生成语义丰富图像特征,替代传统VAE
- 分阶段预训练策略提升理解能力并增强生成效果
- 开源全套资源,适合研究统一多模态模型的开发者
统一图像理解与生成是多模态模型的研究热点。尽管图像理解的设计已较成熟,但兼具生成能力的统一框架在架构与训练方法上仍不充分。本文基于自回归与扩散模型在高质量生成和可扩展性方面的潜力,系统研究了其在统一多模态场景中的应用,重点聚焦图像表征、建模目标与训练策略。我们提出一种新方法:采用扩散Transformer生成语义丰富的CLIP图像特征,相比传统VAE方案显著提升训练效率与生成质量。同时,实验表明先进行图像理解预训练、再进行图像生成预训练的分阶段策略,可在保持理解能力的同时有效提升生成性能。此外,我们通过GPT-4o生成多样化描述,精心构建了高质量指令微调数据集BLIP3o-60k。基于上述创新设计与数据,我们推出BLIP3-o系列统一多模态模型,在多数主流图像理解与生成基准上均达到领先水平。为推动后续研究,我们完全开源模型代码、权重、训练脚本及预训练与指令微调数据集。
原文摘要 · Abstract (English)
Unifying image understanding and generation has gained growing attention in recent research on multimodal models. Although design choices for image understanding have been extensively studied, the optimal model architecture and training recipe for a unified framework with image generation remain underexplored. Motivated by the strong potential of autoregressive and diffusion models for high-quality generation and scalability, we conduct a comprehensive study of their use in unified multimodal settings, with emphasis on image representations, modeling objectives, and training strategies. Grounded in these investigations, we introduce a novel approach that employs a diffusion transformer to generate semantically rich CLIP image features, in contrast to conventional VAE-based representations. This design yields both higher training efficiency and improved generative quality. Furthermore, we demonstrate that a sequential pretraining strategy for unified models-first training on image understanding and subsequently on image generation-offers practical advantages by preserving image understanding capability while developing strong image generation ability. Finally, we carefully curate a high-quality instruction-tuning dataset BLIP3o-60k for image generation by prompting GPT-4o with a diverse set of captions covering various scenes, objects, human gestures, and more. Building on our innovative model design, training recipe, and datasets, we develop BLIP3-o, a suite of state-of-the-art unified multimodal models. BLIP3-o achieves superior performance across most of the popular benchmarks spanning both image understanding and generation tasks. To facilitate future research, we fully open-source our models, including code, model weights, training scripts, and pretraining and instruction tuning datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。