一个模型同时搞定多模态图像生成与理解,省去多个模型切换的麻烦。
MMGen: Unified Multi-modal Image Generation and Understanding in One Go
- 用统一扩散框架整合多种生成与理解任务,一次推理完成多模态输出。
- 在Depth、Normals、Segmentation等任务上达到领先性能,支持跨模态条件生成。
- 适合需要多任务协同的视觉系统开发者,如自动驾驶、机器人感知。
统一的多模态生成与理解扩散框架具有实现无缝可控图像生成及其他跨模态任务的潜力。本文提出MMGen,一个将多种生成任务整合到单一扩散模型中的统一框架。包括:(1) 多模态类别条件生成,通过单次推理同时生成多模态输出;(2) 多模态视觉理解,从RGB图像准确预测深度、表面法线和分割图;(3) 多模态条件生成,根据特定模态条件和其他对齐模态生成对应RGB图像。该方法设计了一种新型扩散变换器,灵活支持多模态输出,并采用简单的模态解耦策略统一各类任务。大量实验与应用表明,MMGen在多种任务与条件下均表现出色,展现出在需同时进行生成与理解的应用中的巨大潜力。
原文摘要 · Abstract (English)
A unified diffusion framework for multi-modal generation and understanding has the transformative potential to achieve seamless and controllable image diffusion and other cross-modal tasks. In this paper, we introduce MMGen, a unified framework that integrates multiple generative tasks into a single diffusion model. This includes: (1) multi-modal category-conditioned generation, where multi-modal outputs are generated simultaneously through a single inference process, given category information; (2) multi-modal visual understanding, which accurately predicts depth, surface normals, and segmentation maps from RGB images; and (3) multi-modal conditioned generation, which produces corresponding RGB images based on specific modality conditions and other aligned modalities. Our approach develops a novel diffusion transformer that flexibly supports multi-modal output, along with a simple modality-decoupling strategy to unify various tasks. Extensive experiments and applications demonstrate the effectiveness and superiority of MMGen across diverse tasks and conditions, highlighting its potential for applications that require simultaneous generation and understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。