统一图像生成与感知任务,保留模型先验能力。
UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception

- 基于MMDiT构建联合训练框架,仅复制图像分支建模多模态分布。
- 在多个任务上达到专用模型水平,且生成与感知互促提升。
- 适合需要多任务统一建模的视觉生成与理解场景。
扩散模型在可控图像生成和密集预测任务中表现优异,但现有方法通常将二者分开处理,忽视了异构分布联合建模的潜力。本文提出UniGP,基于MMDiT框架,通过简单联合训练统一可控生成与密集预测任务,无需复杂特定设计或损失函数,同时保留骨干模型的通用先验。模型通过在不同条件下学习生成与预测,有效捕捉图像-几何对的联合分布。UniGP包含DUGP模块和统一数据集训练策略:前者遵循奥卡姆剃刀原则,仅使用MMDiT的复制图像分支建模超出RGB的密集分布;后者将异构数据集整合进统一训练框架,联合建模生成与感知任务。大量实验表明,该统一模型超越先前统一方法,性能媲美专用模型。此外,多任务联合训练带来互补优势:生成先验丰富感知细节,感知学习增强生成结构一致性。
原文摘要 · Abstract (English)
Recent advances in diffusion models have shown impressive performance in controllable image generation and dense prediction tasks. However, existing approaches typically treat diffusion-based controllable generation and dense prediction as separate tasks, overlooking the potential benefits of jointly modeling the heterogeneous distributions. In this work, we introduce UniGP, a framework built upon MMDiT, which unifies controllable generation and dense prediction through simple joint training, without the need for complex task-specific designs or losses, while preserving the backbone's versatile priors. By learning controllable generation and prediction under different conditions, our model effectively captures the joint distribution of image-geometry pairs. UniGP is capable of versatile controllable generation, dense prediction, and joint generation. Specifically, the proposed UniGP consists of DUGP and a unified dataset training strategy. The former, following the principle of Occam's razor, uses only a copied image branch of MMDiT to model dense distributions beyond RGB, while the latter integrates heterogeneous datasets into a unified training framework to jointly model generation and perception tasks. Extensive experiments demonstrate that our unified model surpasses prior unified approaches and performs on par with specialized methods. Furthermore, we demonstrate that multi-task joint training provides complementary benefits: generative priors enrich perceptual details, while perceptual learning improves structural alignment in generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。