用多层中间表示打通图文理解与生成的桥梁
TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning
- 让扩散模型从MLLM多个中间层获取生成条件
- 图文任务统一性能超越现有方法
- 适合想高效构建多模态系统的研究者
本文提出TBAC-UniImage,一种新型多模态统一理解与生成模型。通过深度整合预训练扩散模型(作为生成阶梯)与多模态大语言模型(MLLM),实现理解与生成的深度融合。以往基于扩散模型的统一模型存在两大局限:一是仅使用MLLM最终隐藏状态作为生成条件,导致连接过浅;二是从零开始预训练统一架构,计算成本过高。为此,本工作探索新范式:不依赖单一输出,而是利用MLLM多个不同层次的中间表示作为扩散模型的生成条件。该方法将预训练生成器视为一座梯子,接收来自MLLM多层次理解过程的引导,从而实现更深层次、更精细的理解与生成统一。
原文摘要 · Abstract (English)
This paper introduces TBAC-UniImage, a novel unified model for multimodal understanding and generation. We achieve this by deeply integrating a pre-trained Diffusion Model, acting as a generative ladder, with a Multimodal Large Language Model (MLLM). Previous diffusion-based unified models face two primary limitations. One approach uses only the MLLM's final hidden state as the generative condition. This creates a shallow connection, as the generator is isolated from the rich, hierarchical representations within the MLLM's intermediate layers. The other approach, pretraining a unified generative architecture from scratch, is computationally expensive and prohibitive for many researchers. To overcome these issues, our work explores a new paradigm. Instead of relying on a single output, we use representations from multiple, diverse layers of the MLLM as generative conditions for the diffusion model. This method treats the pre-trained generator as a ladder, receiving guidance from various depths of the MLLM's understanding process. Consequently, TBAC-UniImage achieves a much deeper and more fine-grained unification of understanding and generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。