arXiv:2605.26111cs.CVcs.AI2026-05

用多模态大模型提升主体驱动图像生成的保真度与指令遵循能力

Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation

论文配图:Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation
图 1 · 摘自论文原文
  • 将文本和参考图联合编码,通过双层聚合模块优化多粒度特征融合
  • 在扩散模型中引入VAE身份条件,实现语义与细节的渐进式去噪平衡
  • 显著减少复制粘贴伪影,生成结果更符合人类偏好

主体驱动图像生成旨在合成保留给定主体身份的同时遵循文本指令的新图像。现有方法通常分别编码文本和参考图像,限制了跨模态推理能力并导致复制粘贴伪影。近期结合多模态模型与扩散模型的框架虽提升了指令遵循能力,但大多忽视身份保持。为此,我们基于联合编码文本与参考图像的多模态大语言模型(MLLM)条件化扩散模型,并引入基于VAE的身份条件机制。设计了新颖的双层聚合(DLA)模块以最优地聚合多层级MLLM特征,并采用多阶段去噪策略,在推理过程中逐步平衡来自MLLM的语义信息与来自VAE的精细身份细节。大量实验表明,该方法在主体驱动图像生成中实现了多模态理解与身份保真的协同,有效缓解复制粘贴问题,并在人类偏好评估中表现更优。

原文摘要 · Abstract (English)

Subject-driven image generation aims to synthesize new images that preserve the identity of the given subject while following textual instructions. Existing approaches often encode text and reference images separately. This limits cross-modal reasoning abilities and causes copy-paste artifacts. Recent frameworks that connect multimodal models and diffusion models improve instruction following, but largely overlook identity preservation. To address these limitations, we condition diffusion models on Multimodal Large Language Models (MLLMs) that jointly encode text and reference images, and augment it with VAE-based identity conditioning. A novel Dual Layer Aggregation (DLA) module is designed to aggregate multi-level MLLM features for optimal conditioning, and a multi-stage denoising strategy is applied to progressively balance the semantic information from MLLM and fine-detail identity from VAE during inference. Extensive experiments demonstrate that our approach harmonizes multimodal understanding with identity preservation, mitigates copy-paste issues, and achieves superior performance regarding human preference on subject-driven image generation. Our project website is available at https://zsh2000.github.io/squeeze-mllm-subject-gen/.

图像生成多模态扩散模型身份保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。