arXiv:2608.09244cs.CV2026-08

让图像生成模型在作图时实时调整,保持主体一致且贴合文字描述。

In-Loop Model Adaptation with Coupled Latent-Noise Guidance for High-Fidelity Subject-Driven Text-to-Image Generation

论文配图:In-Loop Model Adaptation with Coupled Latent-Noise Guidance for High-Fidelity Subject-Driven Text-to-Image Generation
图 1 · 摘自论文原文
  • 生成过程中动态调整模型参数,无需提前训练或微调。
  • 通过隐空间与噪声的联合损失,精准保留主体特征并匹配文本内容。
  • 适合需要快速生成个性化图像的场景,如创意设计、虚拟试衣。

文本到图像扩散模型在根据文本提示生成高质量图像方面取得了显著进展。主体驱动生成的目标是基于参考图像,在不同视觉语境下合成模仿该主体外观的定制化图像。核心挑战在于:当参考图像变化时,扩散模型难以在保持主体身份一致性的同时高效适应不同视觉语境。现有方法要么依赖大规模领域特定数据集进行训练,要么在生成前使用参考图像进行数百次迭代微调。本文提出一种新方法——【环内模型自适应】(In-Loop Model Adaptation, IMA),在实际生成过程的每一步动态调整核心扩散模型,而无需在生成前对参考图像进行预训练。为此,我们构建了一个DDIM反演链,将参考图像映射为一系列隐变量,同时建立仅基于文本提示的图像生成链。随后引入掩码隐变量一致性损失和噪声正则化损失,表征扩散模型与两条链在每一步生成中的隐变量-噪声差异。该耦合隐变量-噪声损失用于引导环内模型自适应,从而在保持参考图像所指定主体身份的同时,精确对齐文本提示,实现高保真度的主体驱动文本到图像生成。大量实验表明,所提IMA方法显著提升了主体驱动文本到图像生成性能。

原文摘要 · Abstract (English)

Text-to-image diffusion models have achieved remarkable success in generating high-quality images from a given text prompt. Subject-driven generation aims to synthesize customized images to mimic the appearance of subjects in given reference images within different visual contexts specified by the text prompts. The central challenge here is that, when the reference image changes, the diffusion model cannot efficiently adapt to different visual contexts while consistently maintaining the subject identity. Existing methods either train the model with a large domain-specific dataset or fine-tune the model using the reference image for hundreds of iterations before actual image generation. In this work, we explore a new approach, called \textit{In-Loop Model Adaptation} (IMA), which adapts the core diffusion model at each generation step during the actual process of image generation, without being trained on the reference image before the generation process. To this end, we establish a DDIM inversion chain that maps the reference image to a sequence of latent, as well as a text-to-image generation chain which generates the image from the text prompt only. We then introduce a masked latent consistency loss and a noise regularization loss to characterize the latent-noise difference between the diffusion model and these two chains at each generation step. This coupled latent-noise loss is used to guide the in-loop model adaptation to preserve the subject identity specified by the reference image while maintaining accurate alignment with the text prompt, resulting in high-fidelity text-to-image generation. Our extensive experiments demonstrate that our proposed IMA method significantly improves the performance of subject-driven text-to-image generation.

图像生成扩散模型主体驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。