arXiv:2411.19182cs.CVcs.AI2024-11被引 4

用可控扩散提升图像生成的语义一致性,让图文生成更精准。

SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation

  • 设计双向扩散机制,实现信息精准传递。
  • 结合MLLM动态调节扩散方向与强度,提升细节保真度。
  • 无需训练即可增强生成图像的上下文连贯性,适合视觉生成研究者。

受物理扩散现象启发,扩散生成模型通过数据空间中的随机游走模拟去噪过程,使信息在图像区域间扩散,产生和谐结果。然而,扩散过程的混沌性常导致区域间干扰,降低细节保留和上下文一致性。本文将无序扩散重构为文本-视觉-图像生成(TV2I)的强大工具,在保持像素级条件一致的同时确保整体视觉与语义连贯。提出循环单向扩散(COW),实现高效单向信息传递并减少干扰。在此基础上,引入选择性单向扩散(SOW),利用多模态大语言模型(MLLMs)解析图像中的语义与空间关系,并结合注意力机制,动态调控扩散的方向与强度。大量实验验证了受控信息扩散的潜力,为无需学习的自适应生成模型开辟新路径。

原文摘要 · Abstract (English)

Originating from the diffusion phenomenon in physics, which describes the random movement and collisions of particles, diffusion generative models simulate a random walk in the data space along the denoising trajectory. This allows information to diffuse across regions, yielding harmonious outcomes. However, the chaotic and disordered nature of information diffusion in diffusion models often results in undesired interference between image regions, causing degraded detail preservation and contextual inconsistency. In this work, we address these challenges by reframing disordered diffusion as a powerful tool for text-vision-to-image generation (TV2I) tasks, achieving pixel-level condition fidelity while maintaining visual and semantic coherence throughout the image. We first introduce Cyclic One-Way Diffusion (COW), which provides an efficient unidirectional diffusion framework for precise information transfer while minimizing disruptive interference. Building on COW, we further propose Selective One-Way Diffusion (SOW), which utilizes Multimodal Large Language Models (MLLMs) to clarify the semantic and spatial relationships within the image. Based on these insights, SOW combines attention mechanisms to dynamically regulate the direction and intensity of diffusion according to contextual relationships. Extensive experiments demonstrate the untapped potential of controlled information diffusion, offering a path to more adaptive and versatile generative models in a learning-free manner.

图像生成扩散模型MLLM上下文一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。