arXiv:2503.12213cs.CV2025-03中稿 · WACV2025被引 6

用布局生成多样逼真图像,风格可控且细节精准。

STAY Diffusion: Styled Layout Diffusion Model for Diverse Layout-to-Image Generation

  • 通过全局条件与语义图调节,实现风格化对象精细控制。
  • 在多样性、准确性和可控性上均超越现有最佳方法。
  • 适合需要高可控性图像生成的设计师与开发人员。

在布局到图像(L2I)合成中,从边界框等粗粒度信息生成复杂场景。该任务因输入布局能强引导生成过程且易于人工调整而具有广泛应用前景。本文提出基于扩散模型的STAY Diffusion,可生成逼真图像,并实现对场景中风格化物体的细粒度控制。方法为每张布局学习全局条件,并利用新型边缘感知归一化(EA Norm)自监督生成语义图以调节权重;引入样式掩码注意力(SM Attention),跨条件融合全局条件与图像特征以捕捉物体间关系。这些设计提供稳定引导,提升生成精度与可控性。大量基准测试表明,所提模型在图像质量、生成多样性、准确性和可控性方面均优于现有最先进方法。

原文摘要 · Abstract (English)

In layout-to-image (L2I) synthesis, controlled complex scenes are generated from coarse information like bounding boxes. Such a task is exciting to many downstream applications because the input layouts offer strong guidance to the generation process while remaining easily reconfigurable by humans. In this paper, we proposed STyled LAYout Diffusion (STAY Diffusion), a diffusion-based model that produces photo-realistic images and provides fine-grained control of stylized objects in scenes. Our approach learns a global condition for each layout, and a self-supervised semantic map for weight modulation using a novel Edge-Aware Normalization (EA Norm). A new Styled-Mask Attention (SM Attention) is also introduced to cross-condition the global condition and image feature for capturing the objects' relationships. These measures provide consistent guidance through the model, enabling more accurate and controllable image generation. Extensive benchmarking demonstrates that our STAY Diffusion presents high-quality images while surpassing previous state-of-the-art methods in generation diversity, accuracy, and controllability.

图像生成扩散模型风格控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。