用结构化场景图生成复杂图像,更准更可控。
Scene Graph Disentanglement and Composition for Generalizable Complex Image Generation
- 先用VAE从场景图提取布局与语义,实现一图多生。
- 结合扩散模型与细粒度属性,生成更合理图像。
- 支持修改场景图却不破坏原有视觉,适合编辑应用。
自然语言或布局条件下的图像生成已取得显著进展,但现有方法在复现复杂场景时仍存在不足,主要源于对多个物体及其关系建模不充分。为此,本文利用场景图这一结构化表示来提升复杂图像生成能力。不同于以往直接使用场景图进行生成的方法,我们以可泛化的形式结合变分自编码器(VAE)和扩散模型的生成能力,将场景图中的视觉线索解耦并组合。具体地,提出语义-布局变分自编码器(SL-VAE),从输入场景图联合推导出布局与语义,实现一对多的多样化、合理化生成。随后设计一种融合布局、语义与细粒度属性的组合掩码注意力机制(CMA),集成至扩散模型中作为生成引导。为实现图形修改同时保持视觉一致性,引入多层采样器(MLS),实现“隔离式”图像编辑效果。大量实验表明,该方法在生成合理性与可控性上均优于基于文本、布局或场景图的最新竞争方法。
原文摘要 · Abstract (English)
There has been exciting progress in generating images from natural language or layout conditions. However, these methods struggle to faithfully reproduce complex scenes due to the insufficient modeling of multiple objects and their relationships. To address this issue, we leverage the scene graph, a powerful structured representation, for complex image generation. Different from the previous works that directly use scene graphs for generation, we employ the generative capabilities of variational autoencoders and diffusion models in a generalizable manner, compositing diverse disentangled visual clues from scene graphs. Specifically, we first propose a Semantics-Layout Variational AutoEncoder (SL-VAE) to jointly derive (layouts, semantics) from the input scene graph, which allows a more diverse and reasonable generation in a one-to-many mapping. We then develop a Compositional Masked Attention (CMA) integrated with a diffusion model, incorporating (layouts, semantics) with fine-grained attributes as generation guidance. To further achieve graph manipulation while keeping the visual content consistent, we introduce a Multi-Layered Sampler (MLS) for an "isolated" image editing effect. Extensive experiments demonstrate that our method outperforms recent competitors based on text, layout, or scene graph, in terms of generation rationality and controllability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。