统一图像生成与编辑,用场景图精准控制物体关系和布局
SimGraph: A Unified Framework for Scene Graph-Based Image Generation and Editing
- 基于场景图统一生成与编辑,通过令牌生成与扩散编辑融合
- 在多个数据集上实现更优的空间一致性和语义连贯性
- 适合需要精细控制图像内容和关系的创作者与研究人员
生成式人工智能的进展显著提升了图像生成与编辑能力。然而,现有方法通常将两项任务分开处理,导致效率低下,并难以保持生成内容与编辑之间的空间一致性与语义连贯性。主要障碍在于缺乏对物体关系和空间布局的结构化控制。场景图方法通过结构化表示对象及其相互关系,为图像生成与编辑提供了更强的组合与交互控制能力。为此,我们提出 SimGraph,一个统一的框架,集成基于场景图的图像生成与编辑,实现对物体交互、布局及空间一致性的精确控制。该框架在单一场景图驱动模型中融合了基于令牌的生成与基于扩散的编辑,确保高质量且一致的结果。大量实验表明,该方法在多个基准上优于现有最先进方法。
原文摘要 · Abstract (English)
Recent advancements in Generative Artificial Intelligence (GenAI) have significantly enhanced the capabilities of both image generation and editing. However, current approaches often treat these tasks separately, leading to inefficiencies and challenges in maintaining spatial consistency and semantic coherence between generated content and edits. Moreover, a major obstacle is the lack of structured control over object relationships and spatial arrangements. Scene graph-based methods, which represent objects and their interrelationships in a structured format, offer a solution by providing greater control over composition and interactions in both image generation and editing. To address this, we introduce SimGraph, a unified framework that integrates scene graph-based image generation and editing, enabling precise control over object interactions, layouts, and spatial coherence. In particular, our framework integrates token-based generation and diffusion-based editing within a single scene graph-driven model, ensuring high-quality and consistent results. Through extensive experiments, we empirically demonstrate that our approach outperforms existing state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。