arXiv:2507.03257cs.CV2025-07ICCV被引 6

让图像生成理解3D布局,支持相机控制和场景整体建模。

LACONIC: A 3D Layout Adapter for Controllable Image Creation

  • 通过轻量适配器让文生图模型具备3D感知能力。
  • 首次实现对场景内内外物体的完整上下文建模。
  • 支持直观的3D编辑,适合需要精确控制的创作场景。

现有的多物体场景引导图像生成方法通常依赖图像或文本空间中的2D控制,难以保持一致的三维几何结构。本文提出一种新型条件控制方式、训练方法及适配网络,可无缝集成至预训练文生图扩散模型中。该方法赋予模型3D感知能力,同时利用其丰富的先验知识。支持相机控制、显式3D几何条件输入,并首次完整建模场景的全部上下文(包括屏幕内外物体),生成语义丰富且合理的图像。尽管具有多模态特性,模型仍保持轻量,所需监督数据合理,且展现出强大泛化能力。我们还引入直观一致的图像编辑与重风格化方法,例如通过位置、旋转或缩放调整场景中单个物体。该方法可良好融入多种图像创作流程,相比以往方法支持更丰富的应用场景。

原文摘要 · Abstract (English)

Existing generative approaches for guided image synthesis of multi-object scenes typically rely on 2D controls in the image or text space. As a result, these methods struggle to maintain and respect consistent three-dimensional geometric structure, underlying the scene. In this paper, we propose a novel conditioning approach, training method and adapter network that can be plugged into pretrained text-to-image diffusion models. Our approach provides a way to endow such models with 3D-awareness, while leveraging their rich prior knowledge. Our method supports camera control, conditioning on explicit 3D geometries and, for the first time, accounts for the entire context of a scene, i.e., both on and off-screen items, to synthesize plausible and semantically rich images. Despite its multi-modal nature, our model is lightweight, requires a reasonable number of data for supervised learning and shows remarkable generalization power. We also introduce methods for intuitive and consistent image editing and restyling, e.g., by positioning, rotating or resizing individual objects in a scene. Our method integrates well within various image creation workflows and enables a richer set of applications compared to previous approaches.

3D生成图像编辑扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。