让图像生成理解物体遮挡关系,实现3D布局精准控制。
SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generation
- 用半透明3D盒子表示物体,显式建模遮挡关系。
- 在多物体场景中实现精确遮挡与相机视角控制。
- 适合需要真实遮挡和3D布局精度的应用场景。
我们识别出遮挡推理是3D布局条件生成中的关键但被忽视的方面,对生成具有深度一致几何和尺度的部分遮挡物体至关重要。现有方法虽能生成符合输入布局的逼真场景,但常无法准确建模物体间的遮挡关系。为此,我们提出SeeThrough3D,一种基于3D布局的生成模型,显式建模遮挡。引入遮挡感知3D场景表示(OSCR),将物体表示为置于虚拟环境中的半透明3D盒,并从指定相机视角渲染。透明度编码隐藏区域,使模型可推理遮挡,渲染视角则提供生成过程中的显式相机控制。通过引入由渲染3D表示生成的视觉令牌,对预训练的基于流的文本到图像生成模型进行条件化。同时,使用掩码自注意力机制,准确绑定每个物体边界框与其对应文本描述,避免属性混淆。为训练模型,构建包含多样多物体场景及强遮挡关系的合成数据集。SeeThrough3D在未见物体类别上具有良好泛化能力,实现具有真实遮挡和一致相机控制的精确3D布局生成。
原文摘要 · Abstract (English)
We identify occlusion reasoning as a fundamental yet overlooked aspect for 3D layout-conditioned generation. It is essential for synthesizing partially occluded objects with depth-consistent geometry and scale. While existing methods can generate realistic scenes that follow input layouts, they often fail to model precise inter-object occlusions. We propose SeeThrough3D, a model for 3D layout conditioned generation that explicitly models occlusions. We introduce an occlusion-aware 3D scene representation (OSCR), where objects are depicted as translucent 3D boxes placed within a virtual environment and rendered from desired camera viewpoint. The transparency encodes hidden object regions, enabling the model to reason about occlusions, while the rendered viewpoint provides explicit camera control during generation. We condition a pretrained flow based text-to-image image generation model by introducing a set of visual tokens derived from our rendered 3D representation. Furthermore, we apply masked self-attention to accurately bind each object bounding box to its corresponding textual description, enabling accurate generation of multiple objects without object attribute mixing. To train the model, we construct a synthetic dataset with diverse multi-object scenes with strong inter-object occlusions. SeeThrough3D generalizes effectively to unseen object categories and enables precise 3D layout control with realistic occlusions and consistent camera control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。