让文字生成图像时精准定位物体,提升空间布局准确性。
SpatialLock: Precise Spatial Control in Text-to-Image Synthesis
- 通过注意力层直接注入位置信息,强化模型对物体位置的感知。
- 在多个数据集上实现超过0.9的IOU,显著提升定位精度。
- 适合需要精确控制物体位置的图像生成场景。
文本到图像(T2I)合成近年来取得显著进展,推动了自动数据集生成等应用。然而,生成图像中物体的精确定位仍面临挑战。现有方法未能充分使用位置信息,导致对物体空间布局理解不足。为此,我们提出SpatialLock框架,利用感知信号与定位信息联合控制空间位置生成。该框架包含两个组件:位置增强注入(PoI)和位置引导学习(PoG)。PoI通过注意力层直接整合空间信息,促进模型有效学习定位信息;PoG采用基于感知的监督进一步优化物体定位。两者结合使模型能够生成具有精确空间布局的物体,并提升图像视觉质量。实验表明,SpatialLock在多个数据集上实现了超过0.9的IOU,达到当前最优水平。
原文摘要 · Abstract (English)
Text-to-Image (T2I) synthesis has made significant advancements in recent years, driving applications such as generating datasets automatically. However, precise control over object localization in generated images remains a challenge. Existing methods fail to fully utilize positional information, leading to an inadequate understanding of object spatial layouts. To address this issue, we propose SpatialLock, a novel framework that leverages perception signals and grounding information to jointly control the generation of spatial locations. SpatialLock incorporates two components: Position-Engaged Injection (PoI) and Position-Guided Learning (PoG). PoI directly integrates spatial information through an attention layer, encouraging the model to learn the grounding information effectively. PoG employs perception-based supervision to further refine object localization. Together, these components enable the model to generate objects with precise spatial arrangements and improve the visual quality of the generated images. Experiments show that SpatialLock sets a new state-of-the-art for precise object positioning, achieving IOU scores above 0.9 across multiple datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。