根据文字描述生成与场景一致的图像,保持原场景结构同时准确放置指定物体。
Setting the Stage: Text-Driven Scene-Consistent Image Generation
- 用真实照片+去实体+图像到视频扩散模型构建训练数据
- 引入对应引导注意力损失,强化文本与场景的空间对齐
- 在多个视角下生成符合指令的图像,适合场景一致性生成任务
我们聚焦于场景搭建(Scene Staging)这一基础任务:给定参考场景图像和包含目标角色类别及其与场景空间关系的文本条件,目标是合成一张保留参考图像场景身份、同时按文本描述正确生成目标角色的输出图像。现有方法在此任务上表现不佳,主要由于高质量成对数据稀缺和生成目标无约束。为突破数据瓶颈,我们提出一种新颖的数据构建流程,结合真实照片、实体移除与图像到视频扩散模型,生成具有多样化场景、视角及正确实体-场景关系的训练对。此外,我们引入一种新的对应引导注意力损失,利用跨视角线索强制与参考场景的空间对齐。在自建的场景一致性基准上的实验表明,我们的方法在自动指标和人类偏好评估中均优于当前最优基线。该方法能生成多样视角与构图的图像,忠实遵循文本指令并保留参考场景身份。
原文摘要 · Abstract (English)
We focus on the foundational task of Scene Staging: given a reference scene image and a text condition specifying an actor category to be generated in the scene and its spatial relation to the scene, the goal is to synthesize an output image that preserves the same scene identity as the reference image while correctly generating the actor according to the spatial relation described in the text. Existing methods struggle with this task, largely due to the scarcity of high-quality paired data and unconstrained generation objectives. To overcome the data bottleneck, we propose a novel data construction pipeline that combines real-world photographs, entity removal, and image-to-video diffusion models to generate training pairs with diverse scenes, viewpoints and correct entity-scene relationships. Furthermore, we introduce a novel correspondence-guided attention loss that leverages cross-view cues to enforce spatial alignment with the reference scene. Experiments on our scene-consistent benchmark show that our approach achieves better scene alignment and text-image alignment than state-of-the-art baselines, according to both automatic metrics and human preference studies. Our method generates images with diverse viewpoints and compositions while faithfully following the textual instructions and preserving the reference scene identity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。