让图像生成精准保留多个物体并严格对齐提示词。
Preserve Anything: Controllable Image Synthesis with Object Preservation
- 用多通道ControlNet实现物体位置、大小、颜色和细节的精准保留
- 生成高保真背景,准确还原光影与语义一致性,FID达15.26
- 支持用户显式控制布局与光照,适合需要精确控制的创作场景
我们提出全新的可控图像生成方法Preserve Anything,解决文本到图像生成中物体保持与语义一致性的关键问题。现有方法常无法同时保证多个物体的高保真保留、提示词语义对齐及场景构图的显式控制。新方法采用N通道ControlNet,集成物体保留模块(支持尺寸与位置无关的保留、颜色细节维持、伪影消除)、背景引导模块(生成高分辨率、语义一致的背景,包含准确阴影、光照与提示词匹配),以及显式布局与光照控制能力。框架还包括光照一致性约束与高频叠加模块,以保留细节并减少异常伪影。我们构建了一个包含24万张自然图像(筛选自美学质量)和1.8万张带元数据的3D渲染合成图像的新基准数据集,弥补现有数据集缺陷。实验表明,本方法在特征空间保真度(FID 15.26)和语义对齐度(CLIP-S 32.85)上均达领先水平,且在用户评估中,相比现有方法在提示词对齐、逼真度、AI伪影、自然美感四项指标上分别提升约25%、19%、13%、14%。
原文摘要 · Abstract (English)
We introduce \textit{Preserve Anything}, a novel method for controlled image synthesis that addresses key limitations in object preservation and semantic consistency in text-to-image (T2I) generation. Existing approaches often fail (i) to preserve multiple objects with fidelity, (ii) maintain semantic alignment with prompts, or (iii) provide explicit control over scene composition. To overcome these challenges, the proposed method employs an N-channel ControlNet that integrates (i) object preservation with size and placement agnosticism, color and detail retention, and artifact elimination, (ii) high-resolution, semantically consistent backgrounds with accurate shadows, lighting, and prompt adherence, and (iii) explicit user control over background layouts and lighting conditions. Key components of our framework include object preservation and background guidance modules, enforcing lighting consistency and a high-frequency overlay module to retain fine details while mitigating unwanted artifacts. We introduce a benchmark dataset consisting of 240K natural images filtered for aesthetic quality and 18K 3D-rendered synthetic images with metadata such as lighting, camera angles, and object relationships. This dataset addresses the deficiencies of existing benchmarks and allows a complete evaluation. Empirical results demonstrate that our method achieves state-of-the-art performance, significantly improving feature-space fidelity (FID 15.26) and semantic alignment (CLIP-S 32.85) while maintaining competitive aesthetic quality. We also conducted a user study to demonstrate the efficacy of the proposed work on unseen benchmark and observed a remarkable improvement of $\sim25\%$, $\sim19\%$, $\sim13\%$, and $\sim14\%$ in terms of prompt alignment, photorealism, the presence of AI artifacts, and natural aesthetics over existing works.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。