无需配对图像,用内在图像实现零样本3D物体合成。
ZeroComp: Zero-shot Object Compositing from Image Intrinsics via Diffusion
- 基于几何、反照率和遮蔽光照的内在图像控制生成
- 在真实场景中合成物体时,阴影自适应且视觉逼真
- 仅用合成室内数据训练,即可处理真实户外场景
我们提出ZeroComp,一种无需成对复合场景图像训练的零样本3D物体合成方法。该方法利用ControlNet从内在图像(几何、反照率、遮蔽光照)进行条件控制,并结合Stable Diffusion模型利用其场景先验,共同构成高效渲染引擎。训练阶段不依赖有无物体的配对图像,仅使用内在图像;训练完成后,可无缝将虚拟3D物体融入场景,并自适应调整阴影以生成逼真复合图像。我们构建了高质量评估数据集,结果表明,ZeroComp在定量指标和人类感知评测中均优于依赖显式光照估计与生成技术的方法。此外,该方法可扩展至真实和户外图像合成,即使仅在合成室内数据上训练也表现良好,充分展现其在图像合成中的有效性。
原文摘要 · Abstract (English)
We present ZeroComp, an effective zero-shot 3D object compositing approach that does not require paired composite-scene images during training. Our method leverages ControlNet to condition from intrinsic images and combines it with a Stable Diffusion model to utilize its scene priors, together operating as an effective rendering engine. During training, ZeroComp uses intrinsic images based on geometry, albedo, and masked shading, all without the need for paired images of scenes with and without composite objects. Once trained, it seamlessly integrates virtual 3D objects into scenes, adjusting shading to create realistic composites. We developed a high-quality evaluation dataset and demonstrate that ZeroComp outperforms methods using explicit lighting estimations and generative techniques in quantitative and human perception benchmarks. Additionally, ZeroComp extends to real and outdoor image compositing, even when trained solely on synthetic indoor data, showcasing its effectiveness in image compositing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。