让物体生成突破掩码限制,自动添加阴影反射,更真实自然。
Thinking Outside the BBox: Unconstrained Generative Object Compositing
- 基于扩散模型,不依赖输入掩码,实现无约束生成
- 可生成超出掩码范围的阴影与反射,提升图像真实感
- 支持空掩码自动放置物体,加速创意合成流程
将物体合成到图像中涉及多个复杂子任务,如物体位置与缩放、色彩/光照协调、视角/几何调整以及阴影与反射生成。现有生成式图像合成方法虽利用扩散模型统一处理这些任务,但受限于训练时对原物体进行掩码遮蔽,导致生成结果被束缚在输入掩码范围内。此外,在新图像中获取精确的物体位置与尺度掩码也极为困难。为此,我们提出全新的无约束生成式物体合成问题——生成不受掩码边界限制。我们在合成配对数据集上训练了一个首创的扩散模型,能够生成超出掩码范围的阴影与反射,显著增强图像真实感。若提供空掩码,模型可自动将物体放置于多样且自然的位置与尺度,大幅提升合成效率。该模型在多种质量指标和用户评估中均优于现有物体定位与合成方法。
原文摘要 · Abstract (English)
Compositing an object into an image involves multiple non-trivial sub-tasks such as object placement and scaling, color/lighting harmonization, viewpoint/geometry adjustment, and shadow/reflection generation. Recent generative image compositing methods leverage diffusion models to handle multiple sub-tasks at once. However, existing models face limitations due to their reliance on masking the original object during training, which constrains their generation to the input mask. Furthermore, obtaining an accurate input mask specifying the location and scale of the object in a new image can be highly challenging. To overcome such limitations, we define a novel problem of unconstrained generative object compositing, i.e., the generation is not bounded by the mask, and train a diffusion-based model on a synthesized paired dataset. Our first-of-its-kind model is able to generate object effects such as shadows and reflections that go beyond the mask, enhancing image realism. Additionally, if an empty mask is provided, our model automatically places the object in diverse natural locations and scales, accelerating the compositing workflow. Our model outperforms existing object placement and compositing models in various quality metrics and user studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。