无需用户提示,一键解析图像中多个遮挡物体的层次关系
CObL: Toward Zero-Shot Ordinal Layering without User Prompting
- 用扩散模型并行生成多层物体,基于输入图反推完整结构
- 零样本泛化到真实桌面照片,可处理未知数量的新物体
- 突破以往需预设物体数或人工干预的限制,适合复杂场景理解
视觉理解依赖于将像素分组为对象并理解其空间关系,包括横向与深度上的排列。本文提出一种场景表示方法,即由遮挡顺序排列的“物体层”堆叠而成,每层包含一个独立且完整的物体。为此,我们引入基于扩散模型的CObL架构,可并行生成物体层,以Stable Diffusion为自然物体先验,并通过推理时引导确保重构回原图。训练仅使用数千张合成的多物体桌面场景图像,模型在未见的真实桌面照片上实现零样本泛化,能处理任意数量的新物体。相比现有方法,CObL无需用户提示,也不需事先知道物体数量;相较于以往无监督物体中心表示学习模型,CObL不受训练世界限制。
原文摘要 · Abstract (English)
Vision benefits from grouping pixels into objects and understanding their spatial relationships, both laterally and in depth. We capture this with a scene representation comprising an occlusion-ordered stack of "object layers," each containing an isolated and amodally-completed object. To infer this representation from an image, we introduce a diffusion-based architecture named Concurrent Object Layers (CObL). CObL generates a stack of object layers in parallel, using Stable Diffusion as a prior for natural objects and inference-time guidance to ensure the inferred layers composite back to the input image. We train CObL using a few thousand synthetically-generated images of multi-object tabletop scenes, and we find that it zero-shot generalizes to photographs of real-world tabletops with varying numbers of novel objects. In contrast to recent models for amodal object completion, CObL reconstructs multiple occluded objects without user prompting and without knowing the number of objects beforehand. Unlike previous models for unsupervised object-centric representation learning, CObL is not limited to the world it was trained in.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。