从扩散模型中提取物体摆放空间先验,构建大规模真实场景标注数据集。
HiddenObjects: Scalable Diffusion-Distilled Spatial Priors for Object Placement

- 通过扩散模型修复生成密集物体摆放位置,实现全自动标注。
- 构建2700万条标注,覆盖2.7万张真实背景图,排名准确率提升45%。
- 轻量模型推理速度提升23万倍,适合实际应用部署。
我们提出一种方法,通过蒸馏文本条件扩散模型中隐含的物体摆放知识,学习显式的、类别相关的空间先验。以往工作依赖人工标注(规模有限)或基于修复的物体移除流程(易引入伪相关)。为此,我们设计了一个完全自动化且可扩展的框架,利用基于扩散的修复管道在高质量真实背景上评估密集物体摆放。该流程构建了HiddenObjects数据集,包含2700万条放置标注,覆盖2.7万张不同场景图像,并为不同图像和物体类别提供排序后的边界框插入结果。实验表明,我们的空间先验在下游图像编辑任务中表现优于稀疏人类标注(VLM-Judge得分3.90 vs. 2.68),显著超越现有摆放基线与零样本视觉语言模型。此外,我们将这些先验蒸馏为轻量模型,实现23万倍的推理加速。
原文摘要 · Abstract (English)
We propose a method to learn explicit, class-conditioned spatial priors for object placement in natural scenes by distilling the implicit placement knowledge encoded in text-conditioned diffusion models. Prior work relies either on manually annotated data, which is inherently limited in scale, or on inpainting-based object-removal pipelines, whose artifacts promote shortcut learning. To address these limitations, we introduce a fully automated and scalable framework that evaluates dense object placements on high-quality real backgrounds using a diffusion-based inpainting pipeline. With this pipeline, we construct HiddenObjects, a large-scale dataset comprising 27M placement annotations, evaluated across 27k distinct scenes, with ranked bounding box insertions for different images and object categories. Experimental results show that our spatial priors outperform sparse human annotations on a downstream image editing task (3.90 vs. 2.68 VLM-Judge), and significantly surpass existing placement baselines and zero-shot Vision-Language Models for object placement. Furthermore, we distill these priors into a lightweight model for fast practical inference (230,000x faster).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。