用2D扩散模型生成3D物体空间关系数据,训练出可泛化的3D布局模型
Learning 3D Object Spatial Relationships from Pre-trained 2D Diffusion Models
- 利用预训练2D扩散模型生成含合理空间关系的图像并转为3D样本
- 基于3D样本训练得分引导的扩散模型,学习物体间相对空间分布
- 支持多物体布局与人体动作生成,适用于未见物体类别
本文提出一种从预训练2D扩散模型中学习物体对空间关系(OOR)的方法。我们假设2D扩散模型生成的图像天然包含真实的物体间空间关系线索,从而可高效构建用于学习各类无界物体类别的3D数据集。通过合成多样化的、蕴含合理空间关系的图像,并将其上采样为3D样本,我们构建了丰富的3D物体对数据集。在此基础上,训练了一个基于得分的OOR扩散模型以学习其相对空间分布。进一步地,通过约束成对关系的一致性并避免物体碰撞,将该方法扩展至多物体场景。大量实验表明,本方法在多种物体-物体空间关系上表现稳健,且适用于3D场景排布任务及人体动作合成。
原文摘要 · Abstract (English)
We present a method for learning 3D spatial relationships between object pairs, referred to as object-object spatial relationships (OOR), by leveraging synthetically generated 3D samples from pre-trained 2D diffusion models. We hypothesize that images synthesized by 2D diffusion models inherently capture realistic OOR cues, enabling efficient collection of a 3D dataset to learn OOR for various unbounded object categories. Our approach synthesizes diverse images that capture plausible OOR cues, which we then uplift into 3D samples. Leveraging our diverse collection of 3D samples for the object pairs, we train a score-based OOR diffusion model to learn the distribution of their relative spatial relationships. Additionally, we extend our pairwise OOR to multi-object OOR by enforcing consistency across pairwise relations and preventing object collisions. Extensive experiments demonstrate the robustness of our method across various object-object spatial relationships, along with its applicability to 3D scene arrangement tasks and human motion synthesis using our OOR diffusion model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。