arXiv:2412.03558cs.CV2024-12CVPR被引 113

单图生成3D场景,一次搞定多个物体的空间关系。

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

论文配图:MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation
图 1 · 摘自论文原文
  • 用多实例扩散模型同时生成多个3D物体。
  • 在真实场景和合成数据上达最优,空间关系准确。
  • 适合需要快速生成复杂3D场景的开发者。

本文提出MIDI,一种从单张图像生成组合式3D场景的新范式。与依赖重建或检索的方法,或分步生成物体的现有方法不同,MIDI将预训练的图像到3D物体生成模型扩展为多实例扩散模型,实现多个3D实例的同步生成,保持精确的空间关系和高泛化能力。核心在于引入新型多实例注意力机制,在生成过程中直接建模物体间交互与空间一致性,无需复杂多阶段流程。输入包含部分物体图像与全局场景上下文,直接在3D生成中完成物体补全。训练时,利用有限场景级数据有效监督3D实例间的交互,并结合单物体数据进行正则化,以保持预训练模型的泛化能力。MIDI在合成数据、真实场景数据及文本到图像扩散模型生成的风格化图像上均表现卓越,达到当前最优水平。

原文摘要 · Abstract (English)

This paper introduces MIDI, a novel paradigm for compositional 3D scene generation from a single image. Unlike existing methods that rely on reconstruction or retrieval techniques or recent approaches that employ multi-stage object-by-object generation, MIDI extends pre-trained image-to-3D object generation models to multi-instance diffusion models, enabling the simultaneous generation of multiple 3D instances with accurate spatial relationships and high generalizability. At its core, MIDI incorporates a novel multi-instance attention mechanism, that effectively captures inter-object interactions and spatial coherence directly within the generation process, without the need for complex multi-step processes. The method utilizes partial object images and global scene context as inputs, directly modeling object completion during 3D generation. During training, we effectively supervise the interactions between 3D instances using a limited amount of scene-level data, while incorporating single-object data for regularization, thereby maintaining the pre-trained generalization ability. MIDI demonstrates state-of-the-art performance in image-to-scene generation, validated through evaluations on synthetic data, real-world scene data, and stylized scene images generated by text-to-image diffusion models.

3D生成扩散模型单图生成多物体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。