用分层2D修补生成逼真可交互的3D场景
Architect: Generating Vivid and Interactive 3D Scenes with Hierarchical 2D Inpainting
- 基于扩散模型的分层迭代图像修补生成3D环境
- 支持从文本、平面图等起点生成或优化场景
- 适合机器人与具身AI研究者构建复杂虚拟环境
构建大规模可交互的3D环境对机器人学和具身AI研究至关重要。现有方法如人工设计、程序化生成、基于扩散的场景生成及大语言模型引导设计,受限于人力成本高、依赖预设规则或训练数据、三维空间推理能力弱等问题。由于预训练2D图像生成模型在场景与物体布局理解上优于大语言模型,我们提出Architect——一种利用基于扩散的2D图像修补生成复杂且真实的3D具身环境的生成框架。具体而言,通过基础视觉感知模型从图像中提取每个生成对象,并借助预训练深度估计模型将2D图像提升至3D空间。该流程进一步扩展为分层迭代的修补机制,持续生成大型家具与小型物品的布局以丰富场景。此迭代结构使方法具备灵活性,可从文本、平面图或已有布局等多种起点出发生成或优化场景。
原文摘要 · Abstract (English)
Creating large-scale interactive 3D environments is essential for the development of Robotics and Embodied AI research. Current methods, including manual design, procedural generation, diffusion-based scene generation, and large language model (LLM) guided scene design, are hindered by limitations such as excessive human effort, reliance on predefined rules or training datasets, and limited 3D spatial reasoning ability. Since pre-trained 2D image generative models better capture scene and object configuration than LLMs, we address these challenges by introducing Architect, a generative framework that creates complex and realistic 3D embodied environments leveraging diffusion-based 2D image inpainting. In detail, we utilize foundation visual perception models to obtain each generated object from the image and leverage pre-trained depth estimation models to lift the generated 2D image to 3D space. Our pipeline is further extended to a hierarchical and iterative inpainting process to continuously generate placement of large furniture and small objects to enrich the scene. This iterative structure brings the flexibility for our method to generate or refine scenes from various starting points, such as text, floor plans, or pre-arranged environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。