用AI从简图生成多视角一致的建筑图,提升设计自动化水平。
Multi-View Depth Consistent Image Generation Using Generative AI Models: Application on Architectural Design of University Buildings
- 三阶段框架融合ControlNet与多视角输入,增强生成一致性。
- 引入风格、结构与角度对齐损失,确保多视图统一性。
- 结合深度图与3D注意力机制,显著提升三维一致性效果。
在建筑设计初期,鞋盒模型常被用作建筑结构的简化表示,但将其转化为详细设计需大量操作。生成式人工智能(AI)为自动化这一过程提供了可能,但保持多视角一致性仍是重大挑战。为此,我们提出一种基于生成式AI模型的三阶段一致图像生成框架,可从鞋盒模型生成多视角一致的建筑图像。该方法优化扩散模型,采用ControlNet作为主干网络,并适配从预设视角获取的多视角鞋盒模型输入。为保证多视角图像在风格和结构上的一致性,提出一种图像空间损失模块,包含风格损失、结构损失与角度对齐损失。随后利用深度估计方法从生成的多视角图像中提取深度图。最后,将建筑图像与深度图配对作为输入,通过深度感知的3D注意力模块进一步提升多视角一致性。实验结果表明,所提框架能从鞋盒模型输入生成风格一致、结构连贯的多视角建筑图像。
原文摘要 · Abstract (English)
In the early stages of architectural design, shoebox models are typically used as a simplified representation of building structures but require extensive operations to transform them into detailed designs. Generative artificial intelligence (AI) provides a promising solution to automate this transformation, but ensuring multi-view consistency remains a significant challenge. To solve this issue, we propose a novel three-stage consistent image generation framework using generative AI models to generate architectural designs from shoebox model representations. The proposed method enhances state-of-the-art image generation diffusion models to generate multi-view consistent architectural images. We employ ControlNet as the backbone and optimize it to accommodate multi-view inputs of architectural shoebox models captured from predefined perspectives. To ensure stylistic and structural consistency across multi-view images, we propose an image space loss module that incorporates style loss, structural loss and angle alignment loss. We then use depth estimation method to extract depth maps from the generated multi-view images. Finally, we use the paired data of the architectural images and depth maps as inputs to improve the multi-view consistency via the depth-aware 3D attention module. Experimental results demonstrate that the proposed framework can generate multi-view architectural images with consistent style and structural coherence from shoebox model inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。