用视觉语言模型生成符合场地和功能需求的建筑体量方案
CoMa: Contextual Massing Generation with Vision-Language Models
- 基于视觉语言模型,根据功能与场地条件生成建筑形态
- 在CoMa-20K数据集上验证,零样本模型也能生成有上下文意识的方案
- 适合建筑智能化、生成式设计方向的研究者参考
建筑与城市规划的概念设计阶段,尤其是建筑体量设计,高度依赖设计师直觉与手工操作。为解决这一问题,我们提出一种基于功能需求与场地背景的自动化建筑体量生成框架。此前数据驱动方法的主要障碍是缺乏合适的数据集,因此我们构建了CoMa-20K数据集,包含详尽的体量几何、经济与使用功能数据,以及项目所在城市环境的视觉图像。我们以视觉语言模型(VLMs)为基础,将体量生成建模为条件任务,评估微调与大规模零样本模型的表现。实验揭示了该任务的内在复杂性,同时展示了VLM在生成具上下文敏感性的体量方案方面的潜力。该数据集与分析建立了基础基准,指明了未来数据驱动建筑设计研究的重要机遇。
原文摘要 · Abstract (English)
The conceptual design phase in architecture and urban planning, particularly building massing, is complex and heavily reliant on designer intuition and manual effort. To address this, we propose an automated framework for generating building massing based on functional requirements and site context. A primary obstacle to such data-driven methods has been the lack of suitable datasets. Consequently, we introduce the CoMa-20K dataset, a comprehensive collection that includes detailed massing geometries, associated economical and programmatic data, and visual representations of the development site within its existing urban context. We benchmark this dataset by formulating massing generation as a conditional task for Vision-Language Models (VLMs), evaluating both fine-tuned and large zero-shot models. Our experiments reveal the inherent complexity of the task while demonstrating the potential of VLMs to produce context-sensitive massing options. The dataset and analysis establish a foundational benchmark and highlight significant opportunities for future research in data-driven architectural design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。