用AI生成住宅建筑数据,解决能源建模缺数据难题。
Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity
- 结合公开图像与文本,用生成AI构建可模拟的建筑数据集
- 合成数据与真实数据重合度超95%,三类变量表现优异
- 适合能源建模、城市仿真等数据稀缺场景的研究者
计算模型已成为建筑与城市尺度能源研究的重要工具,支持多尺度数据驱动分析。然而,这些模型依赖大量建筑参数数据,而这些数据往往难以获取、采集成本高或受隐私限制。本文提出一种模块化框架,利用生成式人工智能从公开记录和图像中构建可直接用于仿真的建筑数据集。为验证框架可靠性,我们评估了其AI组件及整体输出效果:在选定图像上,LLaVA比基于GPT的替代方案具有更强的视觉聚焦能力;与国家级参考数据集对比,合成数据在四个变量中的三个重合度超过95%。本工作旨在减少对昂贵或受限数据源的依赖,降低建筑尺度能源研究与机器学习驱动的城市能源建模门槛,提供可用于能源建模、改造分析及城市级仿真的仿真就绪数据集。
原文摘要 · Abstract (English)
Computational models have emerged as powerful tools for multi-scale energy modeling research at the building and urban scale, supporting data-driven analysis across building and urban energy systems. However, these models require large amounts of building parameter data that is often inaccessible, expensive to collect, or subject to privacy constraints. We introduce a modular framework that applies generative Artificial Intelligence (AI) to construct simulation-ready building datasets from publicly available records and imagery. To improve the reliability of this framework, we evaluate both its AI components and its overall result. Our occlusion analysis demonstrates that for our selected images, LLaVA achieves greater visual focus than a GPT-based alternative for building image processing. We also assess plausibility of our results against a national reference dataset, finding that our synthetic data overlaps more than 95% for three of the four selected variables. This work aims to reduce dependence on costly or restricted data sources, lowering barriers to building-scale energy research and Machine Learning (ML)-driven urban energy modeling, thereby providing simulation-ready datasets intended to support downstream applications such as energy modeling, retrofit analysis, and urban-scale simulation under data scarcity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。