arXiv:2602.20066cs.CVcs.AI2026-02

用卫星图零样本预测城市供暖需求,无需建筑级数据

HeatPrompt: Zero-Shot Vision-Language Modeling of Urban Heat Demand from Satellite Images

  • 用视觉语言模型从卫星图提取屋顶年龄等语义特征
  • 相比基线模型,误差降低30%,决定系数提升93.7%
  • 适合缺乏建筑数据的地区进行轻量级供热规划

精准的供暖需求地图对实现空间供暖脱碳至关重要,但多数市政部门缺乏建筑级别的详细数据。我们提出 HeatPrompt,一种零样本视觉-语言能源建模框架,通过语义特征从卫星图像、基础地理信息系统(GIS)和建筑级特征中估算年度供暖需求。将预训练的大规模视觉语言模型(VLMs)与领域特定提示结合,使其扮演能源规划角色,从RGB卫星图像中提取屋顶年龄、建筑密度等对应热负荷的视觉属性。基于这些描述词训练的多层感知机(MLP)回归器,相较于基线模型,决定系数(R²)提升93.7%,平均绝对误差(MAE)降低30%。定性分析显示高影响关键词与高需求区域对齐,为数据匮乏地区提供轻量级供热规划支持。

原文摘要 · Abstract (English)

Accurate heat-demand maps play a crucial role in decarbonizing space heating, yet most municipalities lack detailed building-level data needed to calculate them. We introduce HeatPrompt, a zero-shot vision-language energy modeling framework that estimates annual heat demand using semantic features extracted from satellite images, basic Geographic Information System (GIS), and building-level features. We feed pretrained Large Vision Language Models (VLMs) with a domain-specific prompt to act as an energy planner and extract the visual attributes such as roof age, building density, etc, from the RGB satellite image that correspond to the thermal load. A Multi-Layer Perceptron (MLP) regressor trained on these captions shows an $R^2$ uplift of 93.7% and shrinks the mean absolute error (MAE) by 30% compared to the baseline model. Qualitative analysis shows that high-impact tokens align with high-demand zones, offering lightweight support for heat planning in data-scarce regions.

视觉语言模型供暖预测卫星图像零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。