用街景语义和时间信息提升卫星图像碳排放预测精度
CarbonCLIP: Enhance Carbon Prediction from Satellite Imagery via Integrated Street-View Semantics and Temporal Context Training

- 通过双分支对比学习融合街景语义与月度时间特征
- 北京和新加坡实验中均显著优于现有方法
- 仅需预训练时用多模态数据,推理时只需卫星图
精准估算城市碳排放对可持续城市规划至关重要,但现有方法因数据源异质性及遥感数据缺乏细粒度语义-时间上下文而难以跨城市一致应用。我们提出CarbonCLIP,一种面向任务的多模态蒸馏框架,通过双分支对比学习将上下文知识迁移到统一卫星表征中,提升基于卫星图像的碳排放预测能力。空间分支利用大模型从街景图像自动生成的细粒度文本描述,提供反映建筑功能、基础设施与城市活动的语义先验;时间分支采用月编码器建模与月度排放变化相关的时序先验。CarbonCLIP仅在预训练阶段需要多模态数据,推理时仅依赖卫星图像,实现无地面数据下的可扩展部署。在北京和新加坡的实验表明,该方法在两地均显著优于基线模型。结果验证了多模态知识有效注入卫星表征,为基于卫星的城市碳建模提供了鲁棒解决方案。
原文摘要 · Abstract (English)
Accurately estimating urban carbon emissions is critical for sustainable urban planning, yet many existing approaches remain difficult to apply consistently across cities due to data-source heterogeneity and the lack of fine-grained semantic-temporal context in remote sensing data. We propose CarbonCLIP, a task-oriented multimodal distillation framework that improves satellite-based carbon emission prediction by transferring contextual knowledge into a unified satellite representation through dual-branch contrastive learning. Unlike conventional methods that rely on static visual features, CarbonCLIP explicitly bridges the gap between top-down satellite views and ground-level human activities. Specifically, the spatial branch uses fine-grained textual descriptions automatically generated from street-view images by Large Multimodal Models (LMMs) to provide semantic priors reflecting building functions, infrastructure, and urban activities, while the temporal branch employs a month encoder to encode temporal priors associated with monthly emission variation. CarbonCLIP requires multimodal data only during the pretraining phase; during inference, it relies solely on satellite imagery, thereby supporting scalable deployment when ground-level data are unavailable at inference. Experiments on Beijing and Singapore demonstrate that CarbonCLIP outperforms baselines in both study cities. The results validate that our method effectively transfers multimodal knowledge into satellite representations, offering a robust solution for satellite-based urban carbon modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。