通过时空对齐提升城市图像语义理解,跨区域泛化更强
CoST: Semantic-Aware Urban Understanding via Spatial-Temporal Alignment

- 用对比学习对齐时空语义,捕捉跨区域通用地理规律
- 在8个城市指标任务中平均比最强方法提升8.7%性能
- 适合做城市分析、遥感解译的科研与工程人员参考
从卫星图像中进行地理空间表征学习是大规模城市分析和实际应用的基础问题。尽管近期取得进展,现有方法因依赖区域特定辅助数据且忽略多时相城市影像中的语义对齐,导致跨区域泛化能力差、语义可解释性弱。为此,我们提出CoST——一种基于对比学习的时空对齐框架,通过将空间上下文与多时相语义对齐,提取跨区域共享的通用地理规律。具体而言,CoST显式建模空间相关性以捕获可迁移的地理结构,并利用多年城市变化语义对齐学习到的表示与高层地理语义。大量实验表明,CoST在多种下游任务及未见场景中均表现优异,在八个城市指标设置下平均相对领先最强方法8.7%。代码已开源。
原文摘要 · Abstract (English)
Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and real-world applications. Despite recent advances, current methods struggle with cross-region generalization and semantic interpretability due to their reliance on region-specific auxiliary data and the neglect of semantic alignment within multi-temporal urban imagery. Therefore, we present CoST, a novel \underline{Co}ntrastive-based \underline{S}patial-\underline{T}emporal framework that aligns spatial context with multi-temporal semantics to extract universal geographic regularities shared across regions. Specifically, CoST explicitly models spatial correlations to capture transferable geographic structures and exploits multi-year urban change semantics to align learned representations with high-level geo-semantics. Extensive experiments demonstrate that CoST consistently achieves superior performance across various downstream tasks and in unseen scenario, yielding an average relative gain of 8.7\% over the strongest competing methods across eight city-indicator settings. The code is available in \href{https://github.com/Arandinglv/CoST}{this repo}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。