融合遥感影像与表格数据,构建统一的地理空间表示模型
GeoViSTA: Geospatial Vision-Tabular Transformer for Multimodal Environment Representation

- 通过双模态交叉注意力对齐图像与人口普查区块数据
- 自监督预训练下在死亡率与火灾风险预测上超越基线
- 适合需要综合自然与社会因素的环境分析研究者
大规模地球观测图像预训练已生成强大的自然与人工环境表征。然而,现有地理空间基础模型大多未直接建模以表格形式存储的结构化社会经济变量,这种模态缺口限制了对完整环境的捕捉能力,而这对理解复杂环境、社会及健康结果至关重要。本文提出GeoViSTA(地理空间视觉-表格变换器),一种从共注册网格化影像与表格数据中学习统一地理空间嵌入的视觉-表格架构。GeoViSTA利用双边交叉注意力在模态间交换空间与语义信息,由地理感知注意力机制引导,将连续图像块与不规则的人口普查区块标记对齐。我们采用自监督联合掩码自编码目标训练GeoViSTA,强制其利用局部空间上下文和跨模态线索恢复缺失的图像块与表格行。实证表明,GeoViSTA的统一嵌入在高影响力下游任务上线性探测性能提升,在未见区域的疾病特异性死亡率与火灾风险频率预测中优于基线模型。结果表明,联合建模物理环境与结构化社会经济背景可生成高度可迁移的表征,实现全面的地理空间推理。
原文摘要 · Abstract (English)
Large-scale pretraining on Earth observation imagery has yielded powerful representations of the natural and built environment. However, most existing geospatial foundation models do not directly model the structured socioeconomic covariates typically stored in tabular form. This modality gap limits their ability to capture the complete total environment, which is critical for reasoning about complex environmental, social, and health-related outcomes. In this work, we propose GeoViSTA (Geospatial Vision-Tabular Transformer), a vision-tabular architecture that learns unified geospatial embeddings from co-registered gridded imagery and tabular data. GeoViSTA utilizes bilateral cross-attention to exchange spatial and semantic information across modalities, guided by a geography-aware attention mechanism that aligns continuous image patches with irregular census-tract tokens. We train GeoViSTA with a self-supervised joint masked-autoencoding objective, forcing it to recover missing image patches and tabular rows using local spatial context and cross-modal cues. Empirically, GeoViSTA's unified embeddings improve linear probing performance on high-impact downstream tasks, outperforming baselines in predicting disease-specific mortality and fire hazard frequency across held-out regions. These results demonstrate that jointly modeling the physical environment alongside structured socioeconomic context yields highly transferable representations for holistic geospatial inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。