arXiv:2602.00110cs.CVcs.LG2026-02

用地理信息引导视觉变压器,提升疾病预测准确率

Observing Health Outcomes Using Remote Sensing Imagery and Geo-Context Guided Visual Transformer

  • 引入地理嵌入机制,将多源地理数据转为空间对齐的嵌入块
  • 通过动态注意力模块融合地理与影像信息,疾病预测准确率显著提升
  • 适合从事地理健康分析、城市规划的研究者使用

视觉变换器在遥感图像分析中取得显著进展,尤其在目标检测和分割任务上。近期的视觉语言与多模态模型通过引入文本描述、问答对和元数据等辅助信息,拓展了应用范围。然而,这些模型通常优化的是视觉与文本之间的语义对齐,而非地理空间理解,难以有效表示或推理结构化地理层。本文提出一种新模型,通过辅助地理空间信息增强遥感图像处理。该方法引入地理嵌入机制,将多样化的地理数据转换为与图像块空间对齐的嵌入块。为促进跨模态交互,设计了引导注意力模块,基于与辅助数据的相关性动态计算注意力权重,引导模型聚焦最相关区域。同时,不同注意力头被赋予不同角色,以捕捉引导信息的互补特征,提升预测可解释性。实验表明,所提框架在预测疾病患病率方面优于现有预训练地理基础模型,验证了其在多模态地理空间理解中的有效性。

原文摘要 · Abstract (English)

Visual transformers have driven major progress in remote sensing image analysis, particularly in object detection and segmentation. Recent vision-language and multimodal models further extend these capabilities by incorporating auxiliary information, including captions, question and answer pairs, and metadata, which broadens applications beyond conventional computer vision tasks. However, these models are typically optimized for semantic alignment between visual and textual content rather than geospatial understanding, and therefore are not suited for representing or reasoning with structured geospatial layers. In this study, we propose a novel model that enhances remote sensing imagery processing with guidance from auxiliary geospatial information. Our approach introduces a geospatial embedding mechanism that transforms diverse geospatial data into embedding patches that are spatially aligned with image patches. To facilitate cross-modal interaction, we design a guided attention module that dynamically integrates multimodal information by computing attention weights based on correlations with auxiliary data, thereby directing the model toward the most relevant regions. In addition, the module assigns distinct roles to individual attention heads, allowing the model to capture complementary aspects of the guidance information and improving the interpretability of its predictions. Experimental results demonstrate that the proposed framework outperforms existing pretrained geospatial foundation models in predicting disease prevalence, highlighting its effectiveness in multimodal geospatial understanding.

遥感分析多模态地理健康

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。