arXiv:2604.14373cs.CVcs.AI2026-04

用视觉语言模型分析卫星图,精准预测乡村脆弱性指数。

SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning

论文配图:SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning
图 1 · 摘自论文原文
  • 结合对比对齐与自举式图像描述生成,适配卫星影像语义。
  • 在县层级预测社会脆弱性指数,识别出屋顶、道路等关键风险特征。
  • 结果可解释,适合政策制定者与城乡规划研究者参考。

农村环境风险受本地条件(如房屋质量、道路可达性、地表形态)影响,但传统脆弱性指数粗略且难以揭示具体风险情境。本文提出 SatBLIP,一种面向卫星影像的视觉-语言框架,用于农村情境理解与特征识别,并实现县层级社会脆弱性指数(SVI)预测。该方法克服了以往遥感流程中手工特征提取、人工虚拟审计及自然图像预训练模型的局限,通过对比图像-文本对齐与针对卫星语义定制的自举式描述生成实现改进。利用 GPT-4o 生成卫星图像块的结构化描述(如屋顶类型/状态、房屋大小、院落属性、绿化与道路环境),再微调适配卫星数据的 BLIP 模型以生成未见图像的描述。这些描述经由 CLIP 编码,并与大语言模型生成的嵌入通过注意力机制融合,用于空间聚合下的 SVI 估计。借助 SHAP 分析,识别出驱动预测的关键属性(如屋顶形式/状态、街道宽度、植被覆盖、车辆与开放空间),实现对农村风险环境的可解释映射。

原文摘要 · Abstract (English)

Rural environmental risks are shaped by place-based conditions (e.g., housing quality, road access, land-surface patterns), yet standard vulnerability indices are coarse and provide limited insight into risk contexts. We propose SatBLIP, a satellite-specific vision-language framework for rural context understanding and feature identification that predicts county-level Social Vulnerability Index (SVI). SatBLIP addresses limitations of prior remote sensing pipelines-handcrafted features, manual virtual audits, and natural-image-trained VLMs-by coupling contrastive image-text alignment with bootstrapped captioning tailored to satellite semantics. We use GPT-4o to generate structured descriptions of satellite tiles (roof type/condition, house size, yard attributes, greenery, and road context), then fine-tune a satellite-adapted BLIP model to generate captions for unseen images. Captions are encoded with CLIP and fused with LLM-derived embeddings via attention for SVI estimation under spatial aggregation. Using SHAP, we identify salient attributes (e.g., roof form/condition, street width, vegetation, cars/open space) that consistently drive robust predictions, enabling interpretable mapping of rural risk environments.

卫星影像视觉语言社会脆弱性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。