arXiv:2608.28567cs.CV2026-08

用通用视觉语言模型直接生成建筑损毁标签和位置,无需专用网络

GeBDA: Building Damage Assessment as Text-Based Sequence Prediction

论文配图:GeBDA: Building Damage Assessment as Text-Based Sequence Prediction
图 1 · 摘自论文原文
  • 将建筑损毁评估转化为文本序列自回归生成任务
  • 仅用双时相卫星图和提示词即可实现精准损毁定位与分级
  • 适合灾后快速评估、资源有限场景下的自动化分析

传统建筑损毁评估通常依赖专用网络架构或微调地理空间图像基础模型。本文探讨通用视觉语言模型(VLM)是否可通过自回归序列生成实现建筑定位与损毁等级判断。我们将BDA建模为预测可变长度的边界框序列,每个框包含坐标和损毁标签。基于开源Gemma模型的初步实现表明,仅需双时相卫星图像与合适文本提示,即可获得令人满意的损毁映射结果。

原文摘要 · Abstract (English)

Conventionally, Building Damage Assessment (BDA) is tackled either with dedicated network architectures or by fine-tuning geospatial image foundation models. In this work, we ask whether a general-purpose Vision-Language Model (VLM) can localize buildings and grade their damage through autoregressive sequence generation alone. We cast BDA as predicting a variable-length set of bounding boxes, each specified by its coordinates and a damage label. Our preliminary implementation, based on the open Gemma model, achieves promising damage mapping results from only bi-temporal satellite images and a suitable text prompt.

损毁评估视觉语言模型卫星图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。