arXiv:2603.20697cs.CVcs.AI2026-03中稿 · presentation at IG…被引 1

用卫星图生成灾后街景,帮救援人员看清实际损毁情况。

Satellite-to-Street: Synthesizing Post-Disaster Views from Satellite Imagery via Generative Vision Models

  • 用视觉语言模型和损伤敏感的专家混合模型,从卫星图生成街景。
  • 控制网模型语义准确率最高(0.71),但易虚构结构细节。
  • 提出新评估框架,兼顾像素质量、语义一致性和视觉感知匹配。

灾后快速获取真实现场信息至关重要。传统卫星图像虽能评估破坏范围,却缺乏地面视角,难以识别具体结构损伤;而地面影像在紧急情况下往往无法获取。本文研究卫星到街景的合成技术,提出两种生成策略:基于视觉语言模型(VLM)引导的方法和损伤敏感的专家混合(MoE)方法。在300个灾情场景上对比通用基线(Pix2Pix、ControlNet),采用多层级评估框架:(1)像素级质量评估,(2)基于ResNet的语义一致性验证,(3)新型VLM作为评判者进行感知对齐评估。实验显示存在真实感与保真度的权衡:扩散模型(如ControlNet)虽具高感知真实感,但常产生虚构结构;标准ControlNet在语义准确率上达0.71,表现最佳;而增强型VLM与MoE模型在纹理合理性上更优,但语义清晰度不足。本工作为可信跨视角合成建立基准,强调视觉逼真未必代表结构信息可靠,对灾后评估有重要警示意义。

原文摘要 · Abstract (English)

In the immediate aftermath of natural disasters, rapid situational awareness is critical. Traditionally, satellite observations are widely used to estimate damage extent. However, they lack the ground-level perspective essential for characterizing specific structural failures and impacts. Meanwhile, ground-level data (e.g., street-view imagery) remains largely inaccessible during time-sensitive events. This study investigates Satellite-to-Street View Synthesis to bridge this data gap. We introduce two generative strategies to synthesize post-disaster street views from satellite imagery: a Vision-Language Model (VLM)-guided approach and a damage-sensitive Mixture-of-Experts (MoE) method. We benchmark these against general-purpose baselines (Pix2Pix, ControlNet) using a proposed Structure-Aware Evaluation Framework. This multi-tier protocol integrates (1) pixel-level quality assessment, (2) ResNet-based semantic consistency verification, and (3) a novel VLM-as-a-Judge for perceptual alignment. Experiments on 300 disaster scenarios reveal a critical realism--fidelity trade-off: while diffusion-based approaches (e.g., ControlNet) achieve high perceptual realism, they often hallucinate structural details. Quantitative results show that standard ControlNet achieves the highest semantic accuracy, 0.71, whereas VLM-enhanced and MoE models excel in textural plausibility but struggle with semantic clarity. This work establishes a baseline for trustworthy cross-view synthesis, emphasizing that visually realistic generations may still fail to preserve critical structural information required for reliable disaster assessment.

图像生成灾后重建视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。