arXiv:2603.19039cs.CV2026-03中稿 · CVPR被引 6

让遥感图像理解更精准,能定位像素级变化并跨时间分析。

TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation

  • 统一模型融合光学与雷达数据,支持多时相变化分析。
  • 在百万样本数据集上实现高精度像素级推理,准确率显著提升。
  • 适合遥感分析、城市规划等需要精确定位的科研与应用。

视觉语言模型(VLMs)在地球观测(EO)中展现潜力,但在需要精确像素级视觉表征的复杂空间推理任务上表现不佳。为此,我们提出TerraScope,一种统一的VLM,具备两大核心能力:(1) 模态灵活推理:可处理单一模态输入(光学或合成孔径雷达SAR),在双模态可用时自适应融合;(2) 多时相推理:整合时间序列,实现多时间点的变化分析。此外,我们构建了包含100万样本的Terra-CoT大规模数据集,其中嵌入像素级掩码的推理链覆盖多个数据源。我们还提出了TerraScope-Bench,首个针对像素级地理空间推理的基准测试,包含六个子任务,评估答案准确率与掩码质量,确保真实像素级推理。实验表明,TerraScope在像素级地理空间推理任务上显著优于现有VLM,并提供可解释的视觉证据。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have shown promise in earth observation (EO), yet they struggle with tasks that require grounding complex spatial reasoning in precise pixel-level visual representations. To address this problem, we introduce TerraScope, a unified VLM that delivers pixel-grounded geospatial reasoning with two key capabilities: (1) modality-flexible reasoning: it handles single-modality inputs (optical or SAR) and adaptively fuses different modalities into the reasoning process when both are available; (2) multi-temporal reasoning: it integrates temporal sequences for change analysis across multiple time points. In addition, we curate Terra-CoT, a large-scale dataset containing 1 million samples with pixel-level masks embedded in reasoning chains across multiple sources. We also propose TerraScope-Bench, the first benchmark for pixel-grounded geospatial reasoning with six sub-tasks that evaluates both answer accuracy and mask quality to ensure authentic pixel-grounded reasoning. Experiments show that TerraScope significantly outperforms existing VLMs on pixel-grounded geospatial reasoning while providing interpretable visual evidence.

遥感分析视觉语言模型像素级推理时空分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。