arXiv:2601.02783cs.CV2026-01被引 8

构建首个面向城市规划的遥感多模态理解与生成框架

EarthVL: A Progressive Earth Vision-Language Understanding and Generation Framework

  • 分阶段设计网络,先分割地物再推理关系
  • 10.9k张亚米级图像,76万+图文对支持问答任务
  • 首次实现遥感图像、掩码、文本三者统一建模

地球视觉在地物识别上已取得突破,但在关系推理方面仍显不足,制约了场景的全面理解。为此,提出一种渐进式地球视觉-语言理解与生成框架,包含多任务数据集 EarthVLSet 与语义引导网络 EarthVLNet。聚焦城市规划应用,EarthVLSet 包含 10.9 千张亚米级遥感图像、地物掩码及 761.5 千组图文对,涵盖多项选择与开放式视觉问答任务。以对象为中心,EarthVLNet 分阶段实现语义分割、关系推理与综合理解:第一阶段进行地物分割,生成用于问答引导的对象语义;基于像素级语义,基于对象感知的大语言模型(LLM)执行关系推理与知识总结以生成答案。优化方面,提出数值差异损失,动态添加差异惩罚以应对不同物体的统计特性。三个基准测试(语义分割、多项选择、开放式 VQA)验证了 EarthVLNet 的优越性,揭示三大方向:1)分割特征即使在跨数据集场景下也持续提升 VQA 性能;2)多项选择任务对视觉编码器更敏感;3)开放式任务需更强的视觉与语言编码器。我们认为该数据集与方法将为「图像-掩码-文本」统一建模提供有益基准,推动地球视觉在地理应用中的发展。

原文摘要 · Abstract (English)

Earth vision has achieved milestones in geospatial object recognition but lacks exploration in object-relational reasoning, limiting comprehensive scene understanding. To address this, a progressive Earth vision-language understanding and generation framework is proposed, including a multi-task dataset (EarthVLSet) and a semantic-guided network (EarthVLNet). Focusing on city planning applications, EarthVLSet includes 10.9k sub-meter resolution remote sensing images, land-cover masks, and 761.5k textual pairs involving both multiple-choice and open-ended visual question answering (VQA) tasks. In an object-centric way, EarthVLNet is proposed to progressively achieve semantic segmentation, relational reasoning, and comprehensive understanding. The first stage involves land-cover segmentation to generate object semantics for VQA guidance. Guided by pixel-wise semantics, the object awareness based large language model (LLM) performs relational reasoning and knowledge summarization to generate the required answers. As for optimization, the numerical difference loss is proposed to dynamically add difference penalties, addressing the various objects' statistics. Three benchmarks, including semantic segmentation, multiple-choice, and open-ended VQA demonstrated the superiorities of EarthVLNet, yielding three future directions: 1) segmentation features consistently enhance VQA performance even in cross-dataset scenarios; 2) multiple-choice tasks show greater sensitivity to the vision encoder than to the language decoder; and 3) open-ended tasks necessitate advanced vision encoders and language decoders for an optimal performance. We believe this dataset and method will provide a beneficial benchmark that connects ''image-mask-text'', advancing geographical applications for Earth vision.

遥感理解视觉问答多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。