用扩散模型实现地理空间内容并行生成,突破传统序列式方法瓶颈。
GeoDiT: A Diffusion-based Vision-Language Model for Geospatial Understanding
- 提出基于扩散模型的并行生成框架,匹配地理空间数据的固有结构。
- 在图像描述、视觉定位和多目标检测任务上显著超越现有方法。
- 适合需要精准结构化输出的地理分析场景,如遥感图像理解。
自回归模型在结构上与地理空间理解的固有并行性不匹配,迫使场景采用僵化的序列叙述,从根本上阻碍了结构化、连贯输出的生成。我们挑战这一范式,将地理空间生成重构为一种并行优化过程,实现全局性的粗粒度到细粒度合成,同时解析所有语义元素。为此,我们提出GeoDiT,首个面向地理空间领域的扩散型视觉-语言模型。大量实验表明,GeoDiT在需结构化、以对象为中心输出的基准测试中达到新SOTA。其在图像描述、视觉定位和多对象检测任务上取得显著提升,恰是自回归模型表现薄弱之处。本工作验证:将生成过程与数据内在结构对齐,是释放复杂地理空间分析卓越性能的关键。
原文摘要 · Abstract (English)
Autoregressive models are structurally misaligned with the inherently parallel nature of geospatial understanding, forcing a rigid sequential narrative onto scenes and fundamentally hindering the generation of structured and coherent outputs. We challenge this paradigm by reframing geospatial generation as a parallel refinement process, enabling a holistic, coarse-to-fine synthesis that resolves all semantic elements simultaneously. To operationalize this, we introduce GeoDiT, the first diffusion-based vision-language model tailored for the geospatial domain. Extensive experiments demonstrate that GeoDiT establishes a new state-of-the-art on benchmarks requiring structured, object-centric outputs. It achieves significant gains in image captioning, visual grounding, and multi-object detection, precisely the tasks where autoregressive models falter. Our work validates that aligning the generative process with the data's intrinsic structure is key to unlocking superior performance in complex geospatial analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。