arXiv:2504.19432cs.CVcs.AI2025-04被引 2

EarthMapper实现卫星图与地图双向可控生成,精度与细节兼备。

EarthMapper: Visual Autoregressive Models for Controllable Bidirectional Satellite-Map Translation

  • 用地理坐标嵌入和多尺度对齐,统一双向翻译训练流程。
  • 在38城数据集上达到最高视觉真实感与结构一致性,优于现有方法。
  • 支持零样本补全与坐标条件生成,适合城市规划等实际应用。

卫星影像与地图是遥感中两种基础模态,分别提供地表直接观测与人类可读的地理抽象。卫星图与地图之间的双向转换(BSMT)在城市规划与灾害响应中有重要应用价值,但面临两大挑战:两模态间缺乏像素级精确对齐,且需同时实现高层地理特征抽象与高质量图像合成,技术复杂度高。为此,本文提出EarthMapper,一种新型自回归框架,用于可控双向卫星-地图转换。EarthMapper利用地理坐标嵌入锚定生成过程,确保区域自适应性,并通过地理条件联合尺度自回归(GJSA)机制实现多尺度特征对齐,统一双向翻译于单一训练周期内。引入语义注入(SI)机制提升特征一致性,提出关键点自适应引导(KPAG)机制,在推理时动态平衡多样性与精度。我们还构建了大规模数据集CNSatMap,包含38个中国城市的302,132对精确对齐的卫星-地图样本,支持稳健评估。在CNSatMap与纽约数据集上的实验表明,EarthMapper在视觉真实感、语义一致性和结构保真度上显著优于当前最优方法。此外,其在零样本图像补全、外推及坐标条件生成任务中表现优异,凸显其通用性。

原文摘要 · Abstract (English)

Satellite imagery and maps, as two fundamental data modalities in remote sensing, offer direct observations of the Earth's surface and human-interpretable geographic abstractions, respectively. The task of bidirectional translation between satellite images and maps (BSMT) holds significant potential for applications in urban planning and disaster response. However, this task presents two major challenges: first, the absence of precise pixel-wise alignment between the two modalities substantially complicates the translation process; second, it requires achieving both high-level abstraction of geographic features and high-quality visual synthesis, which further elevates the technical complexity. To address these limitations, we introduce EarthMapper, a novel autoregressive framework for controllable bidirectional satellite-map translation. EarthMapper employs geographic coordinate embeddings to anchor generation, ensuring region-specific adaptability, and leverages multi-scale feature alignment within a geo-conditioned joint scale autoregression (GJSA) process to unify bidirectional translation in a single training cycle. A semantic infusion (SI) mechanism is introduced to enhance feature-level consistency, while a key point adaptive guidance (KPAG) mechanism is proposed to dynamically balance diversity and precision during inference. We further contribute CNSatMap, a large-scale dataset comprising 302,132 precisely aligned satellite-map pairs across 38 Chinese cities, enabling robust benchmarking. Extensive experiments on CNSatMap and the New York dataset demonstrate EarthMapper's superior performance, achieving significant improvements in visual realism, semantic consistency, and structural fidelity over state-of-the-art methods. Additionally, EarthMapper excels in zero-shot tasks like in-painting, out-painting and coordinate-conditional generation, underscoring its versatility.

卫星图生成地图转换自回归模型地理信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。