用地理矢量图斑生成高精度卫星图像,支持精准空间编辑。
VectorSynth: Fine-Grained Satellite Image Synthesis with Structured Semantics
- 通过矢量语义嵌入引导扩散模型生成图像
- 在真实城市场景中实现更高语义保真度与结构真实感
- 适合需要精确地理空间控制的遥感应用
我们提出VectorSynth,一种基于扩散模型的像素级卫星图像生成框架,可接受带有语义属性的多边形地理标注作为条件。与以往依赖文本或布局提示的方法不同,VectorSynth学习影像与语义矢量几何之间的密集跨模态对应关系,实现细粒度、空间定位准确的编辑。视觉语言对齐模块从多边形语义生成像素级嵌入,指导条件图像生成过程,确保空间范围与语义线索一致。该框架支持混合语言提示与几何感知条件的交互式工作流,适用于快速假设模拟、空间修改和地图驱动的内容生成。为训练与评估,我们构建了一个包含多样城市场景的卫星图像数据集,涵盖建筑与自然要素,并配有像素级配准的多边形标注。实验表明,相比已有方法,其在语义保真度与结构真实性方面有显著提升,且所训练的视觉语言模型展现出精细的空间定位能力。代码与数据已开源。
原文摘要 · Abstract (English)
We introduce VectorSynth, a diffusion-based framework for pixel-accurate satellite image synthesis conditioned on polygonal geographic annotations with semantic attributes. Unlike prior text- or layout-conditioned models, VectorSynth learns dense cross-modal correspondences that align imagery and semantic vector geometry, enabling fine-grained, spatially grounded edits. A vision language alignment module produces pixel-level embeddings from polygon semantics; these embeddings guide a conditional image generation framework to respect both spatial extents and semantic cues. VectorSynth supports interactive workflows that mix language prompts with geometry-aware conditioning, allowing rapid what-if simulations, spatial edits, and map-informed content generation. For training and evaluation, we assemble a collection of satellite scenes paired with pixel-registered polygon annotations spanning diverse urban scenes with both built and natural features. We observe strong improvements over prior methods in semantic fidelity and structural realism, and show that our trained vision language model demonstrates fine-grained spatial grounding. The code and data are available at https://github.com/mvrl/VectorSynth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。