用任意地理矢量数据生成卫星图像,支持不同标注成本的可控布局。
TerraDiT-$Ω$: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial Primitive

- 联合使用精确与粗略地理矢量(多边形、点等)作为控制信号。
- 在各类矢量输入下均优于密集与稀疏控制基线模型。
- 适合城市规划、遥感数据增强等需要矢量驱动生成的场景。
生成模型虽取得显著进展,但应用于卫星影像仍具挑战。与自然图像不同,卫星场景由空间复杂且语义各异的几何结构组成。以往工作通过适配自然图像框架,使用密集栅格或稀疏提示来应对,但牺牲了标注成本与生成保真度,并破坏了与常用地理信息矢量格式的兼容性。我们提出TerraDiT-Ω,一种统一的空间控制框架,可直接从任意原生地理空间矢量生成卫星图像。该模型同时利用精确标注(多边形、折线)与粗略标注(边界框、点),支持在不同标注预算下实现可控布局,拓展至城市规划等设计任务,且天然兼容端到端地理人工智能工作流。为有效利用这些矢量,在生成过程中引入几何感知局部注意力机制,将显式几何线索注入注意力空间。在所有条件输入格式下,本方法持续优于密集控制与稀疏控制基线。此外,这种灵活性使得仅用单一生成模型即可实现可控合成数据增强,提升下游任务在土地覆盖分割、目标检测、道路图提取和场景分类上的表现。代码、数据与权重见https://github.com/mvrl/TerraDiT。
原文摘要 · Abstract (English)
Generative models have achieved remarkable progress, yet applying them to satellite imagery remains challenging. Unlike natural imagery, satellite scenes are structured by spatially complex and semantically distinct geometries. Prior work addresses this complexity by adapting natural image frameworks using dense rasters or sparse prompts, trading off annotation cost and fidelity while breaking compatibility with vector primitives commonly used to represent geographic information. We introduce TerraDiT-$Ω$, a unified spatial control framework that generates satellite imagery directly from any native geospatial primitive. By jointly leveraging precise annotations (polygons, polylines) and coarser ones (bounding boxes, points), the model supports controllable layouts across varying annotation budgets, broadening applicability to design tasks such as urban planning while remaining naturally compatible with end-to-end GeoAI workflows. To effectively leverage these primitives during generation, we propose Geometry-Aware Local Attention, a conditioning mechanism that injects explicit geometric cues into the attention space. Across all conditioning formats, our approach consistently outperforms both dense-control and sparse-control baselines. Furthermore, this flexibility enables controllable synthetic data augmentation using a single generative model, improving downstream performance on land-cover segmentation, object detection, road graph extraction, and scene classification. Code, data, and weights are available at https://github.com/mvrl/TerraDiT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。