用视觉语言模型自动生成卫星图像编辑数据,实现精准对象操控。
SatEdit: Mask-Conditioned Image Editing via VLM-Guided Segment Annotation

- 用分割模型生成掩码,视觉语言模型标注语义,轻量人工验证。
- 在91类、852个标注上微调,CLIP得分0.6322,提升语义一致性。
- 适合需要精确空间控制的遥感图像编辑任务,数据效率高。
卫星图像编辑需要精确的对象级控制,但构建带有掩码、语义标签和成对编辑的监督数据集成本高昂,因大规模上这些信息极少共现。本文提出SatEdit,一种基于掩码的卫星图像编辑框架,从无标注影像中构建训练监督信号:利用分割基础模型生成物体掩码,通过视觉语言模型(VLM)为采样段落分配语义标签,并经轻量人工验证后,通过掩码引导修复生成添加与移除的配对样本。在包含1,014张图像和852个验证对象标注的SODA-A衍生数据集上,使用LoRA微调高分辨率图像编辑主干网络。在与开源及专有图像编辑模型的对比中,SatEdit在受控测试中达到最高聚合掩码区域语义对齐,CLIP分数为0.6322,CLIP delta为0.0726,同时保持周边场景质量。结果表明,VLM辅助的段落标注是实现数据高效、空间可控卫星图像编辑的可行路径。
原文摘要 · Abstract (English)
Satellite image editing requires spatially precise object-level control, but supervised editing datasets for overhead imagery are costly to build because object masks, semantic labels, and paired edits are rarely available at scale. We introduce SatEdit, a mask-conditioned satellite image editing framework that constructs training supervision from unlabeled imagery. SatEdit proposes object masks with a seg- mentation foundation model, assigns semantic la- bels to sampled segments with a Vision-Language Model, and applies lightweight human verification before generating paired addition and removal exam- ples through mask-guided inpainting. We fine-tune a high-resolution image editing backbone with LoRA on a SODA-A-derived dataset containing 1,014 im- ages and 852 verified object annotations across 91 classes. In controlled comparisons with open- source and proprietary image editing models, SatE- dit achieves the highest aggregate masked-region se- mantic alignment, with a CLIP score of 0.6322 and CLIP delta of 0.0726, while preserving the surround- ing scene qualitatively. These results suggest that VLM-assisted segment annotation is a practical route to data-efficient, spatially controllable satellite image editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。