用可控生成合成遥感分割数据,提升小样本复杂场景效果
Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentation
- 用多模态扩散变换器联合控制文本图像掩码生成
- 在少样本和复杂场景下分割精度提升12.3%以上
- 适合遥感图像标注成本高、数据稀缺的研究者
随着可控生成技术的快速发展,训练数据合成已成为扩展遥感(RS)标注数据集、缓解人工标注负担的有前景方法。然而,语义掩码控制的复杂性与采样质量的不确定性常限制合成数据在下游语义分割任务中的应用。为此,本文提出面向任务的数据合成框架TODSynth,包含具备统一三重注意力的多模态扩散变换器(MM-DiT)及基于任务反馈的即插即用采样策略。基于强大的DiT生成基础模型,系统评估不同控制方案,发现结合文本-图像-掩码联合注意力并全量微调图像与掩码分支的方法显著提升了遥感语义分割数据合成的有效性,尤其在少样本与复杂场景下表现优异。此外,提出一种控制-修正流匹配(CRFM)方法,在早期高可塑阶段动态调整采样方向,由语义损失引导,缓解生成图像的不稳定性,缩小合成数据与下游分割任务间的差距。大量实验表明,本方法持续优于现有先进可控生成方法,生成更稳定、任务导向的遥感分割合成数据。
原文摘要 · Abstract (English)
With the rapid progress of controllable generation, training data synthesis has become a promising way to expand labeled datasets and alleviate manual annotation in remote sensing (RS). However, the complexity of semantic mask control and the uncertainty of sampling quality often limit the utility of synthetic data in downstream semantic segmentation tasks. To address these challenges, we propose a task-oriented data synthesis framework (TODSynth), including a Multimodal Diffusion Transformer (MM-DiT) with unified triple attention and a plug-and-play sampling strategy guided by task feedback. Built upon the powerful DiT-based generative foundation model, we systematically evaluate different control schemes, showing that a text-image-mask joint attention scheme combined with full fine-tuning of the image and mask branches significantly enhances the effectiveness of RS semantic segmentation data synthesis, particularly in few-shot and complex-scene scenarios. Furthermore, we propose a control-rectify flow matching (CRFM) method, which dynamically adjusts sampling directions guided by semantic loss during the early high-plasticity stage, mitigating the instability of generated images and bridging the gap between synthetic data and downstream segmentation tasks. Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art controllable generation methods, producing more stable and task-oriented synthetic data for RS semantic segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。