arXiv:2510.21391cs.CV2025-10被引 4

TerraGen统一生成遥感图像布局,支持多任务数据增强。

TerraGen: A Unified Multi-Task Layout Generation Framework for Remote Sensing Data Augmentation

  • 用地理空间布局编码器统一处理框和掩码输入,控制空间结构。
  • 在45000张图像上训练,生成质量优于现有方法。
  • 适合遥感检测、分割等任务的通用数据增强,少样本也有效。

遥感视觉任务需要跨多个相互关联领域的大量标注数据。然而,现有生成式数据增强框架彼此孤立,每个任务需独立训练生成模型,且忽略地理信息与空间约束建模。为此,我们提出TerraGen,一种统一的布局到图像生成框架,可灵活、可控地合成适用于检测、分割、提取等高阶视觉任务的遥感影像。TerraGen引入地理-空间布局编码器,统一处理边界框与分割掩码输入,并结合多尺度注入机制与掩码加权损失,显式编码从全局结构到细粒度的时空约束。同时,我们构建了首个大规模多任务遥感布局生成数据集,包含45,000张图像,并建立了标准化评估协议。实验表明,TerraGen在多种任务中均实现最优生成图像质量。此外,其作为通用数据增强生成器,显著提升下游任务性能,在全数据与少样本场景下均表现出强跨任务泛化能力。

原文摘要 · Abstract (English)

Remote sensing vision tasks require extensive labeled data across multiple, interconnected domains. However, current generative data augmentation frameworks are task-isolated, i.e., each vision task requires training an independent generative model, and ignores the modeling of geographical information and spatial constraints. To address these issues, we propose \textbf{TerraGen}, a unified layout-to-image generation framework that enables flexible, spatially controllable synthesis of remote sensing imagery for various high-level vision tasks, e.g., detection, segmentation, and extraction. Specifically, TerraGen introduces a geographic-spatial layout encoder that unifies bounding box and segmentation mask inputs, combined with a multi-scale injection scheme and mask-weighted loss to explicitly encode spatial constraints, from global structures to fine details. Also, we construct the first large-scale multi-task remote sensing layout generation dataset containing 45k images and establish a standardized evaluation protocol for this task. Experimental results show that our TerraGen can achieve the best generation image quality across diverse tasks. Additionally, TerraGen can be used as a universal data-augmentation generator, enhancing downstream task performance significantly and demonstrating robust cross-task generalisation in both full-data and few-shot scenarios.

遥感数据增强生成模型多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。