arXiv:2506.13307cs.CVcs.AI2025-06被引 6

用预训练扩散模型生成未见雷达图像,支持文本控制与罕见场景模拟。

Quantitative Comparison of Fine-Tuning Techniques for Pretrained Latent Diffusion Models in the Generation of Unseen SAR Images

  • 基于文本到图像模型,通过语义先验对齐雷达物理特性。
  • 混合策略在保持纹理与几何特征上最优,优于全量微调和纯LoRA。
  • 适合遥感数据增强、稀有场景仿真及多模态条件生成任务。

我们提出一个框架,将大型预训练潜在扩散模型适配于高分辨率合成孔径雷达(SAR)图像生成。该方法实现可控合成,并能生成训练集外的罕见或分布外场景。不从零训练专用小模型,而是利用开源文生图基础模型,通过其语义先验将提示词与雷达成像物理特性(侧视几何、斜距投影、具有重尾统计特性的相干斑点)对齐。基于10万张图像的SAR数据集,对比了全量微调与参数高效低秩适应(LoRA)在UNet扩散主干、变分自编码器(VAE)和文本编码器上的表现。评估结合三方面:(i) 与真实SAR幅度分布的统计距离,(ii) 借GLCM描述子衡量纹理相似性,(iii) 使用专用于SAR的CLIP模型评估语义对齐。结果表明,混合策略——全量调整UNet,文本编码器使用LoRA并学习令牌嵌入——最有效保留SAR几何结构与纹理,同时维持提示一致性。该框架支持文本控制与多模态条件输入(如分割图、TerraSAR-X或光学引导),为地球观测中的大规模SAR场景数据增广与未知场景仿真开辟新路径。

原文摘要 · Abstract (English)

We present a framework for adapting a large pretrained latent diffusion model to high-resolution Synthetic Aperture Radar (SAR) image generation. The approach enables controllable synthesis and the creation of rare or out-of-distribution scenes beyond the training set. Rather than training a task-specific small model from scratch, we adapt an open-source text-to-image foundation model to the SAR modality, using its semantic prior to align prompts with SAR imaging physics (side-looking geometry, slant-range projection, and coherent speckle with heavy-tailed statistics). Using a 100k-image SAR dataset, we compare full fine-tuning and parameter-efficient Low-Rank Adaptation (LoRA) across the UNet diffusion backbone, the Variational Autoencoder (VAE), and the text encoders. Evaluation combines (i) statistical distances to real SAR amplitude distributions, (ii) textural similarity via Gray-Level Co-occurrence Matrix (GLCM) descriptors, and (iii) semantic alignment using a SAR-specialized CLIP model. Our results show that a hybrid strategy-full UNet tuning with LoRA on the text encoders and a learned token embedding-best preserves SAR geometry and texture while maintaining prompt fidelity. The framework supports text-based control and multimodal conditioning (e.g., segmentation maps, TerraSAR-X, or optical guidance), opening new paths for large-scale SAR scene data augmentation and unseen scenario simulation in Earth observation.

SAR生成扩散模型微调遥感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。