arXiv:2512.23239cs.CV2025-12被引 4

高比例无训练数据剪枝,提升遥感生成模型训练效率与质量

RS-Prune: Training-Free Data Pruning at High Ratios for Efficient Remote Sensing Diffusion Foundation Models

  • 基于熵值剔除低信息样本,再通过场景感知聚类与分层采样保留多样性
  • 在剪除85%数据后仍实现快速收敛和领先生成质量
  • 适合需要高效构建遥感生成模型的研究者与工程师

基于扩散模型的遥感生成基础模型对下游任务至关重要,但其训练依赖大量全球代表性数据,常含冗余、噪声和类别不平衡,降低训练效率并阻碍收敛。现有方法通常聚合多个分类数据集或采用简单去重,忽视生成建模的分布需求及遥感图像异质性。为此,我们提出一种无需训练的两阶段数据剪枝方法,在高剪枝率下快速选取高质量子集,使初步基础模型快速收敛,并可作为生成、微调等应用的通用骨干。方法结合局部信息量与全局场景级多样性与代表性。首先,基于熵的准则高效去除低信息样本;其次,利用遥感场景分类数据集作为参考基准,实施场景感知聚类与分层采样,提升聚类效果同时降低大规模无标签数据的计算成本;最后,通过平衡簇内均匀性与样本代表性,在高剪枝率下实现细粒度选择,保持整体多样性和代表性。实验表明,即使剪除85%训练数据,本方法仍显著提升收敛速度与生成质量。基于该方法训练的扩散基础模型在超分辨率与语义图像合成等下游任务中持续达到最先进水平。该数据剪枝范式为发展遥感生成基础模型提供了实用指导。

原文摘要 · Abstract (English)

Diffusion-based remote sensing (RS) generative foundation models are cruial for downstream tasks. However, these models rely on large amounts of globally representative data, which often contain redundancy, noise, and class imbalance, reducing training efficiency and preventing convergence. Existing RS diffusion foundation models typically aggregate multiple classification datasets or apply simplistic deduplication, overlooking the distributional requirements of generation modeling and the heterogeneity of RS imagery. To address these limitations, we propose a training-free, two-stage data pruning approach that quickly select a high-quality subset under high pruning ratios, enabling a preliminary foundation model to converge rapidly and serve as a versatile backbone for generation, downstream fine-tuning, and other applications. Our method jointly considers local information content with global scene-level diversity and representativeness. First, an entropy-based criterion efficiently removes low-information samples. Next, leveraging RS scene classification datasets as reference benchmarks, we perform scene-aware clustering with stratified sampling to improve clustering effectiveness while reducing computational costs on large-scale unlabeled data. Finally, by balancing cluster-level uniformity and sample representativeness, the method enables fine-grained selection under high pruning ratios while preserving overall diversity and representativeness. Experiments show that, even after pruning 85\% of the training data, our method significantly improves convergence and generation quality. Furthermore, diffusion foundation models trained with our method consistently achieve state-of-the-art performance across downstream tasks, including super-resolution and semantic image synthesis. This data pruning paradigm offers practical guidance for developing RS generative foundation models.

遥感生成数据剪枝扩散模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。