arXiv:2512.00261cs.CV2025-12被引 1

用少量标注数据让扩散模型高效适配多模遥感图像

UniDiff: Parameter-Efficient Adaptation of Diffusion Models for Land Cover Classification with Multi-Modal Remotely Sensed Imagery and Sparse Annotations

  • 仅用目标域数据微调5%参数,结合时间步-模态条件控制
  • 在两个基准数据集上实现接近监督学习的分类精度
  • 适合标注稀缺的遥感场景,尤其对异构模态有效

稀疏标注严重制约多模态遥感应用:即使最先进的监督方法如MSFMamba也受限于标签数据,影响实际部署。尽管ImageNet预训练模型具备丰富视觉表征,但将其适配至高光谱成像(HSI)和合成孔径雷达(SAR)等异构模态,且无需大量标注数据仍具挑战。本文提出UniDiff,一种参数高效的框架,仅使用目标域数据即可将单一ImageNet预训练扩散模型适配至多种传感模态。该方法结合基于FiLM的时间步-模态条件控制、约5%参数的参数高效微调,以及伪RGB锚定机制,以保留预训练表征并防止灾难性遗忘。实验表明,在两个成熟的多模态基准数据集上,对预训练扩散模型进行无监督适配,能有效缓解标注限制,并实现多模态遥感数据的有效融合。

原文摘要 · Abstract (English)

Sparse annotations fundamentally constrain multimodal remote sensing: even recent state-of-the-art supervised methods such as MSFMamba are limited by the availability of labeled data, restricting their practical deployment despite architectural advances. ImageNet-pretrained models provide rich visual representations, but adapting them to heterogeneous modalities such as hyperspectral imaging (HSI) and synthetic aperture radar (SAR) without large labeled datasets remains challenging. We propose UniDiff, a parameter-efficient framework that adapts a single ImageNet-pretrained diffusion model to multiple sensing modalities using only target-domain data. UniDiff combines FiLM-based timestep-modality conditioning, parameter-efficient adaptation of approximately 5% of parameters, and pseudo-RGB anchoring to preserve pre-trained representations and prevent catastrophic forgetting. This design enables effective feature extraction from remote sensing data under sparse annotations. Our results with two established multi-modal benchmarking datasets demonstrate that unsupervised adaptation of a pre-trained diffusion model effectively mitigates annotation constraints and achieves effective fusion of multi-modal remotely sensed data.

扩散模型遥感少样本学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。