arXiv:2503.13814cs.CV2025-03被引 10

用少量标签训练遥感多模态模型,提升跨数据类型理解能力

FusDreamer: Label-efficient Remote Sensing World Model for Multimodal Data Classification

  • 构建统一表征框架,融合高光谱、激光雷达与文本数据
  • 仅需少量样本即可在新场景中准确分类,性能优于基线方法
  • 适合遥感领域少样本学习与跨模态分析的研究者

世界模型显著提升层次化理解能力,改善数据融合与学习效率。为探索其在遥感(RS)领域的潜力,本文提出一种面向多模态数据融合的标签高效遥感世界模型FusDreamer。FusDreamer以世界模型作为统一表征容器,抽象共性与高层知识,促进不同数据类型——高光谱(HSI)、激光雷达(LiDAR)与文本数据之间的交互。首先,采用新型潜在扩散融合与多模态生成范式(LaMG),具备优异的信息整合与细节保留能力;其次,引入开放世界知识引导的一致性投影(OK-CP)模块,通过提示表示对视觉描述对象进行建模,并利用对比学习对齐语言-视觉特征,从而在有限样本下微调预训练世界模型以弥合领域差异;最后,采用端到端多任务组合优化(MuCO)策略,捕捉细微特征偏差并协同引导扩散过程。在四个典型数据集上的实验验证了FusDreamer的有效性与优势。相关代码将发布于https://github.com/Cimy-wang/FusDreamer。

原文摘要 · Abstract (English)

World models significantly enhance hierarchical understanding, improving data integration and learning efficiency. To explore the potential of the world model in the remote sensing (RS) field, this paper proposes a label-efficient remote sensing world model for multimodal data fusion (FusDreamer). The FusDreamer uses the world model as a unified representation container to abstract common and high-level knowledge, promoting interactions across different types of data, \emph{i.e.}, hyperspectral (HSI), light detection and ranging (LiDAR), and text data. Initially, a new latent diffusion fusion and multimodal generation paradigm (LaMG) is utilized for its exceptional information integration and detail retention capabilities. Subsequently, an open-world knowledge-guided consistency projection (OK-CP) module incorporates prompt representations for visually described objects and aligns language-visual features through contrastive learning. In this way, the domain gap can be bridged by fine-tuning the pre-trained world models with limited samples. Finally, an end-to-end multitask combinatorial optimization (MuCO) strategy can capture slight feature bias and constrain the diffusion process in a collaboratively learnable direction. Experiments conducted on four typical datasets indicate the effectiveness and advantages of the proposed FusDreamer. The corresponding code will be released at https://github.com/Cimy-wang/FusDreamer.

遥感多模态融合少样本学习扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。