arXiv:2502.17066cs.CVcs.LG2025-02ICML被引 5

用图像与激光雷达对齐,生成像素级嵌入,零样本监测环境更高效。

DUNIA: Pixel-Sized Embeddings via Cross-Modal Alignment for Earth Observation Applications

  • 通过图像与全波形激光雷达跨模态对齐,学习像素级嵌入
  • 零样本下在7项任务中表现优于专用监督模型,低数据时仍有效
  • 适合需要少标注数据的遥感环境监测应用

针对地球观测中自监督多模态学习的应用,现有方法多生成粗粒度块级嵌入,限制了与激光雷达等模态的融合。为此,本文提出DUNIA,通过图像与全波形激光雷达数据的跨模态对齐,学习像素级嵌入。模型采用对比学习训练,嵌入可直接用于多种环境监测任务的零样本设置。实验表明,该嵌入在七项任务中表现优异:林冠高度映射、叶面积覆盖率、土地覆盖分类、树种识别、植物面积指数、作物类型分类及逐像素波形垂直结构映射。结果表明,结合零样本分类器,其性能常优于专用监督模型,尤其在低数据场景下。微调设置下,五项任务达到或超过当前最佳水平。

原文摘要 · Abstract (English)

Significant efforts have been directed towards adapting self-supervised multimodal learning for Earth observation applications. However, most current methods produce coarse patch-sized embeddings, limiting their effectiveness and integration with other modalities like LiDAR. To close this gap, we present DUNIA, an approach to learn pixel-sized embeddings through cross-modal alignment between images and full-waveform LiDAR data. As the model is trained in a contrastive manner, the embeddings can be directly leveraged in the context of a variety of environmental monitoring tasks in a zero-shot setting. In our experiments, we demonstrate the effectiveness of the embeddings for seven such tasks: canopy height mapping, fractional canopy cover, land cover mapping, tree species identification, plant area index, crop type classification, and per-pixel waveform-based vertical structure mapping. The results show that the embeddings, along with zero-shot classifiers, often outperform specialized supervised models, even in low-data regimes. In the fine-tuning setting, we show strong performances near or better than the state-of-the-art on five out of six tasks.

遥感多模态像素嵌入零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。