用空间偏置注意力机制,让模型同时理解多尺度、多模态卫星图像。
Multi-modal, multi-scale representation learning for satellite imagery analysis just needs a good ALiBi
- 引入线性偏置注意力,显式建模不同分辨率图像块间关系。
- 在GEO-Bench上性能提升,优于现有方法。
- 适合遥感分析、多源影像融合等研究者使用。
视觉基础模型已在处理卫星影像以生成下游任务可用表征方面表现出色,但同时处理多空间分辨率和多模态数据仍具挑战。本文提出Scale-ALiBi,一种带有空间编码偏置的线性偏置注意力机制,用于建模不同地面采样距离尺度下图像块之间的关系。我们在一个包含对齐的高/低分辨率光学与低分辨率合成孔径雷达(SAR)卫星影像的数据集上,采用三重对比与重建架构实现Scale-ALiBi,并在GEO-Bench基准上取得性能提升,同时公开了新构建的数据集。
原文摘要 · Abstract (English)
Vision foundation models have been shown to be effective at processing satellite imagery into representations fit for downstream tasks, however, creating models which operate over multiple spatial resolutions and modes is challenging. This paper presents Scale-ALiBi, a linear bias transformer attention mechanism with a spatial encoding bias to relationships between image patches at different ground sample distance scales. We provide an implementation of Scale-ALiBi over a dataset of aligned high- and low-resolution optical and low-resolution SAR satellite imagery data using a triple-contrastive and reconstructive architecture, show an improvement on the GEO-Bench benchmark, and release the newly curated dataset publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。