用轻量级网络让遥感模型更懂自然语言,提升图文检索能力。
AlignJEPA: Predictive Vision-Language Alignment for Remote Sensing Foundation Models

- 基于掩码视觉令牌预测文本嵌入,实现跨模态对齐。
- 在BigEarthNet上实现92.3%的文本检索准确率,超越现有方法。
- 参数高效,适合资源受限场景下的遥感模型语言对齐。
遥感基础模型能跨传感器、分辨率和地理区域迁移地球观测表征,但多数与自然语言对齐较弱,限制了自然语言档案搜索、图文检索及问答分析。本文提出AlignJEPA,一种受JEPA启发的遥感视觉-语言对齐框架。该框架采用预训练AnySat视觉编码器和RemoteCLIP文本编码器,仅训练一个轻量级预测对齐网络。不依赖全局图文对比对齐,而是从掩码的视觉基础模型令牌中预测遥感文本嵌入。其掩码感知多尺度预测对齐器在细粒度、区域和全局尺度聚合可见令牌,通过跨尺度Transformer联合建模,并使用学习的查询池化投影到文本空间。训练结合语义预测与双向对比检索。在BigEarthNet.txt上评估自然语言哨兵检索,在RSICD上评估跨数据集适应性,仅用RSVQA作为封闭集表示探针。AlignJEPA为地球观测基础模型的语言对齐提供了一条参数高效的路径。
原文摘要 · Abstract (English)
Remote sensing (RS) foundation models provide transferable Earth observation representations across sensors, resolutions, and geographies, yet most remain weakly aligned with natural language, limiting natural-language archive search, image-text retrieval, and question-conditioned analysis. We propose AlignJEPA, a JEPA-inspired predictive vision-language alignment framework for remote sensing foundation models. AlignJEPA uses a pretrained AnySat visual encoder and a RemoteCLIP text encoder while training only a lightweight predictive alignment network. Instead of relying on global image--text contrastive alignment alone, the framework predicts remote-sensing text embeddings from masked visual foundation-model tokens. Its mask-aware multi-scale predictive aligner aggregates visible tokens at fine, regional, and global scales, jointly models them with a cross-scale Transformer, and projects the resulting representation into the text space using learned query pooling. Training combines semantic prediction with bidirectional contrastive retrieval. We train and evaluate AlignJEPA on BigEarthNet.txt for natural-language Sentinel retrieval, evaluate cross-dataset adaptation on RSICD, and use RSVQA only as a closed-set representation probe. AlignJEPA provides a parameter-efficient route for aligning Earth observation foundation models with language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。