提出多粒度对齐框架,提升遥感图像与文本的精细匹配能力。
GeoAlignCLIP: Enhancing Fine-Grained Vision-Language Alignment in Remote Sensing via Multi-Granular Consistency Learning
- 通过多粒度语义对齐与模态内一致性学习实现细粒度对齐
- 在多个遥感基准上超越现有方法,显著提升细粒度任务性能
- 构建了包含10万样本的精细标注数据集RSFG-100k
视觉-语言预训练模型在连接遥感影像与自然语言方面取得显著进展。然而,现有方法常未能有效整合多粒度视觉与文本信息,主要依赖全局图像-文本对齐,限制了对图像细粒度细节的准确捕捉,进而制约复杂细粒度任务的表现。为此,我们提出GeoAlignCLIP,一种统一框架,通过学习多粒度语义对齐并引入模态内一致性,实现遥感任务中的细粒度对齐,使图像区域与文本概念间实现更精准的视觉-语义对齐。此外,我们构建了RSFG-100k,一个包含场景描述、区域级标注和挑战性难负样本的细粒度遥感数据集,为模型训练提供分层监督。在多个公开遥感基准上的大量实验表明,GeoAlignCLIP在多样任务中持续优于现有遥感专用方法,展现出更强鲁棒性与更高精度的细粒度视觉-语言对齐能力。
原文摘要 · Abstract (English)
Vision-language pretraining models have made significant progress in bridging remote sensing imagery with natural language. However, existing approaches often fail to effectively integrate multi-granular visual and textual information, relying primarily on global image-text alignment. This limitation hinders the model's ability to accurately capture fine-grained details in images, thus restricting its performance in complex, fine-grained tasks. To address this, we propose GeoAlignCLIP, a unified framework that achieves fine-grained alignment in remote sensing tasks by learning multi-granular semantic alignments and incorporating intra-modal consistency, enabling more precise visual-semantic alignment between image regions and text concepts. Additionally, we construct RSFG-100k, a fine-granular remote sensing dataset containing scene descriptions, region-level annotations, and challenging hard-negative samples, providing hierarchical supervision for model training. Extensive experiments conducted on multiple public remote-sensing benchmarks demonstrate that GeoAlignCLIP consistently outperforms existing RS-specific methods across diverse tasks, exhibiting more robust and accurate fine-grained vision-language alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。