提出首个面向遥感影像的组合检索基准,助力灾害监测高效找图。
Benchmarking Composed Image Retrieval for Applied Earth Observation

- 构建统一评测框架,测试六种视觉语言模型在遥感图像上的组合检索能力。
- xView2-CIR数据集聚焦灾后变化,验证场景一致性对检索的关键影响。
- 无需训练的方法已成实用基线,适合应急响应与遥感档案探索场景。
遥感组合图像检索(RSCIR)支持使用参考图像与文本修饰符组成的查询,在大规模卫星图像库中进行搜索。尽管该方法为表达精准检索意图提供了灵活接口,但现代组合方法在地球观测(EO)影像中的可迁移性及其对实际工作流的相关性仍缺乏研究。本文通过统一基准和应用导向研究填补这一空白:首先,系统地在PatternCom数据集上,基于六种视觉-语言主干模型评估代表性组合检索方法,分析其在不同主干、组合策略和查询类型下的表现;其次,引入xView2-CIR数据集,用于灾害与损毁监测,其中检索条件依赖于场景身份和目标灾后状态。结果表明,无需训练的组合方法在地球观测检索中表现强劲且具备可扩展性;而以变化为中心的检索面临不同于属性检索的挑战,尤其在于需保持场景身份不变。本研究建立了一个实用的RSCIR基准,将组合检索定位为遥感图像检索、档案探索与变化分析的互补工具。数据集与代码已开源:https://github.com/billpsomas/rscir。
原文摘要 · Abstract (English)
Remote sensing composed image retrieval (RSCIR) enables search in large satellite image archives using composed queries that combine a reference image with a textual modifier. Although RSCIR offers a flexible interface for expressing targeted retrieval intent, the transferability of modern composition methods to Earth observation (EO) imagery and their relevance to operational EO workflows remain underexplored. We address this gap through a unified benchmark and an application-oriented study. First, we systematically adapt and evaluate representative composed image retrieval methods with six vision-language backbones on PatternCom under a standardized protocol, analyzing their behavior across backbones, composition strategies, and query types. Second, we introduce xView2-CIR, a change-centric dataset for disaster and damage monitoring, where retrieval is conditioned on scene identity and a target post-event state. Our results show that training-free composition methods provide strong and scalable baselines for EO retrieval, while change-centric retrieval presents different challenges from attribute-based retrieval, particularly due to the need to preserve scene identity. Overall, this study establishes a practical benchmark for RSCIR and positions composed retrieval as a complementary tool for remote sensing image retrieval, archive exploration, and change analysis. The dataset and code are available at https://github.com/billpsomas/rscir.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。