arXiv:2511.17442cs.CVcs.AI2025-11

自动选最适合的遥感大模型,还能解释理由。

REMSA: Foundation Model Selection for Remote Sensing via a Constraint-Aware Agent

  • 用结构化数据库+智能代理,根据自然语言查询选模型
  • 在100个真实任务上测试,3000次评估中表现最优
  • 适合遥感从业者快速匹配模型,无需懂技术细节

基础模型(FMs)正被广泛应用于遥感(RS)流程中,涵盖单模态视觉编码器与多模态架构,适用于图像分类、变化检测和视觉问答等任务。然而,因文档分散、格式不一及部署约束复杂,为特定任务选择最合适的遥感基础模型(RSFM)仍具挑战。为此,我们首先构建了首个结构化、基于模式的资源——RSFM数据库(RS-FMD),涵盖超过160个在多种数据模态上训练的RSFM,覆盖不同空间、光谱与时间分辨率,并考虑多种学习范式。基于此,我们提出REMSA,一个约束感知智能体,可基于自然语言查询实现自动化RSFM选择。REMSA结合结构化元数据检索与任务驱动决策流程:解析用户输入、澄清缺失约束、通过上下文学习排序模型,并提供透明解释。系统支持多种遥感任务与数据模态,实现个性化、可复现且高效的模型选择。为评估,我们构建了100个专家验证的遥感查询场景,每项查询在4个系统与3个LLM骨干上评估,顶级3个模型由领域专家手动评分,共产生3000个专家评分的任务-系统-模型配置,采用新型以专家为中心的评估协议。REMSA显著优于多个基线,展现其在真实决策中的实用价值。REMSA仅使用公开开源的RSFM元数据,不接触私有或敏感数据。

原文摘要 · Abstract (English)

Foundation Models (FMs) are increasingly integrated into remote sensing (RS) pipelines. These models include unimodal vision encoders and multimodal architectures. FMs are adapted to diverse perception tasks, such as image classification, change detection, and visual question answering. However, selecting the most suitable remote sensing foundation model (RSFM) for a specific task remains challenging due to scattered documentation, heterogeneous formats, and complex deployment constraints. To address this, we first introduce the RSFM Database (RS-FMD), the first structured and schema-guided resource covering over 160 RSFMs trained on various data modalities, spanning different spatial, spectral, and temporal resolutions, considering different learning paradigms. Built upon RS-FMD, we further present REMSA, a constraint-aware agent that enables automated RSFM selection from natural language queries. REMSA combines structured FM metadata retrieval with a task-driven decision workflow. In detail, it interprets user input, clarifies missing constraints, ranks models via in-context learning, and provides transparent justifications. Our system supports various RS tasks and data modalities, enabling personalized, reproducible, and efficient FM selection. To evaluate REMSA, we construct a benchmark of 100 expert-verified RS query scenarios. Each query is evaluated across 4 systems and 3 LLM backbones, with the top-3 selected models manually assessed by domain experts. This results in 3,000 expert-scored task--system--model configurations under our novel expert-centered evaluation protocol. REMSA outperforms multiple baselines, showing its practical utility in real decision-making applications. REMSA operates entirely on publicly available metadata of open source RSFMs, without accessing private or sensitive data.

遥感大模型智能选型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。