arXiv:2602.13961cs.CVastro-ph.IM2026-02

构建火星尺度地理检索基准,评估视觉语言模型发现地表特征能力

MarsRetrieval: Benchmarking Vision-Language Models for Planetary-Scale Geospatial Retrieval on Mars

  • 设计三类多尺度火星地理检索任务,覆盖图像文本匹配、地貌识别与全球定位
  • 强基底模型在特定地貌区分上表现不佳,表明领域适配至关重要
  • 适用于行星科学、遥感分析与多模态模型评估的研究者

数据驱动方法如深度学习正快速推动行星科学进步,尤其在火星探测领域。尽管已有进展,多数现有基准仍局限于封闭集监督视觉任务,无法支持文本引导的地理空间发现。我们提出 MarsRetrieval,一个用于评估视觉语言模型在火星地理发现中表现的检索基准。该基准包含三项任务:(1) 图像-文本配对检索,(2) 地貌检索,(3) 全球地理定位,涵盖多个空间尺度和多样化的地质成因。我们提出统一的以检索为中心的评估协议,用于测试对比双塔编码器与生成式视觉语言模型等多模态嵌入架构。评估显示,MarsRetrieval 具有挑战性:即使强大的基础模型也难以捕捉特定地貌差异。我们进一步表明,领域特定微调对行星场景下的可泛化地理发现至关重要。代码已开源:https://github.com/ml-stat-Sustech/MarsRetrieval

原文摘要 · Abstract (English)

Data-driven approaches like deep learning are rapidly advancing planetary science, particularly in Mars exploration. Despite recent progress, most existing benchmarks remain confined to closed-set supervised visual tasks and do not support text-guided retrieval for geospatial discovery. We introduce MarsRetrieval, a retrieval benchmark for evaluating vision-language models for Martian geospatial discovery. MarsRetrieval includes three tasks: (1) paired image-text retrieval, (2) landform retrieval, and (3) global geo-localization, covering multiple spatial scales and diverse geomorphic origins. We propose a unified retrieval-centric protocol to benchmark multimodal embedding architectures, including contrastive dual-tower encoders and generative vision-language models. Our evaluation shows MarsRetrieval is challenging: even strong foundation models often fail to capture domain-specific geomorphic distinctions. We further show that domain-specific fine-tuning is critical for generalizable geospatial discovery in planetary settings. Our code is available at https://github.com/ml-stat-Sustech/MarsRetrieval

视觉语言火星探索地理检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。