arXiv:2606.20702cs.CVcs.AI2026-06

用提示工程提升遥感零样本分类,发现简洁描述比大模型生成更稳定

Beyond Templates: Revisiting Zero-Shot Remote Sensing through Meta-Prompting

论文配图:Beyond Templates: Revisiting Zero-Shot Remote Sensing through Meta-Prompting
图 1 · 摘自论文原文
  • 通过元提示设计文本描述,优化遥感图像的零样本识别
  • 轻量级查询嵌入校准使分类与检索性能显著提升
  • 适合关注遥感视觉语言模型应用的研究者和工程师

视觉语言模型(VLMs)在零样本地球观测(EO)任务中引发广泛关注,尤其在适配遥感数据的模型上表现更优。我们基于元提示视觉识别框架(MPVR),在17种VLM变体和12个遥感数据集上系统评估该设置,发现零样本性能对文本设计高度敏感,包括用于引导大语言模型生成类别描述的元提示及其生成内容本身。尽管大语言模型生成的描述语义更丰富,但其在文本嵌入空间中可能引入噪声,降低下游任务鲁棒性。通过白化后的CLIP特征空间中的文本对数似然分析,我们验证了这一现象。进一步研究查询嵌入校准,发现轻量级校准能持续提升零样本分类与检索性能。结果揭示了语义丰富性与鲁棒性间的权衡,并指出嵌入校准是一种简单有效的改进方法。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have sparked growing interest in zero-shot Earth Observation (EO) downstream tasks, with further gains enabled by remote-sensing-adapted models. We examine this setting across 17 VLM variants and 12 remote sensing (RS) datasets under Meta-Prompting for Visual Recognition (MPVR), and show that zero-shot performance remains highly sensitive to textual design choices, from the meta-prompts used to guide the LLM in generating class descriptions to the descriptions themselves. We explore why semantically rich LLM-generated class descriptions do not translate into consistent gains over simple domain-adapted CLIP-style descriptions. While LLM descriptions are more semantically expressive, they can also introduce noise in the text embedding space, reducing robustness in downstream tasks. We support this observation through a text log-likelihood analysis in the whitened CLIP feature space, comparing LLM-generated and template-based descriptions. Building on this finding, we study query embedding calibration and show that lightweight calibration of the query space consistently yields strong improvements in zero-shot classification and retrieval. Overall, our results provide practical insight into the trade-off between semantic richness and robustness, and identify embedding calibration as a simple and effective tool for improving zero-shot remote sensing performance.

零样本学习遥感图像提示工程视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。