arXiv:2510.24321cs.CVcs.AI2025-10被引 1

用提示学习提升遥感图像少样本分类性能,解决标注数据少的难题。

Few-Shot Remote Sensing Image Scene Classification with CLIP and Prompt Learning

  • 采用提示学习轻量适配预训练视觉语言模型,应对遥感领域标签稀缺问题。
  • 在多个遥感数据集上,提示学习方法显著优于零样本和线性探测基线。
  • 自调节约束提示法表现最佳,跨域泛化能力强,适合实际遥感应用。

遥感应用日益依赖深度学习进行场景分类,但其性能常受限于标注数据稀缺及跨地理与传感器域标注成本高。尽管如CLIP等视觉语言模型通过对齐视觉与文本模态实现了可迁移表征,但其直接应用于遥感仍因领域差异大且需任务特定语义适配而效果不佳。为此,我们系统探索提示学习作为少样本遥感图像场景分类的轻量高效适配策略。评估了上下文优化、条件上下文优化、多模态提示学习及自调节约束提示等方法,涵盖从静态优化到条件提示、联合视-语适应及语义正则化等多种设计思路。在多个基准遥感数据集(含跨数据集泛化测试)上进行广泛实验,结果表明提示学习在少样本场景下持续优于零样本CLIP与基于冻结特征的线性探测基线。其中,自调节约束提示法展现出最稳健的跨域性能。研究证实提示学习是弥合卫星与航空影像领域差距的可扩展高效方案,为该领域后续研究奠定基础。

原文摘要 · Abstract (English)

Remote sensing applications increasingly rely on deep learning for scene classification. However, their performance is often constrained by the scarcity of labeled data and the high cost of annotation across diverse geographic and sensor domains. While recent vision-language models like CLIP have shown promise by learning transferable representations at scale by aligning visual and textual modalities, their direct application to remote sensing remains suboptimal due to significant domain gaps and the need for task-specific semantic adaptation. To address this critical challenge, we systematically explore prompt learning as a lightweight and efficient adaptation strategy for few-shot remote sensing image scene classification. We evaluate several representative methods, including Context Optimization, Conditional Context Optimization, Multi-modal Prompt Learning, and Prompting with Self-Regulating Constraints. These approaches reflect complementary design philosophies: from static context optimization to conditional prompts for enhanced generalization, multi-modal prompts for joint vision-language adaptation, and semantically regularized prompts for stable learning without forgetting. We benchmark these prompt-learning methods against two standard baselines: zero-shot CLIP with hand-crafted prompts and a linear probe trained on frozen CLIP features. Through extensive experiments on multiple benchmark remote sensing datasets, including cross-dataset generalization tests, we demonstrate that prompt learning consistently outperforms both baselines in few-shot scenarios. Notably, Prompting with Self-Regulating Constraints achieves the most robust cross-domain performance. Our findings underscore prompt learning as a scalable and efficient solution for bridging the domain gap in satellite and aerial imagery, providing a strong foundation for future research in this field.

遥感图像少样本学习提示学习CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。