用自进化框架让视觉语言模型在少量标注下高效学习卫星图像分类
Semi-Supervised Adaptation of Vision-Language Models for Image Classification

- 通过两阶段迭代挖掘未标注数据中的高置信样本
- 在UCM和NWPU数据集上超越现有半监督方法
- 适合遥感图像少样本场景下的模型适配
视觉语言模型如CLIP在自然图像处理中表现出显著潜力,但在卫星图像上性能受限。尽管存在参数高效的适配技术,其效果常因标注样本稀缺而受限。本文提出自进化CLIP(SE-CLIP),一种用于场景分类的半监督框架,采用双阶段流程:先以少量标注样本进行热启动,再通过迭代方式从无标签数据中识别高置信样本。为保持支持集的类别平衡,引入类均衡选择策略,防止模型被易学类别主导。在UCM与NWPU基准上的实验表明,SE-CLIP显著优于现有半监督方法。该框架为在极少人工干预下适配视觉语言模型至遥感领域提供了可行方案。
原文摘要 · Abstract (English)
Vision-language models like CLIP have shown sig- nificant potential in handling natural images, yet their perfor- mance is often limited by the distinct characteristics of satellite imagery. While parameter-efficient adaptation techniques exist, their efficacy is frequently limited by the scarcity of annotated samples. In this letter, we propose Self-Evolutionary CLIP (SE- CLIP), a semi-supervised framework designed for recursive label mining in scene classification. The approach follows a dual-phase pipeline, where an initial warm-up on a few annotated seeds is followed by a recursive discovery phase that iteratively identifies high-confidence samples from unlabeled pools. To maintain the integrity of the evolving support set, we employ a class-balanced selection strategy that prevents the model from being dominated by easily learned categories. Results on the UCM and NWPU benchmarks indicate that SE-CLIP significantly outperforms existing semi-supervised approaches. The framework provides a viable solution for adapting VLMs to the remote sensing domain with minimal human intervention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。