SelfPrompt通过自适应伪标签提升视觉语言模型在少样本下的鲁棒性。
SelfPrompt: Confidence-Aware Semi-Supervised Tuning for Robust Vision-Language Model Adaptation
- 基于聚类的伪标签方法提高标注准确性
- 在13个数据集上平均提升6.23%(标准半监督)
- 适合资源有限但需高泛化能力的场景
我们提出SelfPrompt,一种用于视觉语言模型(VLMs)在半监督学习设置下的新型提示调优方法。现有方法在半监督场景中受限于模型置信度校准不足导致的伪标签噪声累积。SelfPrompt通过引入聚类引导的伪标签机制提升伪标签准确率,并设计置信度感知的半监督学习模块,结合监督与弱监督学习以最大化未标记数据利用率。此外,我们在主动半监督学习框架下验证方法,提出一种弱监督采样策略,可高效选择多样且有代表性的标注集,无缝集成至现有方法中提升性能。在13个数据集上进行广泛评估,标准半监督下平均提升6.23%,主动半监督下提升6.25%,基线到新类别泛化提升4.9%(2次提示设置)。在单次提示设置中,平均提升达11.78%,展现出优异泛化能力。
原文摘要 · Abstract (English)
We present SelfPrompt, a novel prompt-tuning approach for vision-language models (VLMs) in a semi-supervised learning setup. Existing methods for tuning VLMs in semi-supervised setups struggle with the negative impact of the miscalibrated VLMs on pseudo-labelling, and the accumulation of noisy pseudo-labels. SelfPrompt addresses these challenges by introducing a cluster-guided pseudo-labelling method that improves pseudo-label accuracy, and a confidence-aware semi-supervised learning module that maximizes the utilization of unlabelled data by combining supervised learning and weakly-supervised learning. Additionally, we investigate our method in an active semi-supervised learning setup, where the labelled set is strategically selected to ensure the best utilization of a limited labelling budget. To this end, we propose a weakly-supervised sampling technique that selects a diverse and representative labelled set, which can be seamlessly integrated into existing methods to enhance their performance. We conduct extensive evaluations across 13 datasets, significantly surpassing state-of-the-art performances with average improvements of 6.23% in standard semi-supervised learning, 6.25% in active semi-supervised learning, and 4.9% in base-to-novel generalization, using a 2-shot setup. Furthermore, SelfPrompt shows excellent generalization in single-shot settings, achieving an average improvement of 11.78%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。