从核方法视角改进视觉语言模型少样本适配,性能达新高
ProKeR: A Kernel Perspective on Few-Shot Adaptation of Large Vision-Language Models
- 将缓存式适配视为局部核方法,揭示其理论机制
- 引入全局正则项,在11个数据集上达到顶尖效果
- 无需微调,适合快速部署于多场景少样本任务
对比图像预训练(CLIP)的广泛应用催生了高效少样本适配技术。其中,无需训练的缓存方法(如Tip-Adapter)因轻量性备受关注。本文从核方法视角重新审视Tip-Adapter,揭示其作为局部适配器的本质,并与经典核学习理论建立联系。基于此,我们提出理论洞见并改进基线方法,强调在局部适配中融入全局信息的重要性。为此,我们提出ProKeR(近端核岭回归),在再生核希尔伯特空间(RKHS)中学习一个近端正则项,以CLIP为基学习器。该方法具有闭式解,在标准少样本适配基准的11个数据集上均取得当前最优表现。
原文摘要 · Abstract (English)
The growing popularity of Contrastive Language-Image Pretraining (CLIP) has led to its widespread application in various visual downstream tasks. To enhance CLIP's effectiveness and versatility, efficient few-shot adaptation techniques have been widely adopted. Among these approaches, training-free methods, particularly caching methods exemplified by Tip-Adapter, have gained attention for their lightweight adaptation without the need for additional fine-tuning. In this paper, we revisit Tip-Adapter from a kernel perspective, showing that caching methods function as local adapters and are connected to a well-established kernel literature. Drawing on this insight, we offer a theoretical understanding of how these methods operate and suggest multiple avenues for enhancing the Tip-Adapter baseline. Notably, our analysis shows the importance of incorporating global information in local adapters. Therefore, we subsequently propose a global method that learns a proximal regularizer in a reproducing kernel Hilbert space (RKHS) using CLIP as a base learner. Our method, which we call ProKeR (Proximal Kernel ridge Regression), has a closed form solution and achieves state-of-the-art performances across 11 datasets in the standard few-shot adaptation benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。