用动态对齐提升视觉语言模型细粒度分类效果
Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score
- 通过类描述锚点动态对齐图像与文本特征
- 在13个数据集上比现有方法提升2.78%准确率
- 适合需要无监督细粒度分类的场景
视觉语言模型(如CLIP)通过对比预训练实现零样本学习,但在细粒度分类的无监督适配中,现有方法或依赖固定对齐分数难以捕捉细微类别差异,或采用计算成本高的伪标签策略,限制可扩展性。本文提出细粒度对齐与交互优化(FAIR),通过一组类描述锚点(CDA)动态对齐局部图像特征与文本嵌入,定义可学习的对齐分数(LAS),作为自适应分类器促进跨模态交互,提升自训练性能。同时引入自训练加权机制,缓解类间模糊性。FAIR在13个细粒度数据集上相比最先进方法平均提升2.78%。
原文摘要 · Abstract (English)
Vision-language models (VLMs) like CLIP excel in zero-shot learning by aligning image and text representations through contrastive pretraining. Existing approaches to unsupervised adaptation (UA) for fine-grained classification with VLMs either rely on fixed alignment scores that cannot capture evolving, subtle class distinctions or use computationally expensive pseudo-labeling strategies that limit scalability. In contrast, we show that modeling fine-grained cross-modal interactions during adaptation produces more accurate, class-discriminative pseudo-labels and substantially improves performance over state-of-the-art (SOTA) methods. We introduce Fine-grained Alignment and Interaction Refinement (FAIR), an innovative approach that dynamically aligns localized image features with descriptive language embeddings through a set of Class Description Anchors (CDA). This enables the definition of a Learned Alignment Score (LAS), which incorporates CDA as an adaptive classifier, facilitating cross-modal interactions to improve self-training in unsupervised adaptation. Furthermore, we propose a self-training weighting mechanism designed to refine pseudo-labels in the presence of inter-class ambiguities. Our approach, FAIR, delivers a substantial performance boost in fine-grained unsupervised adaptation, achieving a notable overall gain of 2.78% across 13 fine-grained datasets compared to SOTA methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。