arXiv:2511.18396cs.CV2025-11

用弱模型指导强模型,提升CLIP分类性能。

Exploring Weak-to-Strong Generalization for CLIP-based Classification

  • 用类别原型学习法,在弱监督下优化CLIP的分类能力。
  • 在预训练不足时,相比强基线提升3.67%准确率。
  • 适合资源受限场景下提升多模态模型泛化能力。

将大规模商业模型与用户意图对齐对于防止有害输出至关重要。当前方法依赖人工监督,但随着模型复杂度上升变得不切实际。当模型超越人类知识时,提供精准反馈既困难又低效。近期提出的新方案是利用弱模型监督强模型,借助弱模型的评估能力减轻人工负担。已有研究证明该方法在纯语言模型中有效。本研究将此思路拓展至视觉-语言模型,探索基于CLIP的弱到强泛化。我们提出类别原型学习(CPL)方法,通过学习更具代表性的类别原型来增强CLIP的分类能力。实验表明,尽管使用简单损失函数进行弱监督,CPL在目标场景下仍表现稳健,尤其在预训练数据有限时效果显著。大量实验证明该方法有效,相比强基线实现3.67%的性能提升。

原文摘要 · Abstract (English)

Aligning large-scale commercial models with user intent is crucial to preventing harmful outputs. Current methods rely on human supervision but become impractical as model complexity increases. When models surpass human knowledge, providing accurate feedback becomes challenging and inefficient. A novel solution proposed recently is using a weaker model to supervise a stronger model. This concept leverages the ability of weaker models to perform evaluations, thereby reducing the workload on human supervisors. Previous work has shown the effectiveness of weak-to-strong generalization in the context of language-only models. Extending this concept to vision-language models leverages these insights, adapting the proven benefits to a multi-modal context. In our study, we explore weak-to-strong generalization for CLIP-based classification. We propose a method, class prototype learning (CPL), which aims to enhance the classification capabilities of the CLIP model, by learning more representative prototypes for each category. Our findings indicate that, despite using a simple loss function under weak supervision, CPL yields robust improvements in targeted scenarios, particularly when pretraining is limited. Extensive experiments demonstrate that our approach is effective under these settings, achieving a 3.67% improvement over strong baseline methods.

CLIP弱监督多模态分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。