arXiv:2512.09441cs.CVcs.AI2025-12

用轻量投影器校准特征,解决视觉语言模型增量学习中的类别混淆问题。

Representation Calibration and Uncertainty Guidance for Class-Incremental Learning based on Vision Language Model

  • 通过任务专用适配器和轻量投影混合策略校准特征表示
  • 在多个数据集上显著降低旧类遗忘,提升新类识别准确率
  • 适合需要持续学习且资源受限的图像分类场景

类别增量学习要求模型在持续学习新类别知识的同时,尽可能保留对旧类别的记忆。当前基于视觉语言模型(VLMs)的方法仍存在跨任务类别区分度不足的问题。本文提出一种新的基于VLM的增量学习框架:在预训练且冻结的图像编码器中添加任务特定适配器以学习新知识;采用一种基于轻量级投影器混合的跨任务表征校准策略,帮助在统一特征空间中更好分离已学习的所有类别,缓解任务间的类别混淆;同时设计了一种基于预测不确定性的推理策略,更准确地选择用于分类的图像特征。在多个数据集上的大量实验表明,该方法在多种设置下均优于现有方法。

原文摘要 · Abstract (English)

Class-incremental learning requires a learning system to continually learn knowledge of new classes and meanwhile try to preserve previously learned knowledge of old classes. As current state-of-the-art methods based on Vision-Language Models (VLMs) still suffer from the issue of differentiating classes across learning tasks. Here a novel VLM-based continual learning framework for image classification is proposed. In this framework, task-specific adapters are added to the pre-trained and frozen image encoder to learn new knowledge, and a novel cross-task representation calibration strategy based on a mixture of light-weight projectors is used to help better separate all learned classes in a unified feature space, alleviating class confusion across tasks. In addition, a novel inference strategy guided by prediction uncertainty is developed to more accurately select the most appropriate image feature for class prediction. Extensive experiments on multiple datasets under various settings demonstrate the superior performance of our method compared to existing ones.

增量学习视觉语言模型特征校准不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。