用通用属性描述缓解视觉语言模型持续学习中的遗忘问题
DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual Recognition
- 通过通用属性描述建立视觉-属性-类别三元关联
- 在多个数据集上实现95%以上识别准确率保持
- 适合需要长期更新的视觉识别系统使用
视觉语言模型(VLM)的持续学习旨在利用跨模态预训练知识,增量适应不断扩展的下游任务和数据集,同时应对知识遗忘挑战。现有研究多聚焦于将视觉特征与特定类别文本连接,忽视了通用知识与专有知识之间的潜在关系。我们发现,强制模型优化不恰当的视觉-文本匹配会加剧VLM识别能力的遗忘。为此,我们提出DesCLIP,利用通用属性(GA)描述引导对特定类别物体的理解,使VLM建立稳健的视觉-属性-类别三元关联,而非仅依赖视觉-类别连接。具体地,引入语言助手通过合理提示生成具体的GA描述候选;设计基于锚点的嵌入过滤器获取高度相关的GA描述嵌入,作为视觉-文本实例匹配的配对文本嵌入,从而调整视觉编码器;相应地,类别文本嵌入逐步校准以对齐这些共享的GA描述嵌入。大量实验表明,所提方法在多项指标上优于现有持续学习方法,充分验证其在VLM-based识别中的先进性与有效性。
原文摘要 · Abstract (English)
Continual learning of vision-language models (VLMs) focuses on leveraging cross-modal pretrained knowledge to incrementally adapt to expanding downstream tasks and datasets, while tackling the challenge of knowledge forgetting. Existing research often focuses on connecting visual features with specific class text in downstream tasks, overlooking the latent relationships between general and specialized knowledge. Our findings reveal that forcing models to optimize inappropriate visual-text matches exacerbates forgetting of VLM's recognition ability. To tackle this issue, we propose DesCLIP, which leverages general attribute (GA) descriptions to guide the understanding of specific class objects, enabling VLMs to establish robust vision-GA-class trilateral associations rather than relying solely on vision-class connections. Specifically, we introduce a language assistant to generate concrete GA description candidates via proper request prompts. Then, an anchor-based embedding filter is designed to obtain highly relevant GA description embeddings, which are leveraged as the paired text embeddings for visual-textual instance matching, thereby tuning the visual encoder. Correspondingly, the class text embeddings are gradually calibrated to align with these shared GA description embeddings. Extensive experiments demonstrate the advancements and efficacy of our proposed method, with comprehensive empirical evaluations highlighting its superior performance in VLM-based recognition compared to existing continual learning methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。