动态提示与表征学习,让模型持续识别新类别。
Open-World Dynamic Prompt and Continual Visual Representation Learning

- 用动态提示替代静态提示池,实时生成适配新类别的提示。
- 联合优化提示生成与表征学习,提升对未知类别的识别能力。
- 适合需要持续学习新视觉类别的实际应用场景。
开放世界本质上是动态的,其概念和分布不断演化。在这一动态开放世界环境中进行持续学习(CL)面临巨大挑战,即如何有效泛化到测试时出现的新类别。为此,我们提出一种针对开放世界视觉表征学习的新实用持续学习设置:后续数据流系统性引入与先前训练阶段类别不重叠的新类别,同时又与未见测试类别保持区分。针对此挑战,我们提出动态提示与表征学习器(DPaRL),一种简单而有效的基于提示的持续学习(PCL)方法。与以往仅依赖静态提示池的PCL方法不同,我们的DPaRL在推理时学习生成动态提示。此外,DPaRL在每个训练阶段联合学习动态提示生成与判别性表征,而此前的PCL方法仅在过程中优化提示学习。实验结果表明,该方法在公认的开放世界图像检索基准上,平均提升了4.7%的Recall@1性能,优于现有最先进方法。
原文摘要 · Abstract (English)
The open world is inherently dynamic, characterized by ever-evolving concepts and distributions. Continual learning (CL) in this dynamic open-world environment presents a significant challenge in effectively generalizing to unseen test-time classes. To address this challenge, we introduce a new practical CL setting tailored for open-world visual representation learning. In this setting, subsequent data streams systematically introduce novel classes that are disjoint from those seen in previous training phases, while also remaining distinct from the unseen test classes. In response, we present Dynamic Prompt and Representation Learner (DPaRL), a simple yet effective Prompt-based CL (PCL) method. Our DPaRL learns to generate dynamic prompts for inference, as opposed to relying on a static prompt pool in previous PCL methods. In addition, DPaRL jointly learns dynamic prompt generation and discriminative representation at each training stage whereas prior PCL methods only refine the prompt learning throughout the process. Our experimental results demonstrate the superiority of our approach, surpassing state-of-the-art methods on well-established open-world image retrieval benchmarks by an average of 4.7% improvement in Recall@1 performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。