让个性化视觉语言模型实时学习新概念,无需训练。
Online-PVLM: Advancing Personalized VLMs with Online Concept Learning
- 用双曲空间表示实现测试时零训练生成概念嵌入
- 在1292个概念上达到领先性能,支持大规模快速检索
- 适合需要动态适应用户特定图像的智能系统
个性化视觉语言模型(VLMs)在识别用户特定概念(如用户自行车)方面表现优异。现有方法需为每个新概念单独学习嵌入,无法在测试时实时适应,尤其在大规模场景下难以高效检索概念嵌入。为此,我们提出Online-PVLM,通过双曲表示实现测试时的在线概念学习,无需训练即可生成概念嵌入,使个性化VLM兼具可扩展性与高效性。同时,我们构建了OP-Eval,一个包含1,292个概念和超过3万条高质量样本的大规模基准,涵盖多种问题类型,用于在真实场景中严格评估在线概念学习。大量实验表明,该框架性能达到当前最优。源代码与数据集将公开。
原文摘要 · Abstract (English)
Personalized Visual Language Models (VLMs) are gaining increasing attention for their formidable ability in user-specific concepts aligned interactions (e.g., identifying a user's bike). Existing methods typically require the learning of separate embeddings for each new concept, which fails to support real-time adaptation during testing. This limitation becomes particularly pronounced in large-scale scenarios, where efficient retrieval of concept embeddings is not achievable. To alleviate this gap, we propose Online-PVLM, a framework for online concept learning by leveraging hyperbolic representations. Our approach makes a train-free paradigm for concept embeddings generation at test time, making the use of personalized VLMs both scalable and efficient. In addition, we develop OP-Eval, a comprehensive and large-scale benchmark comprising 1,292 concepts and over 30K high-quality instances with diverse question types, designed to rigorously assess online concept learning in realistic scenarios. Extensive experiments demonstrate the state-of-the-art performance of our proposed framework. Our source code and dataset will be made available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。