arXiv:2511.20056cs.CL2025-11被引 4

让个性化视觉语言模型实时学习新概念,无需训练。

Online-PVLM: Advancing Personalized VLMs with Online Concept Learning

  • 用双曲空间表示实现测试时零训练生成概念嵌入
  • 在1292个概念上达到领先性能,支持大规模快速检索
  • 适合需要动态适应用户特定图像的智能系统

个性化视觉语言模型(VLMs)在识别用户特定概念(如用户自行车)方面表现优异。现有方法需为每个新概念单独学习嵌入,无法在测试时实时适应,尤其在大规模场景下难以高效检索概念嵌入。为此,我们提出Online-PVLM,通过双曲表示实现测试时的在线概念学习,无需训练即可生成概念嵌入,使个性化VLM兼具可扩展性与高效性。同时,我们构建了OP-Eval,一个包含1,292个概念和超过3万条高质量样本的大规模基准,涵盖多种问题类型,用于在真实场景中严格评估在线概念学习。大量实验表明,该框架性能达到当前最优。源代码与数据集将公开。

原文摘要 · Abstract (English)

Personalized Visual Language Models (VLMs) are gaining increasing attention for their formidable ability in user-specific concepts aligned interactions (e.g., identifying a user's bike). Existing methods typically require the learning of separate embeddings for each new concept, which fails to support real-time adaptation during testing. This limitation becomes particularly pronounced in large-scale scenarios, where efficient retrieval of concept embeddings is not achievable. To alleviate this gap, we propose Online-PVLM, a framework for online concept learning by leveraging hyperbolic representations. Our approach makes a train-free paradigm for concept embeddings generation at test time, making the use of personalized VLMs both scalable and efficient. In addition, we develop OP-Eval, a comprehensive and large-scale benchmark comprising 1,292 concepts and over 30K high-quality instances with diverse question types, designed to rigorously assess online concept learning in realistic scenarios. Extensive experiments demonstrate the state-of-the-art performance of our proposed framework. Our source code and dataset will be made available.

个性化模型在线学习双曲表示视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。