arXiv:2502.02452cs.CV2025-02中稿 · Transactions on Ma…被引 7

无需训练即可个性化大模型,支持多概念图像视频识别

Personalization Toolkit: Training Free Personalization of Large Vision Language Models

  • 用预训练模型提取特征,结合检索生成与视觉提示实现无训练个性化
  • 在真实场景多概念任务上超越现有训练方法,效果更优
  • 适合需要快速部署个性化应用的开发者和研究者

大视觉语言模型(LVLM)的个性化旨在使模型能够识别特定用户或物体实例,并生成上下文相关的响应。现有方法依赖针对每个对象的耗时训练,难以在实际中部署,当前评估基准也仅限于以物体为中心的单概念测试。本文提出一种全新的无训练个性化方法 extit{Personalization Toolkit} ( extbf{ours})。我们构建了一个全面的、面向真实场景的评估基准,用于严格评测个性化任务的多个方面。 extbf{ours} 利用预训练视觉基础模型提取独特特征,采用检索增强生成(RAG)技术识别视觉输入中的实例,并通过视觉提示策略引导模型输出。该模型无关的视觉工具包可在不进行任何额外训练的前提下,高效灵活地实现图像与视频的多概念个性化。实验表明,其性能达到当前最优水平,优于已有基于训练的方法。

原文摘要 · Abstract (English)

Personalization of Large Vision-Language Models (LVLMs) involves customizing models to recognize specific users or object instances and to generate contextually tailored responses. Existing approaches rely on time-consuming training for each item, making them impractical for real-world deployment, as reflected in current personalization benchmarks limited to object-centric single-concept evaluations. In this paper, we present a novel training-free approach to LVLM personalization called \ours. We introduce a comprehensive, real-world benchmark designed to rigorously evaluate various aspects of the personalization task. \ours leverages pre-trained vision foundation models to extract distinctive features, applies retrieval-augmented generation (RAG) techniques to identify instances within visual inputs, and employs visual prompting strategies to guide model outputs. Our model-agnostic vision toolkit enables efficient and flexible multi-concept personalization across both images and videos, without any additional training. We achieve state-of-the-art results, surpassing existing training-based methods.

大模型个性化无训练视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。