arXiv:2603.09771cs.CVcs.AI2026-03中稿 · the IEEE/CVF Confe…

用视觉模型内部注意力提取记忆令牌,实现轻量个性化

Ego: Embedding-Guided Personalization of Vision-Language Models

  • 利用模型自身注意力机制提取代表特定概念的视觉令牌
  • 在单/多概念及视频场景中均显著提升个性化效果
  • 无需额外训练,适合快速部署的智能助手应用

支持日常生活的AI助手正日益可行,得益于多模态语言模型的快速发展。核心挑战在于克服模型的通用性,实现个性化体验。现有个性化方法通常依赖额外训练阶段,限制了泛化能力与可扩展性;或采用带有外部预训练模块的工程化流水线,影响部署效率。本文提出一种高效个性化方法,利用模型内在捕捉个性化概念的能力。具体而言,通过模型内部注意力机制提取主要表征目标概念的视觉令牌,这些令牌作为该概念的记忆,使模型在测试图像中再次遇到时能回忆并描述它。我们在单概念、多概念及视频个性化等多种设置下,对本方法与最先进方法进行了全面统一评估,结果表明在极小个性化开销下实现了显著性能提升。

原文摘要 · Abstract (English)

AI assistants that support humans in daily life are becoming increasingly feasible, driven by the rapid advancements in multimodal language models. A key challenge lies in overcoming the generic nature of these models to deliver personalized experiences. Existing approaches to personalizing large vision language models often rely on additional training stages, which limit generality and scalability, or on engineered pipelines with external pre-trained modules, which hinder deployment efficiency. In this work, we propose an efficient personalization method that leverages the model's inherent ability to capture personalized concepts. Specifically, we extract visual tokens that predominantly represent the target concept by utilizing the model's internal attention mechanisms. These tokens serve as a memory of that specific concept, enabling the model to recall and describe it when it appears in test images. We conduct a comprehensive and unified evaluation of our approach and SOTA methods across various personalization settings including single-concept, multi-concept, and video personalization, demonstrating strong performance gains with minimal personalization overhead.

视觉语言模型个性化注意力机制轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。