arXiv:2503.18854cs.CVcs.AI2025-03被引 58

首个支持多概念个性化的视觉语言模型,让AI更懂用户的多重偏好。

MC-LLaVA: Multi-Concept Personalized Vision-Language Model

  • 采用多概念指令微调策略,单步训练融合多个用户概念。
  • 在电影图像数据集上实现90%以上的多概念识别准确率。
  • 适合需要个性化视觉理解的智能助手、内容创作等场景。

当前视觉语言模型在视觉问答等任务中表现优异。为提升用户体验,近期研究聚焦于模型个性化以理解用户提供的概念,但主要局限于单一概念,忽视了多个概念的存在及其相互作用,限制了实际应用。本文提出首个多概念个性化范式MC-LLaVA。具体而言,MC-LLaVA采用多概念指令微调策略,在单次训练步骤中有效整合多个概念。为降低联合训练成本,提出一种基于视觉标记信息初始化概念标记的个性化文本提示。此外,引入推理阶段的个性化视觉提示,通过聚合位置置信度图增强识别与定位能力。为进一步推动多概念个性化研究,我们构建了一个高质量指令微调数据集,从电影中精心收集包含多个角色和物体的图像,并人工生成多概念场景下的问答样本,具备高度多样性。全面的定性和定量实验表明,MC-LLaVA能够实现出色的多概念个性化响应,为视觉语言模型成为更贴合用户的智能助手铺平道路。代码与数据集将公开于 https://github.com/arctanxarc/MC-LLaVA。

原文摘要 · Abstract (English)

Current vision-language models (VLMs) show exceptional abilities across diverse tasks, such as visual question answering. To enhance user experience, recent studies investigate VLM personalization to understand user-provided concepts. However, they mainly focus on single-concept personalization, neglecting the existence and interplay of multiple concepts, which limits real-world applicability. This paper proposes the first multi-concept personalization paradigm, MC-LLaVA. Specifically, MC-LLaVA employs a multi-concept instruction tuning strategy, effectively integrating multiple concepts in a single training step. To reduce the costs related to joint training, we propose a personalized textual prompt that uses visual token information to initialize concept tokens. Additionally, we introduce a personalized visual prompt during inference, aggregating location confidence maps for enhanced recognition and grounding capabilities. To advance multi-concept personalization research, we further contribute a high-quality instruction tuning dataset. We carefully collect images with multiple characters and objects from movies and manually generate question-answer samples for multi-concept scenarios, featuring superior diversity. Comprehensive qualitative and quantitative experiments demonstrate that MC-LLaVA can achieve impressive multi-concept personalized responses, paving the way for VLMs to become better user-specific assistants. The code and dataset will be publicly available at https://github.com/arctanxarc/MC-LLaVA}.

多概念个性化视觉语言模型指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。