让视觉语言模型学会基于用户个人经历理解图片
Contextualized Visual Personalization in Vision-Language Models
- 将个性化图文生成设为核心任务,通过强化学习优化模型
- 在真实用户上下文中显著提升图像描述准确率
- 适合需要个性化的智能助手、社交应用开发者
尽管视觉语言模型(VLMs)取得进展,但现有方法难以根据用户的特定经历生成个性化回应,因其缺乏将视觉输入与用户积累的图文上下文关联的能力。本文首次将此挑战定义为上下文感知的视觉个性化,要求模型在解读新图像时能识别并检索用户的个性化视觉经验。为此,我们提出CoViP框架,将个性化图像描述作为核心任务,并通过基于强化学习的后训练和标注增强生成来提升该能力。我们进一步设计诊断评估,明确排除仅依赖文本捷径的可能,验证模型是否真正利用视觉上下文。大量实验表明,现有开源及专有VLMs存在明显局限,而CoViP不仅提升了个性化图像描述性能,还在下游个性化任务中实现整体提升。结果表明,CoViP是实现鲁棒且通用的上下文感知视觉个性化的关键阶段。
原文摘要 · Abstract (English)
Despite recent progress in vision-language models (VLMs), existing approaches often fail to generate personalized responses based on the user's specific experiences, as they lack the ability to associate visual inputs with a user's accumulated visual-textual context. We newly formalize this challenge as contextualized visual personalization, which requires the visual recognition and textual retrieval of personalized visual experiences by VLMs when interpreting new images. To address this issue, we propose CoViP, a unified framework that treats personalized image captioning as a core task for contextualized visual personalization and improves this capability through reinforcement-learning-based post-training and caption-augmented generation. We further introduce diagnostic evaluations that explicitly rule out textual shortcut solutions and verify whether VLMs truly leverage visual context. Extensive experiments demonstrate that existing open-source and proprietary VLMs exhibit substantial limitations, while CoViP not only improves personalized image captioning but also yields holistic gains across downstream personalization tasks. These results highlight CoViP as a crucial stage for enabling robust and generalizable contextualized visual personalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。