让大模型对话更个性化,能随时添加新概念且无额外开销
Personalized Large Vision-Language Models
- 用预训练视觉编码器对齐指代概念与图像特征
- 对话中可实时识别并引用新概念,无需重训练
- 轻量级设计,计算和参数开销几乎为零
个性化模型在图像生成领域备受关注,但在大型视觉语言模型(LVLMs)中仍研究不足。与通用形式不同,个性化模型能处理包含具体指代的概念(如‘Mike和Susan在说话’),使对话更可定制、更具指代友好性。此外,PLVM可在对话过程中持续添加新概念,且不增加额外成本,显著提升实用性。该方法提出Aligner,一个预训练的视觉编码器,用于将指代概念与查询图像对齐。在对话中,它从参考图像中提取对应概念的特征,并在查询图像中识别这些概念,实现个性化。我们注意到,Aligner在整个框架中的计算开销和参数量可忽略不计。通过全面的定性和定量分析,验证了PLVM的有效性和优越性。
原文摘要 · Abstract (English)
The personalization model has gained significant attention in image generation yet remains underexplored for large vision-language models (LVLMs). Beyond generic ones, with personalization, LVLMs handle interactive dialogues using referential concepts (e.g., ``Mike and Susan are talking.'') instead of the generic form (e.g., ``a boy and a girl are talking.''), making the conversation more customizable and referentially friendly. In addition, PLVM is equipped to continuously add new concepts during a dialogue without incurring additional costs, which significantly enhances the practicality. PLVM proposes Aligner, a pre-trained visual encoder to align referential concepts with the queried images. During the dialogues, it extracts features of reference images with these corresponding concepts and recognizes them in the queried image, enabling personalization. We note that the computational cost and parameter count of the Aligner are negligible within the entire framework. With comprehensive qualitative and quantitative analyses, we reveal the effectiveness and superiority of PLVM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。