arXiv:2504.20998cs.CVcs.AI2025-04CVPR被引 21

让AI学会个性化生成用户指定概念的图文内容

YoChameleon: Personalized Vision and Language Generation

论文配图:YoChameleon: Personalized Vision and Language Generation
图 1 · 摘自论文原文
  • 用3-5张图+软提示微调,让模型记住特定概念
  • 能回答关于该概念的问题并生成细节丰富的图像
  • 适合需要个性化图文创作的用户或场景

大型多模态模型(如 GPT-4、Gemini、Chameleon)已服务数百万用户,但仍是通用模型,缺乏对特定用户概念的个性化知识。以往研究聚焦文本个性化,却未解决图像生成等新模态的适配问题。本文提出 Yo'Chameleon,首个针对大模型多模态个性化的探索。仅需3-5张目标概念图像,通过软提示微调嵌入用户专属信息,实现:(i) 回答与该概念相关的问题;(ii) 在新场景中生成包含像素级细节的图像。模型采用自提示优化机制以平衡多模态性能,并引入“软正例”图像生成方法,在少样本条件下显著提升图像质量。

原文摘要 · Abstract (English)

Large Multimodal Models (e.g., GPT-4, Gemini, Chameleon) have evolved into powerful tools with millions of users. However, they remain generic models and lack personalized knowledge of specific user concepts. Previous work has explored personalization for text generation, yet it remains unclear how these methods can be adapted to new modalities, such as image generation. In this paper, we introduce Yo'Chameleon, the first attempt to study personalization for large multimodal models. Given 3-5 images of a particular concept, Yo'Chameleon leverages soft-prompt tuning to embed subject-specific information to (i) answer questions about the subject and (ii) recreate pixel-level details to produce images of the subject in new contexts. Yo'Chameleon is trained with (i) a self-prompting optimization mechanism to balance performance across multiple modalities, and (ii) a ``soft-positive" image generation approach to enhance image quality in a few-shot setting.

多模态个性化图像生成软提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。