无需用户数据,用文字描述个性化生成图像。
ZIPP:Zero-shot Image Personalization from Personas

- 用大模型将提示词改写为用户画像视角,实现零样本个性化生成。
- 在1.5千用户、4万张图的基准上提升13%-20%,优于微调方法。
- 适合冷启动场景,显著降低群体偏差,提升真实用户偏好匹配度。
文本到图像扩散模型在开放创作场景中广泛应用,但输出缺乏个性,偏向整体美学而非个体偏好。人类审美多元:有人偏爱低饱和怀旧肖像,有人钟情明亮街头摄影或梦幻电影感。现有方法需密集交互历史或用户微调,难以应对冷启动问题,且将上下文依赖偏好简化为静态表征。本文提出零样本图像个性化(ZIPP),仅通过自然语言描述的用户画像(如身份与审美特征)即可引导生成,无需用户数据或权重更新。ZIPP利用大模型将提示词从特定画像视角重写,驱动扩散模型生成个性化图像。为大规模挖掘画像,我们基于2200万用户的Reddit互动图谱训练了一个归纳式图注意力网络,采用双对比目标对齐图结构与视觉行为,并通过多模态大模型将学习表征转化为自然语言画像。我们构建了首个零样本个性化基准ZIPBench,包含1500名用户、图挖掘画像和4万张生成图像。在四个基准和14个大模型(覆盖五个模型家族)上,画像条件化带来13%-20%一致提升,前沿模型收益最大。少样本设置下,ZIPP性能媲美甚至超越在100+样本/用户上微调的基线。其偏好分布差异最小(CMMD 0.16 vs. 0.55),IPF归一化的人口统计评估显示显著减少现有方法中的子群体偏差。人工评估确认其胜过通用生成79%,优于所有微调基线58%-65%。
原文摘要 · Abstract (English)
Text-to-image diffusion models are increasingly deployed in open-ended creative contexts, yet their outputs remain impersonal, optimized for aggregate aesthetics rather than individual taste. Human preferences are pluralistic: one user favoring muted, nostalgic portraits may prefer vibrant street photography, while another gravitates toward dreamy film aesthetics. Existing methods require dense interaction histories or per-user fine-tuning, failing in cold-start settings and collapsing context-dependent preferences into a static representation. We introduce zero-shot image personalization from personas (ZIPP), which conditions image generation on natural-language personas (concise descriptors of a user's identity and aesthetic sensibilities) without any user-specific data or weight updates. ZIPP uses an LLM to rewrite prompts from the perspective of a given persona, steering diffusion models toward personalized outputs. To mine personas at scale, we train an inductive Graph Attention Network over a 22M-user Reddit interaction graph with dual contrastive objectives aligning graph structure with visual behavior, then verbalize learned representations into natural-language personas via an MLLM. We introduce ZIPBench, the first zero-shot personalization benchmark with 1.5K users, graph-mined personas, and 40K generated images. Across four benchmarks and 14 LLMs spanning five model families, persona conditioning yields consistent gains (13-20%), with frontier models benefiting most. In the few-shot setting, ZIPP matches or exceeds fine-tuned baselines trained on 100+ examples per user. ZIPP achieves the lowest preference distributional divergence (CMMD 0.16 vs. 0.55), and IPF-normalized demographic evaluation shows it substantially reduces subpopulation bias present in existing methods. Human evaluation confirms a 79% win rate over generic generation and 58-65% over all fine-tuned baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。