arXiv:2410.14170cs.IRcs.AI2024-10中稿 · publication in WWW…被引 19

用大模型实现个性化图像生成,解决用户偏好捕捉难题

Personalized Image Generation with Large Multimodal Models

  • 基于大模型设计三模块框架,从杂乱历史中提取用户视觉偏好
  • 提出两阶段偏好对齐机制,缓解个性化图像生成数据稀缺问题
  • 在贴纸和电影海报生成中表现优于现有方法,适合创意设计场景

个性化内容过滤(如推荐系统)已成为缓解信息过载的关键基础设施,但仅能筛选现有内容,受限于其多样性不足,难以满足用户多样化需求。为此,个性化内容生成成为新方向。然而,现有研究多聚焦文本生成,图像生成研究较少,且面临从噪声用户交互图像和复杂多模态指令中准确捕捉视觉偏好的挑战,同时缺乏用于训练的标注数据。为克服上述困难,本文提出名为 Pigeon 的个性化图像生成框架,采用先进的大型多模态模型,并设计三个专用模块,从用户历史记录和多模态指令中提取视觉偏好。为缓解数据稀缺问题,引入两阶段偏好对齐策略:掩码偏好重建与成对偏好对齐,以使 Pigeon 更好适配个性化图像生成任务。在个性化贴纸与电影海报生成任务中,实验结果表明,Pigeon 在定量指标与人工评估上均显著优于多种生成基线。

原文摘要 · Abstract (English)

Personalized content filtering, such as recommender systems, has become a critical infrastructure to alleviate information overload. However, these systems merely filter existing content and are constrained by its limited diversity, making it difficult to meet users' varied content needs. To address this limitation, personalized content generation has emerged as a promising direction with broad applications. Nevertheless, most existing research focuses on personalized text generation, with relatively little attention given to personalized image generation. The limited work in personalized image generation faces challenges in accurately capturing users' visual preferences and needs from noisy user-interacted images and complex multimodal instructions. Worse still, there is a lack of supervised data for training personalized image generation models. To overcome the challenges, we propose a Personalized Image Generation Framework named Pigeon, which adopts exceptional large multimodal models with three dedicated modules to capture users' visual preferences and needs from noisy user history and multimodal instructions. To alleviate the data scarcity, we introduce a two-stage preference alignment scheme, comprising masked preference reconstruction and pairwise preference alignment, to align Pigeon with the personalized image generation task. We apply Pigeon to personalized sticker and movie poster generation, where extensive quantitative results and human evaluation highlight its superiority over various generative baselines.

个性化生成多模态模型图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。