arXiv:2510.18433cs.CVcs.AI2025-10ICCV被引 2

构建首个真实场景的生成图像个性化数据集,助力模型理解用户细微偏好。

ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization

  • 收集57K用户真实交互数据,涵盖242K LoRA、3M提示词和5M生成图像。
  • 基于用户偏好训练出更优对齐模型,提升个性化图像检索与推荐效果。
  • 提出隐空间编辑框架,实现扩散模型的个性化快速调整,适合研究者与开发者。

我们提出ImageGem,一个用于研究生成模型理解细粒度个体偏好的数据集。当前生成模型发展受限于缺乏真实场景下的细粒度用户偏好标注。本数据集包含57,000名用户的实际交互数据,累计创建242,000个定制LoRA、撰写300万条文本提示词、生成500万张图像。利用该数据集中的用户偏好标注,我们训练了性能更优的偏好对齐模型。此外,基于个体用户偏好,我们评估了检索模型与视觉-语言模型在个性化图像检索及生成模型推荐中的表现。最后,我们提出一种端到端的隐空间权重编辑框架,可对定制扩散模型进行个性化调整。实验表明,ImageGem首次实现了生成模型个性化的新范式。

原文摘要 · Abstract (English)

We introduce ImageGem, a dataset for studying generative models that understand fine-grained individual preferences. We posit that a key challenge hindering the development of such a generative model is the lack of in-the-wild and fine-grained user preference annotations. Our dataset features real-world interaction data from 57K users, who collectively have built 242K customized LoRAs, written 3M text prompts, and created 5M generated images. With user preference annotations from our dataset, we were able to train better preference alignment models. In addition, leveraging individual user preference, we investigated the performance of retrieval models and a vision-language model on personalized image retrieval and generative model recommendation. Finally, we propose an end-to-end framework for editing customized diffusion models in a latent weight space to align with individual user preferences. Our results demonstrate that the ImageGem dataset enables, for the first time, a new paradigm for generative model personalization.

生成模型个性化数据集偏好对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。