构建个性化图像评价体系,让生成图像更符合个人审美偏好。
Personalizing Text-to-Image Generation to Individual Taste

- 基于7万条用户评分,建立覆盖多领域的个性化评价数据集
- 模型预测个体偏好准确率超过现有主流方法对群体偏好的预测能力
- 适合研究个性化图像生成、审美主观性与人机对齐的学者使用
当前文本到图像(T2I)模型虽能生成高质量图像,但缺乏对个体偏好的敏感性。现有奖励模型通常优化平均人类偏好,难以捕捉审美的主观性。本文提出PAMELA数据集与预测框架,包含5,000张由Flux 2和Nano Banana生成的多样化图像,每张图由15位不同用户评分,共70,000条评价,覆盖艺术、设计、时尚和电影摄影等领域。基于该数据,我们训练了一个联合高质标注与现有美学评估子集的个性化奖励模型。结果表明,该模型在预测个体喜好方面,优于多数当前SOTA方法对群体偏好的预测表现。进一步验证了简单提示优化即可引导生成满足个体偏好。研究强调数据质量与个性化对处理审美主观性的关键作用。数据集与模型已公开,以推动个性化T2I对齐与主观视觉质量评估的标准化研究。
原文摘要 · Abstract (English)
Modern text-to-image (T2I) models generate high-fidelity visuals but remain indifferent to individual user preferences. While existing reward models optimize for "average" human appeal, they fail to capture the inherent subjectivity of aesthetic judgment. In this work, we introduce a novel dataset and predictive framework, called PAMELA, designed to model personalized image evaluations. Our dataset comprises 70,000 ratings across 5,000 diverse images generated by state-of-the-art models (Flux 2 and Nano Banana). Each image is evaluated by 15 unique users, providing a rich distribution of subjective preferences across domains such as art, design, fashion, and cinematic photography. Leveraging this data, we propose a personalized reward model trained jointly on our high-quality annotations and existing aesthetic assessment subsets. We demonstrate that our model predicts individual liking with higher accuracy than the majority of current state-of-the-art methods predict population-level preferences. Using our personalized predictor, we demonstrate how simple prompt optimization methods can be used to steer generations towards individual user preferences. Our results highlight the importance of data quality and personalization to handle the subjectivity of user preferences. We release our dataset and model to facilitate standardized research in personalized T2I alignment and subjective visual quality assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。