用对话式方法学习用户图像偏好,更懂你的审美。
Learning User Preferences for Image Generation Model
- 基于多模态大模型,用对比损失和可学习偏好标记捕捉用户喜好
- 在真实交互数据上预测准确率超越现有方法,能识别相似审美的用户
- 适合个性化图像生成、推荐系统等需要理解用户偏好的场景
用户偏好预测需要全面且准确地理解个人审美,涵盖颜色、风格等表层特征,以及主题、构图等深层内容。然而现有方法多依赖通用人类偏好或静态用户画像,忽视个体差异与审美的动态多样性。为此,我们提出一种基于多模态大语言模型的方法,引入对比偏好损失和可学习偏好标记,从历史交互中学习个性化偏好。对比偏好损失有效区分用户“喜欢”与“不喜欢”,可学习偏好标记则捕捉用户间的共性兴趣,激活群体特定偏好,提升相似用户的生成一致性。大量实验表明,该模型在偏好预测准确率上优于其他方法,能有效识别具有相似审美倾向的用户,并为生成契合个体口味的图像提供更精准引导。
原文摘要 · Abstract (English)
User preference prediction requires a comprehensive and accurate understanding of individual tastes. This includes both surface-level attributes, such as color and style, and deeper content-related aspects, such as themes and composition. However, existing methods typically rely on general human preferences or assume static user profiles, often neglecting individual variability and the dynamic, multifaceted nature of personal taste. To address these limitations, we propose an approach built upon Multimodal Large Language Models, introducing contrastive preference loss and preference tokens to learn personalized user preferences from historical interactions. The contrastive preference loss is designed to effectively distinguish between user ''likes'' and ''dislikes'', while the learnable preference tokens capture shared interest representations among existing users, enabling the model to activate group-specific preferences and enhance consistency across similar users. Extensive experiments demonstrate our model outperforms other methods in preference prediction accuracy, effectively identifying users with similar aesthetic inclinations and providing more precise guidance for generating images that align with individual tastes. The project page is \texttt{https://learn-user-pref.github.io/}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。