基于用户偏好挖掘与群体融合,实现更个性化的图像审美评估。
Personalized Image Aesthetic Assessment via Preference-rich Sample Mining and Cohort Merging

- 通过集体争议与个人偏差识别高价值偏好样本
- 利用偏好嵌入相似性构建共鸣用户群体
- 适合需要个性化审美推荐的场景
个性化图像审美评估(PIAA)旨在预测个体差异化的图像审美评分。审美偏好在不同视觉刺激中表现程度各异,并呈现群体特异性模式。为此,本文提出一种基于多模态大语言模型(MLLM)的方法——偏好丰富样本挖掘与美学共鸣群体融合(PRAC)。PRAC首先通过分析图像的集体争议性与个人偏离度,识别出偏好丰富的样本,最大化有限用户数据的利用效率;随后基于偏好嵌入比较跨用户偏好相似性,构建美学共鸣用户群体;最后采用群体模型融合策略,聚合共鸣用户的偏好模式,进一步提升目标个体的个性化表现。在四个基准PIAA数据集上的大量实验表明,所提PRAC模型优于现有最先进方法。代码与模型将公开于https://github.com/yzc-ippl/PRAC。
原文摘要 · Abstract (English)
Personalized Image Aesthetic Assessment (PIAA) aims to predict aesthetic ratings of images that vary across individuals. The aesthetic preferences manifest to different extents across distinct visual stimuli and exhibit cohort-specific patterns. Motivated by the above fact, this paper presents a Multimodal Large Language Model (MLLM)-based approach, which models individual aesthetic preferences by Preference-Rich sample mining and Aesthetically-resonant Cohort merging (PRAC). Specifically, PRAC first identifies preference-rich samples by analyzing both Collective Controversy and Personalized Deviation of images, maximizing the utility of limited user data. Based upon the preference-rich samples, cross-user preference similarities are measured by comparing preference embeddings. Then, a cohort-based model merging strategy, is proposed by aggregating preference patterns from aesthetically-resonant users, which further enhances the personalization for the target individual. Extensive experiments and comparisons on four benchmark PIAA databases demonstrate the superiority of the proposed PRAC model over the state-of-the-arts. The code and model will be public at https://github.com/yzc-ippl/PRAC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。