用用户画像实现零样本个性化审美评估,无需历史评分数据。
Enhancing Zero-shot Personalized Image Aesthetics Assessment with Profile-aware Multimodal LLM

- 通过用户画像引导多模态大模型,动态融合视觉信息进行个性化判断。
- 在多个基准上达到领先零样本性能,即使画像信息粗略也有效。
- 适合无用户评分数据的场景,如新用户或隐私敏感应用。
个性化图像审美评估(PIAA)旨在预测用户对图像的主观评分,需建模个体审美偏好。现有方法依赖历史评分数据,但在无历史数据时表现不佳。本文提出零样本设置下的画像驱动范式,引入P-MLLM:一种基于画像感知的多模态大模型,在冻结LLM基础上添加可选融合模块,实现受控的视觉信息整合。该模块在画像条件推理过程中,选择性地将视觉信息融入模型隐状态,确保视觉内容以符合用户画像的方式被理解。在最新PIAA基准上的实验表明,P-MLLM实现了具有竞争力的零样本性能,且在使用粗粒度画像信息时仍保持有效性,验证了基于画像的个性化在零样本PIAA中的潜力。
原文摘要 · Abstract (English)
Personalized image aesthetics assessment (PIAA) aims to predict an individual user's subjective rating of an image, which requires modeling user-specific aesthetic preferences. Existing methods rely on historical user ratings for this modeling and therefore struggle when such data are unavailable. We address this zero-shot setting by using user profiles as contextual signals for personalization and adopting a profile-based personalization paradigm. We introduce P-MLLM, a profile-aware multimodal LLM that augments a frozen LLM with selective fusion modules for controlled visual integration. These modules selectively integrate visual information into the model's evolving hidden states during profile-conditioned reasoning, allowing visual information to be incorporated in a profile-aware manner. Experiments on recent PIAA benchmarks show that P-MLLM achieves competitive zero-shot performance and remains effective even with coarse profile information, highlighting the potential of profile-based personalization for zero-shot PIAA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。