首个跨领域个性化图像审美评估数据集,支持多领域偏好建模。
XPASS-Vis: A Dataset for Cross-Domain Personalized Image Aesthetic Assessment

- 构建跨艺术、时尚、风景三领域的个性化审美数据集
- 每用户每领域超200张图,支持跨域偏好学习
- 验证审美偏好在无监督下可部分迁移,仍需改进
个性化图像审美评估(PIAA)旨在建模个体对艺术品和照片的主观审美判断。审美偏好既具个人差异,又在视觉领域间存在部分一致性。然而现有数据集与方法多局限于单一领域,或每位标注者样本过少,难以实现跨领域个性化。为此,我们提出XPASS-Vis,首个专为跨领域PIAA设计的数据集,包含来自艺术、时尚、风景三个视觉领域的6,526张图像,由129名标注者评分,共产生87,836次用户-图像交互,每条包含整体审美分及九项审美-情绪评分。每位标注者在每个领域均评价超过200张图像,具备足够的领域内覆盖以支持跨域个性化。我们建立无监督域适应(UDA)下的基线模型,系统评估主流方法发现,表现最佳模型在完全无监督设置下达到监督上限约60%(斯皮尔曼相关系数ρ = .28),表明个性化审美偏好在一定程度上可在不同视觉领域间迁移。但差距依然显著,凸显亟需针对PIAA的专用适配策略。XPASS-Vis与配套基线为未来跨域PIAA研究奠定基础,所有数据与代码将在接受后公开。
原文摘要 · Abstract (English)
Personalized image aesthetic assessment (PIAA) seeks to model, at the individual level, the subjective nature of aesthetic judgments toward artworks and photographs. Aesthetic preference is known to be both deeply personal and partially consistent across visual domains. Yet existing PIAA datasets and methods are largely confined to a single domain, or provide too few samples per annotator within each domain to enable personalization across domains. Consequently, the cross-domain generalization of personalized aesthetic preferences remains largely unexplored. To address this gap, we introduce XPASS-Vis, the first dataset explicitly designed for cross-domain PIAA. XPASS-Vis comprises 6,526 stimuli from three visual domains -- art, fashion, and landscape -- rated by 129 annotators, yielding 87,836 user-stimulus interactions, each annotated with an overall aesthetic score and nine aesthetic-emotion ratings. Notably, each annotator rated more than 200 stimuli per domain, providing sufficient per-domain coverage to support personalization both within and across domains. Moreover, we establish baseline models for cross-domain PIAA under unsupervised domain adaptation (UDA), where a model trained on a labeled source domain is transferred to an unlabeled target domain. A systematic evaluation of representative UDA approaches shows that the best-performing method recovers approximately 60\% (Spearman's $ρ$ = .28) of the supervised upper bound under a fully unsupervised setting. This provides encouraging evidence that personalized aesthetic preferences are, to a meaningful extent, transferable across visual domains. At the same time, a substantial gap remains, highlighting the need for PIAA-specific adaptation strategies. XPASS-Vis and the accompanying baselines provide a foundation for future research on cross-domain PIAA. All datasets and code will be made publicly available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。