提出细粒度图像美学评估新方法,精准区分细微美感差异。
Fine-grained Image Aesthetic Assessment: Learning Discriminative Scores from Relative Ranks
- 基于相对排序学习,通过差异保持标记等技术提升细粒度判别力。
- 在32,217张图像的10,028个序列上实现高精度美学评分。
- 适合需要精细筛选美感图片的创作与推荐场景。
图像美学评估(IAA)在内容创作、相册管理与推荐系统中应用广泛。实际场景常需从一系列审美差异极小的图像中选出最美观者,此即细粒度IAA。然而现有模型多针对显著差异的粗粒度评估,难以区分细微美感差异。为此,我们构建了FGAesthetics数据集,包含32,217张图像,分属10,028个系列,涵盖自然、AIGC和裁剪等多样类别,通过系列内成对比较获取标注,并采用序列优化与排名校准确保标签可靠性。基于此,我们提出FGAesQ框架,通过差值保持标记(DiffToken)、对比文本辅助对齐(CTAlign)与排名感知回归(RankReg),从相对排名中学习可区分的美学得分。该方法在细粒度评估中表现优异,同时保持粗粒度下的竞争力。大量实验验证其优越性。
原文摘要 · Abstract (English)
Image aesthetic assessment (IAA) has extensive applications in content creation, album management, and recommendation systems, etc. In such applications, it is commonly needed to pick out the most aesthetically pleasing image from a series of images with subtle aesthetic variations, a topic we refer to as fine-grained IAA. Unfortunately, state-of-the-art IAA models are typically designed for coarse-grained evaluation, where images with notable aesthetic differences are evaluated independently on an absolute scale. These models are inherently limited in discriminating fine-grained aesthetic differences. To address the dilemma, we contribute FGAesthetics, a fine-grained IAA database with 32,217 images organized into 10,028 series, which are sourced from diverse categories including Natural, AIGC, and Cropping. Annotations are collected via pairwise comparisons within each series. We also devise Series Refinement and Rank Calibration to ensure the reliability of data and labels. Based on FGAesthetics, we further propose FGAesQ, a novel IAA framework that learns discriminative aesthetic scores from relative ranks through Difference-preserved Tokenization (DiffToken), Comparative Text-assisted Alignment (CTAlign), and Rank-aware Regression (RankReg). FGAesQ enables accurate aesthetic assessment in fine-grained scenarios while still maintains competitive performance in coarse-grained evaluation. Extensive experiments and comparisons demonstrate the superiority of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。