用对比判断预测艺术审美,效率更高且效果接近直接评分。
Modeling Art Evaluations from Comparative Judgments: A Deep Learning Approach to Predicting Aesthetic Preferences
- 用成对比较代替直接打分,降低标注负担。
- 深度模型比基线提升328%的预测准确率,对比模型接近回归表现。
- 适合大规模审美数据构建,尤其节省人工标注时间。
视觉艺术中的审美判断因个体偏好差异大且标注成本高而难以建模。为降低标注成本,本文采用基于成对偏好判断的对比学习框架,利用相对选择认知负担小、一致性高的特性。通过ResNet-50提取画作深层特征,构建深度神经网络回归模型与双分支成对比较模型。研究四个问题:(RQ1)CNN特征的深度回归模型是否优于使用手工特征的线性基线?(RQ2)在无直接评分的情况下,对比学习能否媲美回归预测?(RQ3)能否预测个体评分偏好?(RQ4)直接评分与对比判断在标注时间上的成本差异?结果表明,深度回归模型相比基线实现高达328%的R²提升;对比模型虽无直接评分信息,但性能接近回归模型,验证其实用性;然而,个体偏好预测表现远低于平均评分预测。人类实验显示,对比判断每项标注耗时减少60%,显著提升大规模偏好建模效率。
原文摘要 · Abstract (English)
Modeling human aesthetic judgments in visual art presents significant challenges due to individual preference variability and the high cost of obtaining labeled data. To reduce cost of acquiring such labels, we propose to apply a comparative learning framework based on pairwise preference assessments rather than direct ratings. This approach leverages the Law of Comparative Judgment, which posits that relative choices exhibit less cognitive burden and greater cognitive consistency than direct scoring. We extract deep convolutional features from painting images using ResNet-50 and develop both a deep neural network regression model and a dual-branch pairwise comparison model. We explored four research questions: (RQ1) How does the proposed deep neural network regression model with CNN features compare to the baseline linear regression model using hand-crafted features? (RQ2) How does pairwise comparative learning compare to regression-based prediction when lacking access to direct rating values? (RQ3) Can we predict individual rater preferences through within-rater and cross-rater analysis? (RQ4) What is the annotation cost trade-off between direct ratings and comparative judgments in terms of human time and effort? Our results show that the deep regression model substantially outperforms the baseline, achieving up to $328\%$ improvement in $R^2$. The comparative model approaches regression performance despite having no access to direct rating values, validating the practical utility of pairwise comparisons. However, predicting individual preferences remains challenging, with both within-rater and cross-rater performance significantly lower than average rating prediction. Human subject experiments reveal that comparative judgments require $60\%$ less annotation time per item, demonstrating superior annotation efficiency for large-scale preference modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。