用大模型对话+语义分析,让AI更懂个人审美偏好。
AI Outperforms Humans in Personalized Image Aesthetics Assessment via LLM-Based Interviews and Semantic Feature Extraction

- 通过大模型对话主动收集用户审美偏好,融合低层图像特征与高层语义信息。
- 在高质量图片上表现超越人类与自身重评,误差低于人自身的评价波动。
- 适合个性化推荐、艺术创作辅助等需要精准理解个人审美的场景。
准确预测个体对图像的审美评价是人工智能的核心挑战。现有深度学习模型多基于图像评价数据提取客观低层特征,但审美偏好具有主观性和个体差异性。本研究提出一种集成深度学习与大语言模型(LLM)的系统,通过大模型引导的半结构化访谈主动获取目标个体的偏好信息,并结合低层与高层语义特征进行审美评估。实验对比了该系统与传统模型、人类评估者及个体自身后续重评的表现。结果表明,该系统在所有对比中表现最优,尤其在高分图像上优势显著;其预测误差小于个体内部变异性,而人类评估者误差最大,可能受其自身审美价值观影响。这表明,在特定时刻,AI或比人类自身更能准确捕捉个体审美偏好,引发关于AI是否能成为人类审美感知更深层解读者的思考。
原文摘要 · Abstract (English)
Accurately predicting individual aesthetic evaluation for images is a fundamental challenge for AI. Various deep learning (DL)-based models have been proposed for this task, training on image evaluation data to extract objective low-level features. However, aesthetic preferences are inherently subjective and individual-dependent. Accurate prediction thus requires the extraction of high-level semantic features of images and the active collection of preference information from the target individual. To address this issue, we focus on the utility of Large Language Models (LLMs) pretrained on vast amounts of textual data, and develop an integrated DL-LLM system. The system actively elicits aesthetic preferences through LLM-based semi-structured interviews and predicts aesthetic evaluation by leveraging both low-level and high-level features. In our experiments, we compare the proposed system against conventional systems, human predictors, and the target individual's own re-evaluations after a certain time interval. Our results show that the proposed system outperforms all of them, with particularly strong performance on highly-rated images. Moreover, the prediction error of the proposed system is smaller than within-person variability, while human predictors show the largest error, likely due to the influence of their own aesthetic values. These results suggest that AI may be better positioned than others or one's future self to capture individual aesthetic preferences at a given point. This opens a new question of whether AI could serve as a deeper interpreter of human aesthetic sensibility than humans themselves.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。