用编辑前后差异学习美学,让模型更懂人眼的审美对比。
Beyond Absolute Scores: Relative Edit-induced Difference for Generalizable Image Aesthetic Assessment

- 通过图像编辑生成对比对,学习美学变化的相对差异
- 在多个基准上达到顶尖性能,泛化能力更强
- 适合研究视觉感知与可解释性评估的学者
传统图像美学评估(IAA)方法主要依赖回归绝对均值评分(MOS),但这种范式忽略了人类审美感知固有的动态性——即基于隐含视觉参照的无意识比较。因此,缺乏对美学差异的因果推理,导致模型难以学习通用美学原则,限制了其在多样场景下的泛化能力。本文重新思考IAA任务,提出相对编辑差异美学学习(RED-Aes)框架,利用可控图像编辑模型模拟人类审美推理过程。不同于拟合绝对评分分布,RED-Aes显式学习驱动美学变化的视觉因素。为此,我们构建了RED-20k数据集,包含基于编辑的图像对、定量美学差异及思维链(CoT)推理。此外,我们设计三阶段训练策略,以相对排序一致性奖励引导优化,仅通过相对监督训练模型。大量实验表明,RED-Aes在多个公开基准上达到当前最优表现,展现出卓越的泛化能力。
原文摘要 · Abstract (English)
Traditional Image Aesthetic Assessment (IAA) methods mainly rely on regressing absolute Mean Opinion Scores (MOS). However, such a paradigm overlooks the inherently dynamic nature of human aesthetic perception, which relies on subconscious comparison against implicit visual references. Consequently, the lack of causal reasoning regarding aesthetic differences prevents models from learning generalizable aesthetic principles, thus limiting their generalization across diverse scenarios. In this work, we rethink the IAA task and propose Relative Edit-induced Difference Aesthetic learning (RED-Aes), a novel framework that leverages controllable image editing models to simulate the human aesthetic reasoning process. Instead of fitting absolute score distributions, RED-Aes explicitly learns the visual factors that drive aesthetic changes. To support this paradigm, we construct the RED-20k dataset, which comprises editing-based image pairs, quantitative aesthetic differences, and Chain-of-Thought (CoT) reasoning. Furthermore, we introduce a three-stage training strategy guided by a relative ranking consistency reward, optimizing the model solely via relative supervision. Extensive experiments demonstrate that RED-Aes achieves state-of-the-art performance on multiple public benchmarks, exhibiting superior generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。