用强化学习让AI更懂审美,生成有依据的评分与解释。
Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy Optimization
- 构建高质量审美推理数据链,先教模型讲理由再打分。
- 新算法同时优化分数准确性和图像间偏好排序,提升47.9%相关性。
- 适合需要可解释美学评估的视觉生成与内容推荐场景。
多模态大语言模型(MLLM)具备跨模态理解能力,适合图像审美评估。然而,多模态审美推理数据稀缺且审美判断具有主观性,导致模型难以生成准确且可解释的审美判断。为此,我们提出Aes-R1框架,结合强化学习实现美学推理。具体而言,Aes-R1采用AesCoT流程构建并筛选高质量的思维链审美推理数据,用于冷启动训练。在模型学会生成结构化解释后再进行评分,随后引入相对-绝对策略优化(RAPO)算法,联合优化绝对评分回归与相对排序,显著提升单图评分准确性和跨图偏好判断能力。Aes-R1使MLLM能够生成有依据的评分与解释,在统一框架中增强审美评估与推理性能。大量实验表明,Aes-R1使基线模型平均PLCC/SRCC提升47.9%/34.8%,超越同类规模最优基线。消融实验验证了其在有限监督及分布外场景下的鲁棒泛化能力。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) are well suited to image aesthetic assessment, as they can capture high-level aesthetic features leveraging their cross-modal understanding capacity. However, the scarcity of multimodal aesthetic reasoning data and the inherently subjective nature of aesthetic judgment make it difficult for MLLMs to generate accurate aesthetic judgments with interpretable rationales. To this end, we propose Aes-R1, a comprehensive aesthetic reasoning framework with reinforcement learning (RL). Concretely, Aes-R1 integrates a pipeline, AesCoT, to construct and filter high-quality chain-of-thought aesthetic reasoning data used for cold-start. After teaching the model to generate structured explanations prior to scoring, we then employ the Relative-Absolute Policy Optimization (RAPO), a novel RL algorithm that jointly optimizes absolute score regression and relative ranking order, improving both per-image accuracy and cross-image preference judgments. Aes-R1 enables MLLMs to generate grounded explanations alongside faithful scores, thereby enhancing aesthetic scoring and reasoning in a unified framework. Extensive experiments demonstrate that Aes-R1 improves the backbone's average PLCC/SRCC by 47.9%/34.8%, surpassing state-of-the-art baselines of similar size. More ablation studies validate Aes-R1's robust generalization under limited supervision and in out-of-distribution scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。