评测大模型审美能力,发现现有方法远不如专家。
Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?

- 用对比选择替代打分,更贴近真实审美判断。
- 20个前沿模型在400个任务中仅26.5%答对,人类达68.9%。
- 数据集含1195张图,适合评估视觉审美模型性能。
多模态大语言模型已广泛用于视觉理解、生成与筛选,其中许多应用需明确的审美判断。现有方法多将审美评价简化为单图打分,但我们发现这种评分难以反映真实偏好:在8位专家参与的对照实验中,打分排名与直接比较结果偏差显著,而直接排序的标注者间一致性更高。为此,我们提出视觉审美基准(VAB),将审美评估转化为同主题候选图像间的对比选择。VAB包含400个任务、1195张图像,涵盖美术、摄影与插画,每项任务由10位独立专家达成共识标注。测试20个前沿多模态模型及6个专用视觉质量奖励模型,最强系统在三轮随机排列下仅26.5%任务正确识别最佳与最差图像,远低于人类专家的68.9%准确率。对350亿参数模型用2000个专家样本微调后,其表现接近3970亿参数开源模型,表明VAB中的对比信号具备可迁移性。这些结果揭示了当前多模态模型与专家审美判断间的显著差距,而VAB提供了首个基于集合、专家标注的测评基准,可用于持续追踪与弥补这一差距。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) are now routinely deployed for visual understanding, generation, and curation. A substantial fraction of these applications require an explicit aesthetic judgment. Most existing solutions reduce this judgment to predicting a scalar score for a single image. We first ask whether such scores faithfully capture comparative preference: in a controlled study with eight expert annotators, score-derived rankings align poorly with the same annotators' direct comparisons, while direct ranking yields substantially higher inter-annotator agreement on best- and worst-image labels. Motivated by this finding, we introduce the Visual Aesthetic Benchmark (VAB), which casts aesthetic evaluation as comparative selection over candidate sets with matched subject matter. VAB contains 400 tasks and 1,195 images across fine art, photography, and illustration, with labels derived from the consensus of 10 independent expert judges per task. Evaluating 20 frontier MLLMs and six dedicated visual-quality reward models, we find that the strongest system identifies both the best and the worst image correctly across three random permutations of the candidate order in only 26.5% of tasks, far below the 68.9% achieved by human experts. Fine-tuning a 35B-parameter model on 2,000 expert examples brings its accuracy close to that of a 397B-parameter open-weight model, suggesting that the comparative signal in VAB is transferable. Together, these results expose a clear and measurable gap between current multimodal models and expert aesthetic judgment, and VAB provides the first set-based, expert-grounded testbed on which that gap can be tracked and closed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。