arXiv:2603.01083cs.CV2026-03被引 3

首个系统评估视觉语言模型设计美学能力的基准与数据集

Can Vision Language Models Assess Graphic Design Aesthetics? A Benchmark, Evaluation, and Dataset Perspective

  • 构建涵盖四维度十二指标的综合评测框架
  • 发现主流模型在美学判断上显著落后于人类表现
  • 用人工引导标注生成大规模训练数据,提升模型适配性

评估图形设计的审美质量是视觉传播的核心,但当前视觉语言模型(VLMs)对此研究不足。现有工作存在三大局限:评测基准局限于狭窄原则且评价方式粗糙、缺乏系统的VLM对比、训练数据有限。本文提出AesEval-Bench,一个涵盖四个维度、十二个指标、三个可量化任务(审美判断、区域选择、精确定位)的综合性基准。系统评估了专有、开源及推理增强型VLMs,揭示其在审美评估中存在明显性能差距。此外,构建了一个用于领域微调的训练数据集,通过人工引导的VLM标注实现规模化任务标签生成,并基于指标导向的推理将抽象指标与具体设计区域关联。本工作建立了首个系统化的图形设计美学评估框架。代码与数据集将公开于:https://github.com/arctanxarc/AesEval-Bench

原文摘要 · Abstract (English)

Assessing the aesthetic quality of graphic design is central to visual communication, yet remains underexplored in vision language models (VLMs). We investigate whether VLMs can evaluate design aesthetics in ways comparable to humans. Prior work faces three key limitations: benchmarks restricted to narrow principles and coarse evaluation protocols, a lack of systematic VLM comparisons, and limited training data for model improvement. In this work, we introduce AesEval-Bench, a comprehensive benchmark spanning four dimensions, twelve indicators, and three fully quantifiable tasks: aesthetic judgment, region selection, and precise localization. Then, we systematically evaluate proprietary, open-source, and reasoning-augmented VLMs, revealing clear performance gaps against the nuanced demands of aesthetic assessment. Moreover, we construct a training dataset to fine-tune VLMs for this domain, leveraging human-guided VLM labeling to produce task labels at scale and indicator-grounded reasoning to tie abstract indicators to concrete design regions.Together, our work establishes the first systematic framework for aesthetic quality assessment in graphic design. Our code and dataset will be released at: \href{https://github.com/arctanxarc/AesEval-Bench}{https://github.com/arctanxarc/AesEval-Bench}

图像评估视觉语言模型设计美学数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。