arXiv:2604.04192cs.CVcs.AI2026-04被引 6

首个全面评估AI图形设计能力的基准,涵盖布局、排版等五大任务

Graphic-Design-Bench: A Comprehensive Benchmark for Evaluating AI on Graphic Design Tasks

  • 构建50个真实设计模板任务,覆盖布局到动画全链条
  • 现有模型在矢量生成与精细排版上表现不佳,空间推理能力弱
  • 适合研究人机协同设计与评估生成式AI的开发者

我们提出GraphicDesignBench(GDB),首个专为评估AI在专业图形设计任务中表现而设计的综合性基准套件。不同于聚焦自然图像理解或通用文生图的现有基准,GDB针对专业设计的核心挑战:将沟通意图转化为结构化布局、实现精准排版、操作分层构图、生成有效矢量图形以及动画时序推理。该套件包含50项任务,按布局、排版、信息图、模板与设计语义、动画五维划分,每项在理解与生成两种设置下评估,并基于LICA分层构图数据集中的真实设计模板。我们采用标准化指标体系(空间精度、感知质量、文本保真度、语义对齐、结构有效性)评估前沿闭源模型。结果表明,当前模型在复杂布局的空间推理、矢量代码忠实生成、细粒度排版感知及动画时序分解方面仍严重不足。尽管高层语义理解已初步达成,但随着对精度、结构与组合意识要求提升,差距急剧扩大。GDB为追踪具备设计协作能力的AI系统进展提供了严谨可复现的测试平台,完整评估框架已公开。

原文摘要 · Abstract (English)

We introduce GraphicDesignBench (GDB), the first comprehensive benchmark suite designed specifically to evaluate AI models on the full breadth of professional graphic design tasks. Unlike existing benchmarks that focus on natural-image understanding or generic text-to-image synthesis, GDB targets the unique challenges of professional design work: translating communicative intent into structured layouts, rendering typographically faithful text, manipulating layered compositions, producing valid vector graphics, and reasoning about animation. The suite comprises 50 tasks organized along five axes: layout, typography, infographics, template & design semantics and animation, each evaluated under both understanding and generation settings, and grounded in real-world design templates drawn from the LICA layered-composition dataset. We evaluate a set of frontier closed-source models using a standardized metric taxonomy covering spatial accuracy, perceptual quality, text fidelity, semantic alignment, and structural validity. Our results reveal that current models fall short on the core challenges of professional design: spatial reasoning over complex layouts, faithful vector code generation, fine-grained typographic perception, and temporal decomposition of animations remain largely unsolved. While high-level semantic understanding is within reach, the gap widens sharply as tasks demand precision, structure, and compositional awareness. GDB provides a rigorous, reproducible testbed for tracking progress toward AI systems that can function as capable design collaborators. The full evaluation framework is publicly available.

图形设计基准测试生成模型人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。