构建多维度设计偏好数据集,让AI理解设计的细节优劣。
TASTE: A Designer-Annotated Multi-Dimensional Preference Dataset for AI-Generated Graphic Design

- 九个设计维度由专业设计师标注,包含排版、配色等
- 发现设计师共识中等,但显著高于随机水平
- 为图像生成模型提供可训练的评估基准
文本到图像模型已能以生产规模生成图形设计,但其训练仍依赖以照片风格为主的单维度偏好数据集,仅提供整体评价。设计师在评估设计时会从多个独立维度(如排版、布局、色彩和谐)进行判断,而单一偏好标签会掩盖这些差异。我们发布 TASTE(Typography, Aesthetics, Spatial, Tone, Etc.),一个由两组各五名专业设计师组成的多维度偏好数据集,对四种主流文生图模型的输出在九个维度上进行评分,并标注每张图的幻觉情况。同时,我们提出一种基于肯德尔τ、多数投票概率和康多塞循环的无准则信号验证框架,分析显示设计师间存在显著但中等程度的一致性,所有TASTE维度均拒绝随机打分假设。我们在TASTE上基准测试了偏好模型,发现现有VLM评判器和专用文生图评分器无法达到设计师小组的多数一致,而直接在TASTE上训练的小型MLP头显著缩小与单人上限的差距,为未来模型提供了基准。
原文摘要 · Abstract (English)
Text-to-image models now generate graphic design at production scale, yet their supervision still comes primarily from photo-style preference datasets with a single overall verdict per comparison. Designers evaluate designs along several distinct axes (e.g., typography, layout, color harmony) that a single preference label collapses. We release \emph{TASTE} \textit{(Typography, Aesthetics, Spatial, Tone, Etc.)}, a multi-dimensional preference dataset in which two disjoint cohorts of five professional designers each ranked outputs from four current text-to-image models across nine criteria along with per-image hallucination flags. We pair the dataset with two contributions. First, a criterion-agnostic signal-validation framework based on Kendall's $τ$, majority-vote probability, and Condorcet cycles against exact iid-uniform nulls; the analysis reveals significant but moderate designer agreement, with every TASTE criterion rejecting the random-rater null. Second, we benchmark preference models on TASTE and find that off-the-shelf VLM judges and dedicated T2I scorers fail to reach majority agreement with the designer panel, while a small MLP head trained directly on TASTE substantially narrows the gap to the single-rater ceiling, setting a baseline for future TASTE-trained preference models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。