构建首个个性化视觉设计偏好数据集,支持精准建模设计师个人审美。
DesignPref: Capturing Personal Preferences in Visual Design Generation
- 收集20位专业设计师对1.2万组界面设计的多层级偏好判断。
- 发现设计师间偏好分歧显著(一致性仅0.25),传统多数投票无效。
- 个性化微调或融合特定标注,用20倍少数据仍超越通用模型。
生成式模型如大语言模型和文生图扩散模型正被广泛用于生成用户界面(UI)和演示幻灯片等视觉设计。现有微调与评估常依赖人工标注的设计偏好数据集,但视觉设计具有高度主观性和个性化特征,个体间偏好差异显著。本文提出DesignPref,一个包含1.2万组界面设计配对比较的标注数据集,由20位专业设计师在多层级偏好评分下完成。研究发现,训练有素的设计师间存在显著分歧(二元偏好一致系数Krippendorff's alpha = 0.25)。设计师提供的自然语言理由表明,分歧源于对设计要素重要性的不同认知及个人偏好。基于DesignPref,我们证明传统多数投票法训练的聚合评判模型难以准确反映个体偏好。为此,我们探索多种个性化策略,尤其是将设计师特定标注融入RAG流程或进行微调。结果表明,个性化模型在预测个体设计师偏好方面持续优于聚合基线模型,即使仅使用20倍更少的样本。本工作首次提供研究个性化视觉设计评估的数据集,为未来建模个体设计品味奠定基础。
原文摘要 · Abstract (English)
Generative models, such as large language models and text-to-image diffusion models, are increasingly used to create visual designs like user interfaces (UIs) and presentation slides. Finetuning and benchmarking these generative models have often relied on datasets of human-annotated design preferences. Yet, due to the subjective and highly personalized nature of visual design, preference varies widely among individuals. In this paper, we study this problem by introducing DesignPref, a dataset of 12k pairwise comparisons of UI design generation annotated by 20 professional designers with multi-level preference ratings. We found that among trained designers, substantial levels of disagreement exist (Krippendorff's alpha = 0.25 for binary preferences). Natural language rationales provided by these designers indicate that disagreements stem from differing perceptions of various design aspect importance and individual preferences. With DesignPref, we demonstrate that traditional majority-voting methods for training aggregated judge models often do not accurately reflect individual preferences. To address this challenge, we investigate multiple personalization strategies, particularly fine-tuning or incorporating designer-specific annotations into RAG pipelines. Our results show that personalized models consistently outperform aggregated baseline models in predicting individual designers' preferences, even when using 20 times fewer examples. Our work provides the first dataset to study personalized visual design evaluation and support future research into modeling individual design taste.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。