构建首个面向版面设计的人类偏好数据集,提升AI生成布局的审美一致性。
DesignSense: A Human Preference Dataset and Reward Modeling Framework for Graphic Layout Generation
- 通过五阶段流程生成10235组高质量版面对比对,确保空间布局多样性。
- 基于视觉语言模型的评分器在多指标上超越主流模型54.6%的宏平均F1。
- 适用于需高美学质量的版面生成任务,尤其适合训练与评估设计类AI系统。
版面设计是跨渠道视觉传播的重要载体。尽管近期布局生成模型表现突出,但常无法契合人类细腻的审美判断。现有基于文本到图像生成的偏好数据集和评分模型难以泛化至版面评估,因相同元素的空间排列决定质量。为填补这一空白,我们提出DesignSense-10k,一个包含10,235组人工标注偏好对的大规模版面评价数据集。采用五阶段精炼流程——语义分组、布局预测、过滤、聚类与视觉语言模型(VLM)优化,生成跨不同宽高比的视觉一致布局变换。偏好标注采用四分类方案(左优、右优、均好、均差),捕捉主观模糊性。基于该数据集,我们训练了DesignSense,一种基于视觉语言模型的分类器,在综合评估中显著优于现有开源与专有模型,宏平均F1提升54.6%。分析显示,前沿视觉语言模型整体仍不可靠,且在完整四分类任务中出现灾难性失效,凸显专用、偏好感知模型的必要性。除数据集外,我们的评分模型在下游任务中带来实际增益:强化学习训练中使用该判别器使生成器胜率提升约3%;推理阶段多候选采样选择则带来3.6%的提升。结果表明,专用的版面感知偏好建模可显著提升真实世界布局生成质量。
原文摘要 · Abstract (English)
Graphic layouts serve as an important and engaging medium for visual communication across different channels. While recent layout generation models have demonstrated impressive capabilities, they frequently fail to align with nuanced human aesthetic judgment. Existing preference datasets and reward models trained on text-to-image generation do not generalize to layout evaluation, where the spatial arrangement of identical elements determines quality. To address this critical gap, we introduce DesignSense-10k, a large-scale dataset of 10,235 human-annotated preference pairs for graphic layout evaluation. We propose a five-stage curation pipeline that generates visually coherent layout transformations across diverse aspect ratios, using semantic grouping, layout prediction, filtering, clustering, and VLM-based refinement to produce high-quality comparison pairs. Human preferences are annotated using a 4-class scheme (left, right, both good, both bad) to capture subjective ambiguity. Leveraging this dataset, we train DesignSense, a vision-language model-based classifier that substantially outperforms existing open-source and proprietary models across comprehensive evaluation metrics (54.6% improvement in Macro F1 over the strongest proprietary baseline). Our analysis shows that frontier VLMs remain unreliable overall and fail catastrophically on the full four-class task, underscoring the need for specialized, preference-aware models. Beyond the dataset, our reward model DesignSense yields tangible downstream gains in layout generation. Using our judge during RL based training improves generator win rate by about 3%, while inference-time scaling, which involves generating multiple candidates and selecting the best one, provides a 3.6% improvement. These results highlight the practical impact of specialized, layout-aware preference modeling on real-world layout generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。