用美学偏好对齐让大模型生成更符合人类审美的设计布局。
AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models
- 通过美学感知的偏好对齐训练多模态大模型,提升布局美感。
- 在Crello和Webui上分别超越当前最优方法17%和16%。
- 适用于不同规模与架构的大模型,适合设计自动化场景。
视觉布局在广告、海报和网页界面等图形设计领域至关重要。近年来,生成模型被用于内容感知的布局生成,但这些模型难以理解布局设计的上下文美学需求,且不匹配人类偏好,通常将其视为预测任务而忽略最终呈现效果。为此,我们提出美学感知偏好对齐(AAPA)技术,利用多模态大语言模型(MLLM)的美学偏好,通过直接偏好优化(DPO)训练其进行布局预测。我们设计了一套基于布局质量启发式的数据过滤协议,确保仅使用高质量布局进行训练。此外,提出一种新评估指标,借助另一MLLM根据美学标准计算生成布局相对于真实布局的胜率。我们在两个挑战性基准数据集Crello和Webui上验证了该方法的有效性,结果显示在两项任务中分别比现有最优方法提升17%和16%,证明了MLLM在美学感知布局生成中的潜力。
原文摘要 · Abstract (English)
Visual layouts are essential in graphic design fields such as advertising, posters, and web interfaces. The application of generative models for content-aware layout generation has recently gained traction. However, these models fail to understand the contextual aesthetic requirements of layout design and do not align with human-like preferences, primarily treating it as a prediction task without considering the final rendered output. To overcome these problems, we offer Aesthetic-Aware Preference Alignment(AAPA), a novel technique to train a Multi-modal Large Language Model (MLLM) for layout prediction that uses MLLM's aesthetic preferences for Direct Preference Optimization over graphic layouts. We propose a data filtering protocol utilizing our layout-quality heuristics for AAPA to ensure training happens on high-quality layouts. Additionally, we introduce a novel evaluation metric that uses another MLLM to compute the win rate of the generated layout against the ground-truth layout based on aesthetics criteria. We also demonstrate the applicability of AAPA for MLLMs of varying scales (1B to 8B parameters) and LLM families (Qwen, Phi, InternLM). By conducting thorough qualitative and quantitative analyses, we verify the efficacy of our approach on two challenging benchmarks - Crello and Webui, showcasing 17%, and 16 improvement over current State-of-The-Art methods, thereby highlighting the potential of MLLMs in aesthetic-aware layout generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。