arXiv:2505.07843cs.GRcs.LG2025-05CVPR被引 20

用语言模型生成通用海报布局,支持多样元素和设计意图。

PosterO: Structuring Layout Trees to Enable Language Models in Generalized Content-Aware Layout Generation

  • 将布局结构化为树形SVG,通过意图向量化实现层次化表示。
  • 基于上下文学习选择示例,生成符合设计意图的新布局树。
  • 适用于多场景海报设计,尤其适合复杂形状与多样化需求。

在海报设计中,内容感知布局生成对于自动排布图像中的图文元素至关重要。现有方法受限于训练数据,多聚焦于图像增强,忽视了布局多样性,难以应对形状可变元素及多样化设计意图。为此,本文提出以布局为核心的PosterO方法,利用大语言模型(LLM)中隐含的布局知识,实现面向多种用途的通用海报生成。具体地,将数据集中布局以统一形状、设计意图向量化与分层节点表示的方式结构化为SVG语言的树形结构;推理时,通过上下文学习并结合意图对齐的示例选择,由LLM预测新布局树;生成后可通过与LLM对话直接编辑,无缝转化为实际海报。大量实验表明,PosterO在多个基准上达到最新性能,生成布局视觉效果优越。为进一步探索其泛化能力,我们构建了首个包含多用途海报与多种形状元素的数据集PStylish7,为先进研究提供挑战性测试平台。

原文摘要 · Abstract (English)

In poster design, content-aware layout generation is crucial for automatically arranging visual-textual elements on the given image. With limited training data, existing work focused on image-centric enhancement. However, this neglects the diversity of layouts and fails to cope with shape-variant elements or diverse design intents in generalized settings. To this end, we proposed a layout-centric approach that leverages layout knowledge implicit in large language models (LLMs) to create posters for omnifarious purposes, hence the name PosterO. Specifically, it structures layouts from datasets as trees in SVG language by universal shape, design intent vectorization, and hierarchical node representation. Then, it applies LLMs during inference to predict new layout trees by in-context learning with intent-aligned example selection. After layout trees are generated, we can seamlessly realize them into poster designs by editing the chat with LLMs. Extensive experimental results have demonstrated that PosterO can generate visually appealing layouts for given images, achieving new state-of-the-art performance across various benchmarks. To further explore PosterO's abilities under the generalized settings, we built PStylish7, the first dataset with multi-purpose posters and various-shaped elements, further offering a challenging test for advanced research.

海报生成语言模型布局结构SVG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。