arXiv:2410.12844cs.CLcs.LG2024-10EMNLP被引 10

用文本指令定制大模型,自动生成海报等图形布局。

TextLap: Customizing Language Models for Text-to-Layout Planning

  • 基于指令数据微调大模型,让它像设计师一样理解文本布局需求。
  • 在图文设计基准上超越GPT-4等强基线方法,生成效果更优。
  • 适合需要快速生成视觉布局的UI/UX、广告设计人员使用。

自动生成图形布局对海报、传单、广告和图形用户界面等实际应用至关重要。鉴于大语言模型(LLM)在自然语言理解和生成方面的强大能力,我们提出TextLap(基于文本的布局规划),通过一个精心构建的基于指令的布局规划数据集(InsLap)来定制LLM,使其具备图形设计能力。实验表明,TextLap在图像生成和图形设计基准测试中均优于多个强基线方法,包括基于GPT-4的方法,验证了其有效性。

原文摘要 · Abstract (English)

Automatic generation of graphical layouts is crucial for many real-world applications, including designing posters, flyers, advertisements, and graphical user interfaces. Given the incredible ability of Large language models (LLMs) in both natural language understanding and generation, we believe that we could customize an LLM to help people create compelling graphical layouts starting with only text instructions from the user. We call our method TextLap (text-based layout planning). It uses a curated instruction-based layout planning dataset (InsLap) to customize LLMs as a graphic designer. We demonstrate the effectiveness of TextLap and show that it outperforms strong baselines, including GPT-4 based methods, for image generation and graphical design benchmarks.

文本生成布局规划大模型定制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。