arXiv:2411.11435cs.CV2024-11被引 3

用多模态大模型生成更美观的文本标志布局,提升设计效率。

GLDesigner: Leveraging Multi-Modal LLMs as Designer for Enhanced Aesthetic Text Glyph Layouts

  • 基于视觉语言模型,融合多模态输入与用户约束生成布局。
  • 构建比现有数据集大5倍的文本标志数据集,支持指令微调。
  • 在几何美学和人偏好测试中显著优于现有方法,适合实际应用。

文本标志设计高度依赖专业设计师的创意与经验,其中元素排布是关键步骤,但该任务长期未受重视,常被文档或海报布局等更广泛的生成任务掩盖。本文提出一种基于视觉语言模型(VLM)的框架,通过整合多模态输入与用户定义约束,生成内容感知的文本标志布局,实现更灵活、鲁棒的布局生成。我们引入两项技术,降低同时处理多个字形图像的计算成本,且不损害性能。为支持模型指令微调,我们构建了两个大规模文本标志数据集,规模是现有公开数据集的五倍。除几何标注(如文本掩码和字符识别)外,数据集还包含自然语言描述的详细布局信息,使模型能更好应对复杂设计与自定义输入。实验表明,所提框架与数据集在评估几何美感与人类偏好的多个基准上均表现优异。

原文摘要 · Abstract (English)

Text logo design heavily relies on the creativity and expertise of professional designers, in which arranging element layouts is one of the most important procedures. However, this specific task has received limited attention, often overshadowed by broader layout generation tasks such as document or poster design. In this paper, we propose a Vision-Language Model (VLM)-based framework that generates content-aware text logo layouts by integrating multi-modal inputs with user-defined constraints, enabling more flexible and robust layout generation for real-world applications. We introduce two model techniques that reduce the computational cost for processing multiple glyph images simultaneously, without compromising performance. To support instruction tuning of our model, we construct two extensive text logo datasets that are five times larger than existing public datasets. In addition to geometric annotations (\textit{e.g.}, text masks and character recognition), our datasets include detailed layout descriptions in natural language, enabling the model to reason more effectively in handling complex designs and custom user inputs. Experimental results demonstrate the effectiveness of our proposed framework and datasets, outperforming existing methods on various benchmarks that assess geometric aesthetics and human preferences.

文本生成多模态设计自动化布局优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。