arXiv:2510.07665cs.CV2025-10

自动排版中,模型如何智能放置文字框更美观高效?

Automatic Text Box Placement for Supporting Typographic Design

  • 用Transformer处理多图信息,提升布局理解能力
  • 标准Transformer在Crello数据集上表现优于大视觉语言模型
  • 小文字和密集布局仍是挑战,适合设计自动化研究者

在广告和网页布局设计中,视觉吸引力与传播效率的平衡至关重要。本研究探讨不完整布局下的自动文本框放置,对比了标准Transformer模型、小型视觉语言模型Phi3.5-vision、大型预训练视觉语言模型Gemini以及扩展的多图像处理Transformer。在Crello数据集上的评估显示,标准Transformer模型总体优于基于VLM的方法,尤其在引入更丰富的外观信息时表现更佳。然而,所有方法在处理极小文字或高密度布局时仍面临困难。研究结果凸显任务专用架构的优势,并为自动化布局设计的进一步优化提供了方向。

原文摘要 · Abstract (English)

In layout design for advertisements and web pages, balancing visual appeal and communication efficiency is crucial. This study examines automated text box placement in incomplete layouts, comparing a standard Transformer-based method, a small Vision and Language Model (Phi3.5-vision), a large pretrained VLM (Gemini), and an extended Transformer that processes multiple images. Evaluations on the Crello dataset show the standard Transformer-based models generally outperform VLM-based approaches, particularly when incorporating richer appearance information. However, all methods face challenges with very small text or densely populated layouts. These findings highlight the benefits of task-specific architectures and suggest avenues for further improvement in automated layout design.

自动排版视觉语言模型布局生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。