arXiv:2412.04237cs.CV2024-12被引 15

用视觉反馈让大模型自动修正版面设计,效果更准。

VASCAR: Content-Aware Layout Generation via Visual-Aware Self-Correction

  • 大模型生成版面后,通过看渲染图迭代优化
  • 在海报设计任务中达当前最优效果
  • 不需训练,适配GPT-4o和Gemini等多模型

大型语言模型(LLM)因其生成结构化描述(如HTML或JSON)的能力,在版面生成任务中表现良好。然而,由于无法感知图像,其在需要视觉内容理解的任务(如内容感知版面生成)中存在局限。为此,我们探索将大型视觉语言模型(LVLM)应用于该任务。受设计师迭代修改与启发式评估流程的启发,提出无需训练的视觉感知自修正版面生成方法VASCAR。VASCAR使LVLM(如GPT-4o和Gemini)能基于渲染后的版面图像(以彩色边界框显示在海报背景上)进行迭代优化。大量实验与用户研究证明,VASCAR在版面生成质量上达到当前最优水平,且在GPT-4o与Gemini间具有良好的泛化能力。

原文摘要 · Abstract (English)

Large language models (LLMs) have proven effective for layout generation due to their ability to produce structure-description languages, such as HTML or JSON. In this paper, we argue that while LLMs can perform reasonably well in certain cases, their intrinsic limitation of not being able to perceive images restricts their effectiveness in tasks requiring visual content, e.g., content-aware layout generation. Therefore, we explore whether large vision-language models (LVLMs) can be applied to content-aware layout generation. To this end, inspired by the iterative revision and heuristic evaluation workflow of designers, we propose the training-free Visual-Aware Self-Correction LAyout GeneRation (VASCAR). VASCAR enables LVLMs (e.g., GPT-4o and Gemini) iteratively refine their outputs with reference to rendered layout images, which are visualized as colored bounding boxes on poster background (i.e., canvas). Extensive experiments and user study demonstrate VASCAR's effectiveness, achieving state-of-the-art (SOTA) layout generation quality. Furthermore, the generalizability of VASCAR across GPT-4o and Gemini demonstrates its versatility.

版面生成视觉修正大模型自迭代

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。