arXiv:2512.12596cs.CVcs.AI2025-12

用视觉语言模型生成更懂内容的广告布局

Content-Aware Ad Banner Layout Generation with Two-Stage Chain-of-Thought in Vision Language Models

  • 先分析图像内容生成放置计划,再转为网页代码
  • 相比传统方法,布局质量显著提升
  • 适合需要精准图文排版的广告设计场景

本文提出一种基于视觉语言模型(VLM)的图像类广告布局生成方法。传统广告布局多依赖显著性图检测背景图像中的显著区域,但难以充分考虑图像的细节构成与语义内容。为此,本方法利用VLM识别背景中的产品及其他元素,并据此指导文本与标志的布局。整体流程分为两步:第一步,VLM分析图像以识别物体类型及其空间关系,生成基于文本的‘放置计划’;第二步,将该计划渲染为最终布局的HTML格式代码。通过定量与定性对比实验验证,结果显示,通过显式考虑背景图像内容,本方法生成的广告布局质量明显更高。

原文摘要 · Abstract (English)

In this paper, we propose a method for generating layouts for image-based advertisements by leveraging a Vision-Language Model (VLM). Conventional advertisement layout techniques have predominantly relied on saliency mapping to detect salient regions within a background image, but such approaches often fail to fully account for the image's detailed composition and semantic content. To overcome this limitation, our method harnesses a VLM to recognize the products and other elements depicted in the background and to inform the placement of text and logos. The proposed layout-generation pipeline consists of two steps. In the first step, the VLM analyzes the image to identify object types and their spatial relationships, then produces a text-based "placement plan" based on this analysis. In the second step, that plan is rendered into the final layout by generating HTML-format code. We validated the effectiveness of our approach through evaluation experiments, conducting both quantitative and qualitative comparisons against existing methods. The results demonstrate that by explicitly considering the background image's content, our method produces noticeably higher-quality advertisement layouts.

广告生成视觉语言模型布局优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。