用产品信息自动生成广告图,流程完整且视觉效果更优。
T-Stars-Poster: A Framework for Product-Centric Advertising Image Design
- 分四步生成:提示词、布局、背景图、图形渲染,全流程闭环。
- 在5万+标注图像上训练,实测生成图像更美观。
- 适合电商广告自动化,提升设计效率。
广告图像制作通常耗时耗力。能否仅凭产品前景图、标语和目标尺寸等基础信息自动完成?现有方法多聚焦局部问题,缺乏整体解决方案。为此,我们提出一种以产品为中心的广告图像设计框架T-Stars-Poster,包含四个连续阶段:提示词生成、布局生成、背景图像生成与图形渲染。针对前三个阶段,分别设计并训练了专家模型:首先使用视觉语言模型(VLM)生成匹配产品的背景提示词;其次,基于VLM的布局生成模型将产品前景、图形元素(标语与装饰底图)及非图形元素(来自背景提示的物体)进行合理排布;最后,基于SDXL的模型可同时接收提示词、布局和前景控制信号,生成最终图像。为支持该框架,我们构建了两个包含超过5万张标注图像的数据集。大量实验与线上A/B测试表明,T-Stars-Poster能生成更具视觉吸引力的广告图像。
原文摘要 · Abstract (English)
Creating advertising images is often a labor-intensive and time-consuming process. Can we automatically generate such images using basic product information like a product foreground image, taglines, and a target size? Existing methods mainly focus on parts of the problem and lack a comprehensive solution. To bridge this gap, we propose a novel product-centric framework for advertising image design called T-Stars-Poster. It consists of four sequential stages to highlight product foregrounds and taglines while achieving overall image aesthetics: prompt generation, layout generation, background image generation, and graphics rendering. Different expert models are designed and trained for the first three stages: First, a visual language model (VLM) generates background prompts that match the products. Next, a VLM-based layout generation model arranges the placement of product foregrounds, graphic elements (taglines and decorative underlays), and various nongraphic elements (objects from the background prompt). Following this, an SDXL-based model can simultaneously accept prompts, layouts, and foreground controls to generate images. To support T-Stars-Poster, we create two corresponding datasets with over 50,000 labeled images. Extensive experiments and online A/B tests demonstrate that T-Stars-Poster can produce more visually appealing advertising images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。