arXiv:2507.05601cs.CV2025-07ICCV被引 9

将AI生成的图像转为可编辑分层设计,支持文本优化与风格延续。

Rethinking Layered Graphic Design Generation with a Top-Down Approach

  • 自顶向下分解视觉整体,用提示词引导多阶段生成
  • 在DesignIntention上实现文本到模板等任务的优秀表现
  • 适合需要快速迭代设计的创意工作者使用

图形设计对传达思想至关重要。设计师通常将作品组织为对象、背景和矢量文本层以方便编辑,但此流程需较高专业能力。随着生成式AI的发展,像素格式的高质量设计大量涌现,但缺乏可编辑性。尽管如此,非分层设计仍能启发人类设计师,在版式与文字风格上提供参考。为此,我们提出Accordion框架,首次尝试将AI生成的设计转换为可编辑的分层结构,并利用用户提示优化无意义的AI生成文本。该框架基于视觉语言模型(VLM),在三个定制阶段分别执行不同任务。不同于现有自底向上的方法(如COLE和Open-COLE),本方法以视觉和谐的参考图作为全局指导,进行分层分解。同时引入SAM和元素移除模型等视觉专家辅助生成。我们在自建数据集Design39K基础上,结合由定制修复模型生成的增强版真实标签进行训练。实验与设计师用户研究显示,Accordion在DesignIntention基准测试中表现优异,涵盖文本到模板、向背景加文、文本去渲染等任务,且擅长生成设计变体。

原文摘要 · Abstract (English)

Graphic design is crucial for conveying ideas and messages. Designers usually organize their work into objects, backgrounds, and vectorized text layers to simplify editing. However, this workflow demands considerable expertise. With the rise of GenAI methods, an endless supply of high-quality graphic designs in pixel format has become more accessible, though these designs often lack editability. Despite this, non-layered designs still inspire human designers, influencing their choices in layouts and text styles, ultimately guiding the creation of layered designs. Motivated by this observation, we propose Accordion, a graphic design generation framework taking the first attempt to convert AI-generated designs into editable layered designs, meanwhile refining nonsensical AI-generated text with meaningful alternatives guided by user prompts. It is built around a vision language model (VLM) playing distinct roles in three curated stages. For each stage, we design prompts to guide the VLM in executing different tasks. Distinct from existing bottom-up methods (e.g., COLE and Open-COLE) that gradually generate elements to create layered designs, our approach works in a top-down manner by using the visually harmonious reference image as global guidance to decompose each layer. Additionally, it leverages multiple vision experts such as SAM and element removal models to facilitate the creation of graphic layers. We train our method using the in-house graphic design dataset Design39K, augmented with AI-generated design images coupled with refined ground truth created by a customized inpainting model. Experimental results and user studies by designers show that Accordion generates favorable results on the DesignIntention benchmark, including tasks such as text-to-template, adding text to background, and text de-rendering, and also excels in creating design variations.

图形生成分层设计视觉语言模型可编辑生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。