用自然语言一键生成可编辑的多模态设计图层
IGD: Instructional Graphic Design with Multimodal Layer Generation
- 通过多模态大模型理解指令,自动规划图文排布
- 支持端到端训练,生成结果可编辑且文本清晰可读
- 适合设计师快速原型制作与自动化内容生产
图形设计通过文字、图像和图形的组合来可视化传达信息。现有两阶段布局生成方法缺乏创意与智能,仍依赖人工;基于扩散模型的方法在图像层面生成不可编辑的设计文件,视觉文本可读性差,难以实现实用化自动化设计。本文提出指令式图形设计框架IGD,仅需自然语言指令即可快速生成具有可编辑性的多模态图层。IGD采用参数化渲染与图像资产生成的新范式:首先构建设计平台并建立多场景设计文件的标准化格式,为数据规模化奠定基础;其次,利用多模态大模型(MLLM)完成图层属性预测、顺序排列与布局规划,并结合扩散模型生成资产图像内容。通过端到端训练,IGD在复杂设计任务中具备良好的可扩展性与可拓展性。实验结果表明,IGD提供了一种全新的自动化图形设计解决方案。
原文摘要 · Abstract (English)
Graphic design visually conveys information and data by creating and combining text, images and graphics. Two-stage methods that rely primarily on layout generation lack creativity and intelligence, making graphic design still labor-intensive. Existing diffusion-based methods generate non-editable graphic design files at image level with poor legibility in visual text rendering, which prevents them from achieving satisfactory and practical automated graphic design. In this paper, we propose Instructional Graphic Designer (IGD) to swiftly generate multimodal layers with editable flexibility with only natural language instructions. IGD adopts a new paradigm that leverages parametric rendering and image asset generation. First, we develop a design platform and establish a standardized format for multi-scenario design files, thus laying the foundation for scaling up data. Second, IGD utilizes the multimodal understanding and reasoning capabilities of MLLM to accomplish attribute prediction, sequencing and layout of layers. It also employs a diffusion model to generate image content for assets. By enabling end-to-end training, IGD architecturally supports scalability and extensibility in complex graphic design tasks. The superior experimental results demonstrate that IGD offers a new solution for graphic design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。