用语言模型生成可编辑的多样化设计,支持风格与布局自定义。
CPT: Controllable and Editable Design Variations with Language Models
- 通过创意标记语言CML表示设计结构与风格细节
- 在专业设计模板上微调,实现颜色字体的上下文相关生成
- 输出可编辑文档,适合设计师快速迭代个性化方案
设计高质量且视觉多样化的作品仍是人工主导、耗时费力的过程,限制了创意工作流的扩展性和个性化。我们提出一个基于解码器仅语言模型的系统——创意预训练变换器(CPT),用于生成可编辑的设计变体。核心是新提出的创意标记语言(CML),一种紧凑、机器学习友好的格式,能捕捉画布级结构、页面布局及元素级细节(文本、图像、矢量图),包括内容与风格信息。我们在大量由专业设计师创作的设计模板上对CPT进行微调,使其能够学习颜色搭配、字体选择等属性的语境感知预测。模型生成语义结构清晰且风格一致的输出,保持元素间内部一致性。与生成图像的模型不同,本系统产出的是完全可编辑的设计文档,而非像素图像,使用户可在设计编辑器中自由修改与定制。实验表明,该方法能为现有模板生成符合上下文的颜色与字体变体,并在保持设计原则的前提下调整布局,展现出良好潜力。
原文摘要 · Abstract (English)
Designing visually diverse and high-quality designs remains a manual, time-consuming process, limiting scalability and personalization in creative workflows. We present a system for generating editable design variations using a decoder-only language model, the Creative Pre-trained Transformer (CPT), trained to predict visual style attributes in design templates. At the core of our approach is a new representation called Creative Markup Language (CML), a compact, machine-learning-friendly format that captures canvas-level structure, page layout, and element-level details (text, images, and vector graphics), including both content and style. We fine-tune CPT on a large corpus of design templates authored by professional designers, enabling it to learn meaningful, context-aware predictions for attributes such as color schemes and font choices. The model produces semantically structured and stylistically coherent outputs, preserving internal consistency across elements. Unlike generative image models, our system yields fully editable design documents rather than pixel-only images, allowing users to iterate and personalize within a design editor. In experiments, our approach generates contextual color and font variations for existing templates and shows promise in adjusting layouts while maintaining design principles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。