arXiv:2409.19051cs.CVcs.AI2024-09中稿 · ACM Multimedia 202…被引 17

用混合标记与图像表示设计文档,实现统一的自动补全。

Multimodal Markup Document Models for Graphic Design Completion

  • 将设计拆分为标记语言与图像交织的文档结构。
  • 在三个任务中均生成符合上下文的合理设计内容。
  • 适合需要文本与图像协同生成的设计自动化场景。

我们提出MarkupDM,一种多模态标记文档模型,将图形设计表示为标记语言与图像交错的文档。不同于依赖元素-属性网格的全局方法,该表示支持可变长度元素、类型相关的属性和文本内容。受代码生成中‘填空式’训练启发,模型通过上下文补全缺失部分,统一处理多种设计任务。模型还支持图像生成,通过专用分词器预测离散图像标记,支持透明度。我们在三个任务(属性值、图像、文本补全)上评估,结果表明其能生成与上下文一致的合理设计。进一步在新的指令引导设计补全任务中测试,微调后的MarkupDM在文本补全方面优于当前最优图像编辑模型。这些发现表明,结合该文档表示的多模态语言模型可作为广泛设计自动化的通用基础。

原文摘要 · Abstract (English)

We introduce MarkupDM, a multimodal markup document model that represents graphic design as an interleaved multimodal document consisting of both markup language and images. Unlike existing holistic approaches that rely on an element-by-attribute grid representation, our representation accommodates variable-length elements, type-dependent attributes, and text content. Inspired by fill-in-the-middle training in code generation, we train the model to complete the missing part of a design document from its surrounding context, allowing it to treat various design tasks in a unified manner. Our model also supports image generation by predicting discrete image tokens through a specialized tokenizer with support for image transparency. We evaluate MarkupDM on three tasks, attribute value, image, and text completion, and demonstrate that it can produce plausible designs consistent with the given context. To further illustrate the flexibility of our approach, we evaluate our approach on a new instruction-guided design completion task where our instruction-tuned MarkupDM compares favorably to state-of-the-art image editing models, especially in textual completion. These findings suggest that multimodal language models with our document representation can serve as a versatile foundation for broad design automation.

图形设计多模态文档建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。