用视觉语言模型自动生成简洁有力的视觉表达,提升信息传达效率。
Generative Visual Communication in the Era of Vision-Language Models
- 限定模型输出空间并引入任务特定正则化,提升设计灵活性。
- 在草图、字体、动画等场景中实现从复杂概念到抽象视觉的转化。
- 适合需要快速生成高质量视觉内容的设计与创意工作者。
视觉传播自史前岩画起便存在,是通过视觉元素传递思想与信息的方式。在当今视觉信息泛滥的时代,有效设计需融合图形设计原则、视觉叙事、人类心理及将复杂信息提炼为清晰视觉的能力。本论文探讨如何利用近期视觉语言模型(VLMs)自动化生成高效视觉传播设计。尽管生成模型在文本生成图像方面取得显著进展,但仍难以将复杂概念简化为清晰、抽象的视觉表达,且受限于像素级输出,缺乏对多种设计任务的灵活性。为此,本文限制模型操作空间并引入任务特定正则化。研究涵盖视觉传播多个方面:草图与视觉抽象、排版、动画及视觉灵感生成。
原文摘要 · Abstract (English)
Visual communication, dating back to prehistoric cave paintings, is the use of visual elements to convey ideas and information. In today's visually saturated world, effective design demands an understanding of graphic design principles, visual storytelling, human psychology, and the ability to distill complex information into clear visuals. This dissertation explores how recent advancements in vision-language models (VLMs) can be leveraged to automate the creation of effective visual communication designs. Although generative models have made great progress in generating images from text, they still struggle to simplify complex ideas into clear, abstract visuals and are constrained by pixel-based outputs, which lack flexibility for many design tasks. To address these challenges, we constrain the models' operational space and introduce task-specific regularizations. We explore various aspects of visual communication, namely, sketches and visual abstraction, typography, animation, and visual inspiration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。