自动设计能表达语义又易读的字体,支持多语言
Khattat: Enhancing Readability and Concept Representation of Semantic Typography
- 用大模型生成创意,用FontCLIP选字体,扩散模型局部变形
- 引入OCR损失函数,实现多字同时美化且保持可读性
- 适合跨语言字体设计,尤其抽象概念如'自由'的视觉化
设计既能视觉传达词义又保持可读性的表达性字体是一项复杂任务,称为语义排版。它涉及构思概念、选择合适字体,并在创意与易读性之间取得平衡。我们提出一个端到端系统来自动化该过程:首先,大语言模型(LLM)为词语生成视觉构思,适用于抽象概念如‘自由’;接着,基于预训练的FontCLIP模型,根据字体语义属性自动选择合适字体;系统识别词语中适合变形的区域,并使用预训练扩散模型进行迭代变形;关键创新在于提出的基于OCR的损失函数,显著提升可读性,并支持多个字符同时风格化。我们在多种语言和书写系统上对比基线方法,验证了本方法在可读性增强和跨语言适用性方面的优越表现。
原文摘要 · Abstract (English)
Designing expressive typography that visually conveys a word's meaning while maintaining readability is a complex task, known as semantic typography. It involves selecting an idea, choosing an appropriate font, and balancing creativity with legibility. We introduce an end-to-end system that automates this process. First, a Large Language Model (LLM) generates imagery ideas for the word, useful for abstract concepts like freedom. Then, the FontCLIP pre-trained model automatically selects a suitable font based on its semantic understanding of font attributes. The system identifies optimal regions of the word for morphing and iteratively transforms them using a pre-trained diffusion model. A key feature is our OCR-based loss function, which enhances readability and enables simultaneous stylization of multiple characters. We compare our method with other baselines, demonstrating great readability enhancement and versatility across multiple languages and writing scripts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。