用视觉Transformer一键生成多语言字体,支持从未见过的字形。
One-Shot Multilingual Font Generation Via ViT
- 基于ViT和掩码自编码预训练,无需复杂设计模块
- 仅需一个样本即可生成高质量多语言字体,支持用户自创字形
- 引入检索增强引导模块,提升风格迁移与实际应用能力
汉字、日文、韩文等表意文字的字体设计面临独特挑战,因需手工绘制数千个独立字符。本文提出一种基于视觉变压器(ViT)的多语言字体生成模型,有效应对表意与拼音文字的复杂性。通过采用强视觉预训练任务(掩码自编码,MAE),模型无需依赖先前框架中的复杂设计组件,即可实现全面且泛化能力强的生成效果。令人瞩目的是,该模型能为未见过、未知甚至用户自创的字符生成高质量多语言字体。此外,我们集成检索增强引导(RAG)模块,动态检索并适配风格参考,显著提升可扩展性与实际应用价值。我们在多种字体生成任务中验证了该方法的有效性、适应性与可扩展性。
原文摘要 · Abstract (English)
Font design poses unique challenges for logographic languages like Chinese, Japanese, and Korean (CJK), where thousands of unique characters must be individually crafted. This paper introduces a novel Vision Transformer (ViT)-based model for multi-language font generation, effectively addressing the complexities of both logographic and alphabetic scripts. By leveraging ViT and pretraining with a strong visual pretext task (Masked Autoencoding, MAE), our model eliminates the need for complex design components in prior frameworks while achieving comprehensive results with enhanced generalizability. Remarkably, it can generate high-quality fonts across multiple languages for unseen, unknown, and even user-crafted characters. Additionally, we integrate a Retrieval-Augmented Guidance (RAG) module to dynamically retrieve and adapt style references, improving scalability and real-world applicability. We evaluated our approach in various font generation tasks, demonstrating its effectiveness, adaptability, and scalability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。