用生成式智能体实现少样本多语言字体迁移,支持中日韩等复杂文字。
LoGAN: Multilingual Font Localization with Generative Agents

- 基于视觉语言模型的智能体框架,分步处理字形、风格、间距与纹理迁移。
- 在27种语言上实现高保真字形生成,风格与字距一致性优于主流图像编辑模型。
- 适合需要跨语言字体设计的工业应用,尤其擅长中日韩等复杂文字系统。
将字体适配至新语言是一项复杂任务,需精准调整字形、颜色/纹理及间距/字距,现有方法多仅限单字形生成且难以支持多语言渲染。本文提出LoGAN,一种基于视觉语言模型的生成式智能体框架,可在少量样本下完成多语言字体本地化,输入少量原始字形或标志字母即可生成目标语言完整字符集。该框架包含字形级扩散模型、风格微调模块、间距与字距迁移算法及纹理扩展模型,并由视觉语言模型智能体统一协调。在涵盖超过27种语言的字体与真实标志数据集上评估,相比专门的字体生成方法及先进图像编辑模型(如FLUX、Nano-Banana),LoGAN在定量与定性评价中均展现出更高字形保真度,同时保持更优的风格、纹理与字距一致性,适用于包括中文、日文、韩文在内的多种语言系统。
原文摘要 · Abstract (English)
Localizing a font into new languages is a highly intricate task requiring precise design adaptation of glyphs, color/texture, and spacing/kerning, from source to target languages. Most existing methods focus on single glyph generation with limited capability in handling multilingual font rendering. In this work, we propose LoGAN, a VLM-based agentic framework for few-shot multilingual font localization, which takes in a small number of individual glyphs from a font or letters from a logo and uses them to generate complete character sets in other languages. LoGAN breaks down this task into multiple components: a glyph-level diffusion model, a style finetuning module, a spacing and kerning transfer algorithm, and a texture expansion model, with a VLM agent coordinator. LoGAN achieves broad language coverage for font localization with various styles, including Chinese/Korean/Japanese (CJK). We evaluate our approach on both font and real-world logo datasets spanning more than 27 languages and compare it against both specialized font generation and state-of-the-art image editing models with strong text rendering capabilities (e.g., FLUX, Nano-Banana). Our approach yields higher glyph fidelity while maintaining better style, texture, and kerning consistency according to both quantitative and qualitative evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。