用全局感知的自回归模型,让文字生成更忠实于风格意图。
Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation
- 引入全局感知分词器,同时捕捉局部结构与整体风格模式。
- 通过轻量语言适配器实现文本风格控制,无需大规模多模态预训练。
- 适合需要精准风格还原的字体设计场景,尤其支持文本引导生成。
手动字体设计是将视觉风格概念转化为连贯字形集的复杂过程。这一挑战在少样本字体生成(FFG)中依然存在,现有模型常难以在有限参考下同时保持结构完整性和风格一致性。尽管自回归模型具备强大生成能力,但其在FFG中的应用受限于传统的局部分块标记化方法,忽略了对连贯字体合成至关重要的全局依赖关系。此外,现有方法仍局限于仅依赖视觉参考的图像到图像范式,忽视了语言在传达字体设计风格意图中的作用。为此,我们提出GAR-Font,一种新颖的多模态少样本字体生成自回归框架。GAR-Font引入全局感知分词器,有效捕捉局部结构与全局风格模式;设计轻量级语言-风格适配器,实现灵活风格控制,无需高强度多模态预训练;并构建后置优化流程,进一步提升结构准确性和风格一致性。大量实验表明,GAR-Font在保持全局风格忠实度方面优于现有方法,且在文本风格引导下生成质量更高。
原文摘要 · Abstract (English)
Manual font design is an intricate process that transforms a stylistic visual concept into a coherent glyph set. This challenge persists in automated Few-shot Font Generation (FFG), where models often struggle to preserve both the structural integrity and stylistic fidelity from limited references. While autoregressive (AR) models have demonstrated impressive generative capabilities, their application to FFG is constrained by conventional patch-level tokenization, which neglects global dependencies crucial for coherent font synthesis. Moreover, existing FFG methods remain within the image-to-image paradigm, relying solely on visual references and overlooking the role of language in conveying stylistic intent during font design. To address these limitations, we propose GAR-Font, a novel AR framework for multimodal few-shot font generation. GAR-Font introduces a global-aware tokenizer that effectively captures both local structures and global stylistic patterns, a multimodal style encoder offering flexible style control through a lightweight language-style adapter without requiring intensive multimodal pretraining, and a post-refinement pipeline that further enhances structural fidelity and style coherence. Extensive experiments show that GAR-Font outperforms existing FFG methods, excelling in maintaining global style faithfulness and achieving higher-quality results with textual stylistic guidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。