arXiv:2603.22054cs.CV2026-03被引 1

用视觉元素驱动字体生成,实现高保真风格控制。

FontCrafter: High-Fidelity Element-Driven Artistic Font Creation with Visual In-Context Generation

  • 以元素为视觉上下文,通过像素级风格迁移生成字体。
  • 零样本生成下保持结构与纹理高保真,支持风格混合。
  • 适合需要精细风格控制的设计师和创意工作者。

艺术字体生成旨在基于参考风格合成样式化字形。然而,现有方法存在风格多样性有限和控制粗糙的问题。本文探索了以元素驱动的艺术字体生成,将元素视为字体的基本视觉单元,作为目标风格的参考图像。我们将其分为具象元素(如花朵、石头)和无定形元素(如火焰、云朵)。提出 FontCrafter 框架,并构建大规模数据集 ElementFont,包含多样元素类型与高质量字形图像。为实现元素纹理与结构的高保真重建,引入视觉上下文生成策略,将元素图像作为视觉上下文,利用图像修复模型在像素级转移风格。设计轻量级上下文感知掩码适配器(CMA)注入形状信息,提出免训练注意力重定向机制,实现区域感知风格控制并抑制笔画幻觉。此外,采用边缘重绘使边界更自然。大量实验表明,FontCrafter 在零样本生成中表现优异,尤其在结构与纹理保真度方面,同时支持灵活的风格混合控制。

原文摘要 · Abstract (English)

Artistic font generation aims to synthesize stylized glyphs based on a reference style. However, existing approaches suffer from limited style diversity and coarse control. In this work, we explore the potential of element-driven artistic font generation. Elements are the fundamental visual units of a font, serving as reference images for the desired style. Conceptually, we categorize elements into object elements (e.g., flowers or stones) with distinct structures and amorphous elements (e.g., flames or clouds) with unstructured textures. We introduce FontCrafter, an element-driven framework for font creation, and construct a large-scale dataset, ElementFont, which contains diverse element types and high-quality glyph images. However, achieving high-fidelity reconstruction of both texture and structure of reference elements remains challenging. To address this, we propose an in-context generation strategy that treats element images as visual context and uses an inpainting model to transfer element styles into glyph regions at the pixel level. To further control glyph shapes, we design a lightweight Context-aware Mask Adapter (CMA) that injects shape information. Moreover, a training-free attention redirection mechanism enables region-aware style control and suppresses stroke hallucination. In addition, edge repainting is applied to make boundaries more natural. Extensive experiments demonstrate that FontCrafter achieves strong zero-shot generation performance, particularly in preserving structural and textural fidelity, while also supporting flexible controls such as style mixture.

字体生成风格控制视觉上下文图像修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。