arXiv:2506.05843cs.CV2025-06被引 1

让图像中的文字秒变任意字体,无需训练

FontAdapter: Instant Font Adaptation in Visual Text Generation

  • 分两阶段学习字体特征,先提取字形风格再融合背景
  • 生成新字体仅需几秒,支持未见过的字体
  • 可实现字体编辑、混合与跨语言迁移

文本生成图像的扩散模型已显著提升视觉文字在各类图像中的自然融合。现有方法通过预定义字体字典微调来增强字体控制,但适应字典外的新字体计算成本高,常需数十分钟,难以实现实时定制。本文提出FontAdapter,一种仅需参考字形图像即可在数秒内生成新字体的框架。我们发现直接在字体数据集上训练无法捕捉细腻的字体属性,限制了对新字形的泛化能力。为此,提出两阶段课程学习:首先从孤立字形中提取字体特征,再将其融入多样自然背景。为此构建了针对各阶段的合成数据集,有效利用大规模在线字体资源。实验表明,FontAdapter可在无需推理时额外微调的情况下,高质量、鲁棒地实现未见字体的定制。同时支持视觉文字编辑、字体风格融合与跨语言字体迁移,是通用的字体定制框架。

原文摘要 · Abstract (English)

Text-to-image diffusion models have significantly improved the seamless integration of visual text into diverse image contexts. Recent approaches further improve control over font styles through fine-tuning with predefined font dictionaries. However, adapting unseen fonts outside the preset is computationally expensive, often requiring tens of minutes, making real-time customization impractical. In this paper, we present FontAdapter, a framework that enables visual text generation in unseen fonts within seconds, conditioned on a reference glyph image. To this end, we find that direct training on font datasets fails to capture nuanced font attributes, limiting generalization to new glyphs. To overcome this, we propose a two-stage curriculum learning approach: FontAdapter first learns to extract font attributes from isolated glyphs and then integrates these styles into diverse natural backgrounds. To support this two-stage training scheme, we construct synthetic datasets tailored to each stage, leveraging large-scale online fonts effectively. Experiments demonstrate that FontAdapter enables high-quality, robust font customization across unseen fonts without additional fine-tuning during inference. Furthermore, it supports visual text editing, font style blending, and cross-lingual font transfer, positioning FontAdapter as a versatile framework for font customization tasks.

文本生成字体迁移扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。