用扩散模型实现自由风格文字图像定制,精准还原笔触与字形。
Calligrapher: Freestyle Text Image Customization
- 自蒸馏构建风格基准,减少对人工标注数据依赖。
- 局部风格注入+上下文生成,实现参考图与文字的精细对齐。
- 适合数字艺术、品牌设计等需要高质量定制字体的场景。
我们提出Calligrapher,一种基于扩散模型的新框架,将高级文本定制与艺术排版结合,应用于数字书法与设计。针对排版定制中风格控制精度低和数据依赖强的问题,提出三项关键技术:首先,利用预训练文本到图像生成模型与大语言模型自蒸馏,自动构建以风格为中心的排版基准;其次,引入可训练风格编码器(包含Qformer与线性层),从参考图像中提取鲁棒风格特征,并通过上下文生成机制直接嵌入去噪过程,强化目标风格对齐;在多种字体与设计场景下,定量与定性评估均验证其能精准复现复杂风格细节与精确字形定位。该框架实现高质量、视觉一致的排版自动化,超越传统模型,赋能数字艺术、品牌设计与情境化排版创作。
原文摘要 · Abstract (English)
We introduce Calligrapher, a novel diffusion-based framework that innovatively integrates advanced text customization with artistic typography for digital calligraphy and design applications. Addressing the challenges of precise style control and data dependency in typographic customization, our framework incorporates three key technical contributions. First, we develop a self-distillation mechanism that leverages the pre-trained text-to-image generative model itself alongside the large language model to automatically construct a style-centric typography benchmark. Second, we introduce a localized style injection framework via a trainable style encoder, which comprises both Qformer and linear layers, to extract robust style features from reference images. An in-context generation mechanism is also employed to directly embed reference images into the denoising process, further enhancing the refined alignment of target styles. Extensive quantitative and qualitative evaluations across diverse fonts and design contexts confirm Calligrapher's accurate reproduction of intricate stylistic details and precise glyph positioning. By automating high-quality, visually consistent typography, Calligrapher surpasses traditional models, empowering creative practitioners in digital art, branding, and contextual typographic design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。