可控生成多语言艺术字,细节清晰无模糊
AnyArtisticGlyph: Multilingual Controllable Artistic Glyph Generation
- 基于扩散模型,融合字体与图文特征实现精细控制
- 在多种语言下生成高保真艺术字,细节真实无失真
- 适合需要精准文字视觉设计的创作者与设计师
艺术字图像生成(AGIG)不同于现有以创意为主的生成模型,其核心在于实现可精确控制的确定性生成,将参考图像风格迁移至源文字的同时保持内容不变。尽管前景广阔,现有方法在细节层面常出现模糊或纹理错误等问题。为此,我们提出 AnyArtisticGlyph——一种基于扩散模型的多语言可控艺术字生成模型。该模型包含字体融合与嵌入模块,用于生成细节结构的潜在特征;以及视觉-文本融合与嵌入模块,利用 CLIP 模型编码参考图像,并将其与变换提示词嵌入融合,实现全局图像的无缝生成。此外,引入粗粒度特征级损失以提升生成精度。实验表明,该模型在多种语言上均能生成自然、细节丰富的艺术字图像,性能达到当前最优水平。项目代码已开源至 https://github.com/jiean001/AnyArtisticGlyph,旨在推动文本生成技术发展。
原文摘要 · Abstract (English)
Artistic Glyph Image Generation (AGIG) differs from current creativity-focused generation models by offering finely controllable deterministic generation. It transfers the style of a reference image to a source while preserving its content. Although advanced and promising, current methods may reveal flaws when scrutinizing synthesized image details, often producing blurred or incorrect textures, posing a significant challenge. Hence, we introduce AnyArtisticGlyph, a diffusion-based, multilingual controllable artistic glyph generation model. It includes a font fusion and embedding module, which generates latent features for detailed structure creation, and a vision-text fusion and embedding module that uses the CLIP model to encode references and blends them with transformation caption embeddings for seamless global image generation. Moreover, we incorporate a coarse-grained feature-level loss to enhance generation accuracy. Experiments show that it produces natural, detailed artistic glyph images with state-of-the-art performance. Our project will be open-sourced on https://github.com/jiean001/AnyArtisticGlyph to advance text generation technology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。