arXiv:2412.17225cs.CV2024-12被引 14

CharGen通过多模态编码提升文字生成精度,尤其擅长复杂汉字的准确渲染。

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder

  • 采用字符级多模态编码器,同步提取文字与字形特征
  • 引入感知损失,使生成文字笔画更精准,错误率下降超8%
  • 可嵌入现有扩散模型,中文文字生成准确率提升5.5%

近年来,基于扩散模型的视觉文本生成取得显著进展。尽管这些方法在文本渲染效果上不断提升,但在生成复杂视觉文字时仍存在字符与笔画不准确的问题。本文提出CharGen,一种高精度字符级视觉文本生成与编辑模型。CharGen采用字符级多模态编码器,不仅能提取字符级文本嵌入,还可逐字符编码字形图像,从而更有效地捕捉细粒度跨模态特征。此外,我们引入新的感知损失,强化字符形状监督,解决生成文本中笔画不准的问题。值得注意的是,CharGen可无缝集成至现有扩散模型中,实现高精度视觉文本生成。在AnyText-benchmark和MARIO-Eval等公开基准测试中,其文本渲染准确率分别提升超过8%和6%,在中文测试集上更是达到5.5%的准确率提升。

原文摘要 · Abstract (English)

Recently, significant advancements have been made in diffusion-based visual text generation models. Although the effectiveness of these methods in visual text rendering is rapidly improving, they still encounter challenges such as inaccurate characters and strokes when rendering complex visual text. In this paper, we propose CharGen, a highly accurate character-level visual text generation and editing model. Specifically, CharGen employs a character-level multimodal encoder that not only extracts character-level text embeddings but also encodes glyph images character by character. This enables it to capture fine-grained cross-modality features more effectively. Additionally, we introduce a new perceptual loss in CharGen to enhance character shape supervision and address the issue of inaccurate strokes in generated text. It is worth mentioning that CharGen can be integrated into existing diffusion models to generate visual text with high accuracy. CharGen significantly improves text rendering accuracy, outperforming recent methods in public benchmarks such as AnyText-benchmark and MARIO-Eval, with improvements of more than 8% and 6%, respectively. Notably, CharGen achieved a 5.5% increase in accuracy on Chinese test sets.

文本生成扩散模型多模态中文生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。