让手写体生成更可靠,减少重复与失真。
Autoregressive Styled Text Image Generation, but Make it Reliable
- 将生成任务改为多模态提示控制,用特殊文本标记对齐视觉内容。
- 相比之前方法,输入更少、风格泛化更强、文字内容更忠实于提示。
- 适合需要高内容准确性的手写体图像生成场景。
生成真实可读的手写体图像(尤其在手写体生成领域)仍是一个开放问题,广泛应用于平面设计、文档理解与图像编辑。现有研究致力于复现特定作者的书写风格,近期基于自回归变压器的方法在风格保真度和泛化性上表现良好。但该方法需额外输入,缺乏有效终止机制,易陷入重复循环并产生视觉伪影。本文重新思考自回归建模方式,将手写体生成视为多模态提示条件生成任务,引入特殊文本输入标记以增强与视觉标记的对齐。此外,提出基于无分类器引导的策略优化自回归模型。大量实验证明,所提方法Eruku相较以往方案所需输入更少,对未见风格泛化能力更强,且更忠实地遵循文本提示,显著提升内容一致性。
原文摘要 · Abstract (English)
Generating faithful and readable styled text images (especially for Styled Handwritten Text generation - HTG) is an open problem with several possible applications across graphic design, document understanding, and image editing. A lot of research effort in this task is dedicated to developing strategies that reproduce the stylistic characteristics of a given writer, with promising results in terms of style fidelity and generalization achieved by the recently proposed Autoregressive Transformer paradigm for HTG. However, this method requires additional inputs, lacks a proper stop mechanism, and might end up in repetition loops, generating visual artifacts. In this work, we rethink the autoregressive formulation by framing HTG as a multimodal prompt-conditioned generation task, and tackle the content controllability issues by introducing special textual input tokens for better alignment with the visual ones. Moreover, we devise a Classifier-Free-Guidance-based strategy for our autoregressive model. Through extensive experimental validation, we demonstrate that our approach, dubbed Eruku, compared to previous solutions requires fewer inputs, generalizes better to unseen styles, and follows more faithfully the textual prompt, improving content adherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。