arXiv:2410.10133cs.CV2024-10NeurIPS被引 51

通过引导控制实现文本编辑中风格与内容的精准平衡

TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control

  • 引入风格-结构双重引导,提升文本样式一致性
  • 使用自适应互注意力机制增强图像风格保真度
  • 构建首个真实场景图文对数据集,支持公平评估

尽管近期在文本到图像生成和文本驱动图像操作方面取得显著进展,场景文本编辑(STE)仍面临内容修改与风格保持的挑战。基于GAN的方法普遍存在泛化能力差的问题,而基于扩散模型的方法则容易产生风格偏移。为此,我们提出TextCtrl,一种基于扩散模型并引入先验引导控制的文本编辑方法。该方法包含两个关键组件:(i) 通过构建细粒度文本风格解耦与鲁棒的文本字形结构表示,将风格-结构引导显式融入模型设计与训练,显著提升文本风格一致性和渲染精度;(ii) 提出字形自适应互注意力机制,解构源图像的隐含细粒度特征,从而在推理过程中增强风格一致性与视觉质量。此外,为填补真实场景文本编辑评估基准的空白,我们构建了首个真实世界图像对数据集ScenePair,用于公平比较。实验表明,TextCtrl在风格保真度与文本准确性方面均优于现有方法。

原文摘要 · Abstract (English)

Centred on content modification and style preservation, Scene Text Editing (STE) remains a challenging task despite considerable progress in text-to-image synthesis and text-driven image manipulation recently. GAN-based STE methods generally encounter a common issue of model generalization, while Diffusion-based STE methods suffer from undesired style deviations. To address these problems, we propose TextCtrl, a diffusion-based method that edits text with prior guidance control. Our method consists of two key components: (i) By constructing fine-grained text style disentanglement and robust text glyph structure representation, TextCtrl explicitly incorporates Style-Structure guidance into model design and network training, significantly improving text style consistency and rendering accuracy. (ii) To further leverage the style prior, a Glyph-adaptive Mutual Self-attention mechanism is proposed which deconstructs the implicit fine-grained features of the source image to enhance style consistency and vision quality during inference. Furthermore, to fill the vacancy of the real-world STE evaluation benchmark, we create the first real-world image-pair dataset termed ScenePair for fair comparisons. Experiments demonstrate the effectiveness of TextCtrl compared with previous methods concerning both style fidelity and text accuracy.

文本编辑扩散模型风格保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。