用笔画级编码器提升场景文字编辑的精度与真实感
GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editing
- 设计笔画注意力模块,捕捉笔画到字符的跨层级结构关系
- 在多语言场景文本编辑中实现18.02%句级准确率提升
- 适合需要高保真文字生成的图像编辑与广告设计场景
场景文本编辑需在保持图像风格一致性和环境视觉连贯性的前提下修改图像中的文字。尽管基于扩散模型的方法在文本生成方面展现出潜力,但其仍难以生成高质量结果,尤其在处理如中文等复杂文字时易产生扭曲或不可识别的字符。这些文字由精细的笔画结构和空间关系构成,必须精确保留。本文提出GlyphMastero,一种专为笔画级精准控制设计的字形编码器,引导潜在扩散模型生成具有笔画级精度的文字。关键洞察在于:现有方法虽使用预训练OCR提取特征,却未能捕捉从单个笔画到笔画间交互再到整体字符结构的层次化特性。为此,我们引入新型笔画注意力模块,显式建模局部字符与全局文本行之间的跨层级交互;同时,采用特征金字塔网络融合多尺度OCR骨干特征以增强全局表征。通过跨层级与多尺度融合,获得更细致的字形感知引导,实现对场景文本生成过程的精准控制。实验表明,本方法在多语言场景文本编辑基准上相较最先进方法提升18.02%句级准确率,同时将文本区域弗雷歇起始距离降低53.28%。
原文摘要 · Abstract (English)
Scene text editing, a subfield of image editing, requires modifying texts in images while preserving style consistency and visual coherence with the surrounding environment. While diffusion-based methods have shown promise in text generation, they still struggle to produce high-quality results. These methods often generate distorted or unrecognizable characters, particularly when dealing with complex characters like Chinese. In such systems, characters are composed of intricate stroke patterns and spatial relationships that must be precisely maintained. We present GlyphMastero, a specialized glyph encoder designed to guide the latent diffusion model for generating texts with stroke-level precision. Our key insight is that existing methods, despite using pretrained OCR models for feature extraction, fail to capture the hierarchical nature of text structures - from individual strokes to stroke-level interactions to overall character-level structure. To address this, our glyph encoder explicitly models and captures the cross-level interactions between local-level individual characters and global-level text lines through our novel glyph attention module. Meanwhile, our model implements a feature pyramid network to fuse the multi-scale OCR backbone features at the global-level. Through these cross-level and multi-scale fusions, we obtain more detailed glyph-aware guidance, enabling precise control over the scene text generation process. Our method achieves an 18.02\% improvement in sentence accuracy over the state-of-the-art multi-lingual scene text editing baseline, while simultaneously reducing the text-region Fréchet inception distance by 53.28\%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。