解决场景文字编辑中风格不一致和长度变化难题
Global-Local Aware Scene Text Editing
- 融合全局上下文与局部特征,提升编辑一致性
- 支持任意长度文本编辑,保持风格与背景和谐
- 风格向量可迁移,适配不同尺寸目标文本
场景文字编辑(STE)旨在替换图像中的文字内容,同时保留原有文字风格和背景纹理。现有方法面临两大挑战:编辑区域与周围环境的不一致,以及对文本长度变化敏感。为此,本文提出端到端的全局-局部感知场景文字编辑框架GLASTE,通过设计全局-局部组合结构、联合损失函数,并增强文本图像特征,实现局部区域内风格一致性及局部与全局间的协调性。此外,将文本风格表示为与图像尺寸无关的向量,可迁移至不同尺寸的目标文本图像。采用仿射融合方式填充目标文本图像,保持其长宽比不变。在真实世界数据集上的大量实验表明,GLASTE在定量指标和定性结果上均优于现有方法,有效缓解了上述两个核心问题。
原文摘要 · Abstract (English)
Scene Text Editing (STE) involves replacing text in a scene image with new target text while preserving both the original text style and background texture. Existing methods suffer from two major challenges: inconsistency and length-insensitivity. They often fail to maintain coherence between the edited local patch and the surrounding area, and they struggle to handle significant differences in text length before and after editing. To tackle these challenges, we propose an end-to-end framework called Global-Local Aware Scene Text Editing (GLASTE), which simultaneously incorporates high-level global contextual information along with delicate local features. Specifically, we design a global-local combination structure, joint global and local losses, and enhance text image features to ensure consistency in text style within local patches while maintaining harmony between local and global areas. Additionally, we express the text style as a vector independent of the image size, which can be transferred to target text images of various sizes. We use an affine fusion to fill target text images into the editing patch while maintaining their aspect ratio unchanged. Extensive experiments on real-world datasets validate that our GLASTE model outperforms previous methods in both quantitative metrics and qualitative results and effectively mitigates the two challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。