TextMaster实现高精度文本编辑,支持风格可控与布局自适应。
TextMaster: A Unified Framework for Realistic Text Editing via Glyph-Style Dual-Control
- 引入字形与风格双控机制,提升复杂文本渲染精度
- 通过感知损失与边界框回归损失,保证文本布局准确
- 支持任意图像区域的风格可调文本编辑,适合视觉设计场景
在图像编辑任务中,高质量的文本编辑能力能显著降低人力与资源成本。现有方法在复杂文本的笔画精度和生成风格的可控性方面存在明显局限。为此,我们提出TextMaster,一种可在多种场景和图像区域中精确编辑文本,同时保证布局合理与风格可控的统一框架。该方法通过引入高分辨率标准字形信息,并在文本编辑区域应用感知损失,提升了文本渲染的准确性和保真度。此外,利用注意力机制计算每个字符的中间层边界框回归损失,使模型能够学习不同上下文中的文本布局。进一步提出一种新颖的风格注入技术,实现注入文本的可控风格迁移。通过全面实验,验证了该方法在多项指标上达到当前最优性能。
原文摘要 · Abstract (English)
In image editing tasks, high-quality text editing capabilities can significantly reduce both human and material resource costs. Existing methods, however, face significant limitations in terms of stroke accuracy for complex text and controllability of generated text styles. To address these challenges, we propose TextMaster, a solution capable of accurately editing text across various scenarios and image regions, while ensuring proper layout and controllable text style. Our method enhances the accuracy and fidelity of text rendering by incorporating high-resolution standard glyph information and applying perceptual loss within the text editing region. Additionally, we leverage an attention mechanism to compute intermediate layer bounding box regression loss for each character, enabling the model to learn text layout across varying contexts. Furthermore, we propose a novel style injection technique that enables controllable style transfer for the injected text. Through comprehensive experiments, we demonstrate the state-of-the-art performance of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。