提升商品海报文字编辑的准确性与视觉一致性。
TextRefine: Improving Textual Fidelity, Spatial Placement, and Glyph Rendering for Text Editing in Product Posters

- 分阶段优化文字插入与替换的语义、位置与字形表现。
- 在10万张海报数据上,显著降低文字遗漏与错位率。
- 适合电商设计、广告生成等需要精准文本编辑的场景。
商品海报中的文字编辑需在不破坏产品外观、背景内容和整体构图的前提下,插入或替换文字。尽管指令式图像编辑已有进展,通用模型在此任务中仍不可靠:常遗漏或错误渲染目标文字,将其放置在关键产品或已有内容之上,导致字形结构扭曲或视觉不一致。我们提出 extbf{TextRefine},一种任务对齐的后训练框架,结合监督微调与操作特异性奖励优化,解决这些互补性失败模式。针对文字插入,文本段级奖励联合评估语义保真度与目标跨度覆盖,惩罚与产品及现有文字的空间冲突,并使用门控结构约束保持非文本区域;针对文字替换,字形级奖励利用目标字符的连接时序分类(CTC)后验提供细粒度监督,捕捉缺失笔画、结构变形及相似字符混淆等缺陷。我们还引入 extbf{OpenTextEdit},一个包含10万张图像的数据集,涵盖多文字布局、详细文本属性、产品掩码及低频字符。大量实验表明,TextRefine 在文字保真度、位置可靠性与字形质量上均优于对比基线,同时更好保留源图像内容。
原文摘要 · Abstract (English)
Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition. Despite recent progress in instruction-based image editing, general-purpose models remain unreliable in this setting: they often omit or incorrectly render the target text, place it over salient products or pre-existing content, and produce structurally distorted or visually inconsistent glyphs. We introduce \textbf{TextRefine}, a task-aligned post-training framework that combines supervised fine-tuning with operation-specific reward optimization to address these complementary failure modes. For text insertion, our text-span-level reward jointly assesses semantic fidelity and target-span coverage, penalizes spatial conflicts with products and existing text, and employs a gated structural constraint to preserve non-text regions. For text replacement, our glyph-level reward leverages the connectionist temporal classification (CTC) posterior of the target character to provide graded supervision for fine-grained defects, including missing strokes, structural deformations, and confusion among visually similar characters. We further introduce \textbf{OpenTextEdit}, a dataset comprising 100K images for text editing in product posters, with multi-text layouts, detailed text attributes, product masks, and challenging low-frequency characters. Extensive experiments on both insertion and replacement demonstrate that TextRefine consistently outperforms the evaluated image editing baselines in textual fidelity, placement reliability, and glyph quality while better preserving source-image content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。