arXiv:2606.05730cs.CV2026-06

一个模型搞定场景文字删除、生成和替换,精准控制样式与背景。

TextWand: A Unified Framework for Scene Text Editing

论文配图:TextWand: A Unified Framework for Scene Text Editing
图 1 · 摘自论文原文
  • 将文字编辑拆解为渲染与擦除两个基本操作,实现精细控制。
  • 在多个任务中超越现有开源与闭源模型,保持布局一致性和图像质量。
  • 新设计的定位编码与抑制策略提升像素级精度,适合多场景文字处理。

我们提出 TextWand,一个通用框架,将场景文字删除、生成和替换统一于单一模型。通过将复杂编辑任务分解为渲染与擦除两个原子操作,TextWand 实现对文字外观和背景完整性的精确控制。具体而言,引入一种新型设计——覆盖参考位置编码(Overlay-Reference Positional Encoding, ORPE),以保证像素级版式保真度和示例驱动的风格控制;同时提出区域自适应抑制策略(Region-Adaptive Suppression, RAS),确保干净的文字擦除。针对现有单任务数据集缺乏综合性基准的问题,我们构建了 TextWand-Bench。大量实验证明,TextWand 在场景文字删除、生成与替换任务中均优于当前领先的开源与闭源模型,在文本内容准确性、版式与风格一致性及整体图像质量方面表现更优。

原文摘要 · Abstract (English)

We propose TextWand, a general-purpose framework that unifies scene text removal, generation, and replacement into a single model. By decomposing complex editing tasks into the atomic primitives of rendering and erasure, TextWand achieves precise control over both text appearance and background integrity. Specifically, we introduce a novel design, Overlay-Reference Positional Encoding (ORPE), to enforce pixel-level layout fidelity and exemplar-driven style control, alongside a new strategy, Region-Adaptive Suppression (RAS), to ensure clean text erasure. To address the absence of a comprehensive benchmark for general-purpose scene text editing among existing single-task datasets, we construct TextWand-Bench. Extensive experiments demonstrate that TextWand outperforms existing leading open-source and closed-source models by delivering superior text content accuracy, layout and style consistency, and overall image quality across scene text removal, generation and replacement tasks.

文字编辑图像修复统一框架生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。