arXiv:2605.15523cs.CV2026-05

无需额外编码器,直接从图像生成提示,实现跨语言文本风格一致编辑

Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning

论文配图:Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning
图 1 · 摘自论文原文
  • 通过自动生成风格与字形提示,摆脱预训练编码器限制
  • 在多语言数据集上达到当前最佳的文本准确率与风格一致性
  • 适合需要开放词汇编辑和真实场景还原的研究者与开发者

场景文本编辑旨在修改图像目标区域中的文本,同时保持周围背景的风格与纹理。现有方法仅依赖图像背景信息,忽略目标区域的视觉细节,导致丢失原文本的样式特征,本质上退化为文本渲染任务。此外,预训练字形编码器带来的条件限制也缩小了可编辑文本范围。为此,本文提出一种自提示扩散变换器方法,直接从原始图像构建风格与字形提示,无需引入额外编码器。采用两阶段训练策略:先在大规模自监督数据上训练扩散变换器,再用少量成对图像微调。借助多模态扩散变换器(MM-DiT)的上下文学习能力,实现开放词汇与风格一致的文本编辑。在多种语言上的实验结果表明,该方法在文本准确率与风格一致性方面均达到当前最优水平。

原文摘要 · Abstract (English)

Scene text editing aims to modify text in a target region of an image while preserving surrounding background style and texture. Existing methods rely solely on image background information while neglecting the visual details of target regions, which discards stylistic features in the original text and essentially degrades the task to text rendering. Moreover, the conditions imposed by pre-trained glyph encoder limit the scope of editable text. To address these issues, this paper proposes a self-prompting scene text editing method that constructs style and glyph prompts directly from the original image, without introducing additional style or glyph encoders. We employ a two-stage training strategy: the diffusion transformer is first trained on large-scale self-supervised data and then refined using a small set of paired images. By leveraging the in-context learning capability of the Multi-Modal Diffusion Transformer (MM-DiT), it achieves open-vocabulary and style-consistent text editing. Experimental results on various languages demonstrate that our method achieves the state-of-the-art performance in both text accuracy and style consistency. Our project page: hongxiii.github.io/mstedit.

文本编辑扩散模型开放词汇风格一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。