arXiv:2411.00355cs.CVcs.AI2024-11

无需训练和标注,用扩散模型彻底擦除图像中的异常文字。

TextDestroyer: A Training- and Annotation-Free Diffusion Method for Destroying Anomal Text from Images

  • 用高斯分布打乱文本区域的潜在编码,实现无损背景恢复。
  • 三阶段层级流程生成精确文本掩码,消除可见残留痕迹。
  • 适用于真实场景与生成图像,通用性强且无需额外数据。

本文提出TextDestroyer,首个基于预训练扩散模型的无训练、无标注场景文字擦除方法。现有文字移除模型需复杂标注与重训练,常留下可辨识的微弱文本痕迹,影响隐私保护与内容隐藏效果。TextDestroyer通过三阶段层级流程获取精确文本掩码,在重建前使用高斯分布打乱初始潜在编码中的文本区域。扩散去噪过程中,自注意力的键与值参考原始潜在编码以恢复受损背景。每个反演步骤保存的潜在编码用于重建替换,确保背景完美复原。优势包括:(1)免除繁琐的数据标注与耗资源的训练;(2)实现更彻底的文字清除,防止可识别痕迹残留;(3)具备更强泛化能力,在真实场景与生成图像上均表现良好。

原文摘要 · Abstract (English)

In this paper, we propose TextDestroyer, the first training- and annotation-free method for scene text destruction using a pre-trained diffusion model. Existing scene text removal models require complex annotation and retraining, and may leave faint yet recognizable text information, compromising privacy protection and content concealment. TextDestroyer addresses these issues by employing a three-stage hierarchical process to obtain accurate text masks. Our method scrambles text areas in the latent start code using a Gaussian distribution before reconstruction. During the diffusion denoising process, self-attention key and value are referenced from the original latent to restore the compromised background. Latent codes saved at each inversion step are used for replacement during reconstruction, ensuring perfect background restoration. The advantages of TextDestroyer include: (1) it eliminates labor-intensive data annotation and resource-intensive training; (2) it achieves more thorough text destruction, preventing recognizable traces; and (3) it demonstrates better generalization capabilities, performing well on both real-world scenes and generated images.

图像修复扩散模型隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。