arXiv:2511.22499cs.CVcs.CL2025-11

优化文本移除的掩码形状,提升复杂文档图像修复效果。

What Shape Is Optimal for Masks in Text Removal?

  • 用贝叶斯优化学习灵活掩码参数,实现字符级精准遮挡。
  • 实验表明,最小覆盖掩码并非最优,形状影响修复质量。
  • 为实际文档去文提供可操作的掩码设计指南,适合工业应用。

生成模型显著提升了图像修复的准确性。尤其在从文档图像中移除特定文字时,重建原始图像对工业应用至关重要。然而,现有方法多针对户外场景中的简单文字,缺乏对密集复杂文本图像的研究。为此,我们构建了包含大量文本的大规模文本移除基准数据集。分析发现,文本移除性能对掩码轮廓扰动敏感,因此实际任务中需精确调整掩码形状。本研究提出一种建模高度灵活掩码轮廓的方法,并通过贝叶斯优化学习其参数。结果表明,最优掩码为字符级掩码,且最小覆盖区域并非最佳。该研究有望为人工掩码设计提供直观实用的指导。

原文摘要 · Abstract (English)

The advent of generative models has dramatically improved the accuracy of image inpainting. In particular, by removing specific text from document images, reconstructing original images is extremely important for industrial applications. However, most existing methods of text removal focus on deleting simple scene text which appears in images captured by a camera in an outdoor environment. There is little research dedicated to complex and practical images with dense text. Therefore, we created benchmark data for text removal from images including a large amount of text. From the data, we found that text-removal performance becomes vulnerable against mask profile perturbation. Thus, for practical text-removal tasks, precise tuning of the mask shape is essential. This study developed a method to model highly flexible mask profiles and learn their parameters using Bayesian optimization. The resulting profiles were found to be character-wise masks. It was also found that the minimum cover of a text region is not optimal. Our research is expected to pave the way for a user-friendly guideline for manual masking.

图像修复文本移除掩码优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。