arXiv:2510.02787cs.CV2025-10

构建复杂背景文本移除数据集,解决真实场景下隐私保护难题

OTR: Synthesizing Overlay Text Dataset for Text Removal

  • 用视觉语言模型生成内容+物体感知布局合成复杂背景文本
  • 数据集含干净真值标签,支持跨域文本移除评估
  • 适合图像编辑、隐私保护领域研究者使用

文本移除是计算机视觉中的关键任务,广泛应用于隐私保护、图像编辑和媒体重用。现有研究多聚焦于自然图像中的场景文本移除,但当前数据集存在局限性,难以支持域外泛化或准确评估。例如,常用基准SCUT-EnsText因人工编辑产生真值伪影、文本背景过于简单,且评估指标无法有效反映生成结果质量。为此,我们提出一种合成文本移除基准的方法,适用于非场景文本的多种应用场景。该数据集通过物体感知布局将文本渲染在复杂背景上,并利用视觉语言模型生成内容,确保真值清晰且移除挑战性强。数据集已发布于https://huggingface.co/datasets/cyberagent/OTR。

原文摘要 · Abstract (English)

Text removal is a crucial task in computer vision with applications such as privacy preservation, image editing, and media reuse. While existing research has primarily focused on scene text removal in natural images, limitations in current datasets hinder out-of-domain generalization or accurate evaluation. In particular, widely used benchmarks such as SCUT-EnsText suffer from ground truth artifacts due to manual editing, overly simplistic text backgrounds, and evaluation metrics that do not capture the quality of generated results. To address these issues, we introduce an approach to synthesizing a text removal benchmark applicable to domains other than scene texts. Our dataset features text rendered on complex backgrounds using object-aware placement and vision-language model-generated content, ensuring clean ground truth and challenging text removal scenarios. The dataset is available at https://huggingface.co/datasets/cyberagent/OTR .

文本移除数据集隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。