arXiv:2605.17309cs.CVcs.AI2026-05中稿 · CVPR

构建大规模风格化文本修复数据集,支持精准评估文本可读性与视觉一致性。

StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting

论文配图:StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting
图 1 · 摘自论文原文
  • 自动化流水线生成带风格保持的图文掩码三元组。
  • 引入标准化OCR指标与CLIP相似度评估,提升评测可复现性。
  • 适合文本修复、图像生成与跨模态对齐研究者使用。

我们提出StyleText,一个用于局部场景文本修复并保留风格的大规模数据集与基准。该数据集包含28,518个图像-掩码-提示三元组,归类为9,932个场景家族,可在共享场景上下文中实现对文本可读性和视觉一致性的受控评估。数据集通过自动化流程构建:结合大语言模型提示模板、基于Flux的源图像生成(含键值缓存注入)、基于OCR的语义过滤、多边形掩码提取以及掩码条件下的FluxFill增强。我们定义了可复现的评估协议,采用归一化的OCR指标(单词准确率和字符错误率)及明确预处理的CLIP图像-图像相似度。在StyleText上训练的FluxFill+LoRA基线模型显著提升OCR准确率,同时保持场景风格一致性,为未来研究提供了强参考基准。

原文摘要 · Abstract (English)

We present StyleText, a large-scale dataset and benchmark for localized scene-text inpainting with style preservation. StyleText contains 28,518 image-mask-prompt triplets grouped into 9,932 scene families, enabling controlled evaluation of text legibility and visual consistency under shared scene context. We construct the dataset with an automated pipeline that combines LLM prompt templating, Flux-based source generation with key-value (KV) cache injection, OCR-based semantic filtering, polygon mask extraction, and mask-conditioned FluxFill augmentation. We define a reproducible evaluation protocol using normalized OCR metrics (word accuracy and character error rate) and CLIP image-image similarity with explicit preprocessing. A FluxFill+LoRA baseline trained on StyleText improves OCR accuracy substantially over initialization while maintaining scene style consistency, establishing a strong reference point for future comparisons.

文本修复图像生成数据集风格保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。