arXiv:2602.06450cs.CV2026-02中稿 · CVPR被引 2

提出强合成引擎UnionST,让合成文本数据更逼真、效果超真实数据。

What Is Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution

  • 构建覆盖复杂场景的合成数据集UnionST-S,提升字体、布局多样性
  • 模型在合成数据上训练后,在部分场景超越真实数据表现
  • 自进化标注框架仅需9%真实标签即可达到竞争性性能

大规模且类别平衡的文本数据对训练有效的场景文本识别(STR)模型至关重要,但真实数据收集困难。合成数据虽成本低且标签完美,但性能常落后,暴露真实与合成数据间的显著域差距。本文系统分析主流基于渲染的合成数据集,发现其在语料、字体和版式多样性上的不足,限制了复杂场景下的真实性。为此,我们提出UnionST合成引擎,生成涵盖挑战样本的文本数据,更好匹配真实世界复杂性。进而构建大规模合成数据集UnionST-S,改进复杂场景模拟。此外,开发自进化学习(SEL)框架,实现高效真实数据标注。实验表明,基于UnionST-S训练的模型显著优于现有合成数据集,在某些场景甚至超越真实数据表现;结合SEL,仅需9%真实标签即达竞争力水平。代码已开源。

原文摘要 · Abstract (English)

Large-scale and categorical-balanced text data is essential for training effective Scene Text Recognition (STR) models, which is hard to achieve when collecting real data. Synthetic data offers a cost-effective and perfectly labeled alternative. However, its performance often lags behind, revealing a significant domain gap between real and current synthetic data. In this work, we systematically analyze mainstream rendering-based synthetic datasets and identify their key limitations: insufficient diversity in corpus, font, and layout, which restricts their realism in complex scenarios. To address these issues, we introduce UnionST, a strong data engine synthesizes text covering a union of challenging samples and better aligns with the complexity observed in the wild. We then construct UnionST-S, a large-scale synthetic dataset with improved simulations in challenging scenarios. Furthermore, we develop a self-evolution learning (SEL) framework for effective real data annotation. Experiments show that models trained on UnionST-S achieve significant improvements over existing synthetic datasets. They even surpass real-data performance in certain scenarios. Moreover, when using SEL, the trained models achieve competitive performance by only seeing 9% of real data labels. Code is available at https://github.com/YesianRohn/UnionST.

文本识别合成数据自进化数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。