arXiv:2609.03595cs.CLcs.AI2026-09

用合成数据训练出的泰国文字识别模型,性能接近真实标注数据。

How Far Can Synthetic Data Take Thai OCR?

论文配图:How Far Can Synthetic Data Take Thai OCR?
图 1 · 摘自论文原文
  • 分离字体、结构、手写笔迹等变量,控制实验验证各因素影响
  • 合成数据在页级训练下错误率仅1.82%,接近真实标注的1.31%
  • 无需真实标签即可适配,适合无标注数据的低资源语言场景

我们研究合成OCR标注为何能迁移到真实泰语文档,并据此构建了Wayu-Paxa-OCR-Zero——一个无需真实OCR标签即可从真实泰语文档中适配的模型。合成数据可大规模提供精确标签,但“真实感”混杂了源域、页面上下文、排版、空间结构和字形变化等因素。我们通过受控文档重建流程解耦这些因素,在印刷体与手写体泰语文档上进行页级与裁剪级训练。非文本上下文影响有限,而字体多样性、二维结构和真实手写字形显著提升迁移效果;此外,源域匹配依赖训练粒度:在页级训练下,同域重建接近真实印刷标注(中位字符错误率1.82%对比1.31%),但在裁剪级训练下,跨域重建表现更优(15.59%对比5.52%)。基于此,我们利用45,723张合成页将0.9B参数的PaddleOCR-VL-1.6适配为Wayu-Paxa-OCR-Zero:印刷页中位错误率从6.64%降至1.24%,手写页从74.87%降至20.55%,且在五个评估集上均优于Typhoon OCR v1 7B,证明纯合成数据训练具备竞争力。

原文摘要 · Abstract (English)

We investigate what makes synthetic OCR supervision transfer to real Thai documents and use the resulting insights to build Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages. Synthetic data provide exact labels at scale, but "realism" conflates source domain, page context, typography, spatial structure, and glyph variation. We disentangle these factors with a controlled document-reconstruction pipeline and evaluate each variant under page- and crop-level training on printed and handwritten Thai documents. Non-text context has little consistent effect, whereas typeface diversity, two-dimensional structure, and real handwriting glyphs improve transfer; moreover, source-domain matching depends on training granularity, with in-domain reconstruction approaching real printed supervision under page-level training (1.82% versus 1.31% median character error rate) but underperforming out-of-domain reconstruction under crop-level training (15.59% versus 5.52%). Guided by these findings, we adapt the 0.9B-parameter PaddleOCR-VL-1.6 into Wayu-Paxa-OCR-Zero using 45,723 synthetic pages: relative to its base checkpoint, it reduces median character error rate from 6.64% to 1.24% on printed pages and from 74.87% to 20.55% on handwriting and outperforms Typhoon OCR v1 7B on all five evaluation sets, showing that synthetic-only training can be competitive.

OCR合成数据泰语无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。