arXiv:2511.04699cs.CLcs.CV2025-11被引 1

构建250万样本跨语言合成数据集,提升阿拉伯文文档识别能力

Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding

  • 用真实扫描背景+双语排版生成逼真阿拉伯文档图像
  • 在多个阿拉伯文本基准上,WER和CER显著降低
  • 适合做多语言文档理解、OCR与图表识别的研究者使用

Cross-Lingual SynthDocs 是一个大规模合成语料库,旨在解决阿拉伯文光学字符识别(OCR)与文档理解(DU)资源匮乏的问题。该数据集包含超过250万样本,包括150万条文本数据、27万张完整标注的表格及数十万张基于真实数据的图表。其生成流程采用真实扫描背景、双语布局和带符号的字体,以捕捉阿拉伯文档的排版与结构复杂性。此外,数据还涵盖多样化的图表与表格渲染风格。在SynthDocs上微调Qwen-2.5-VL模型,在多个公开阿拉伯文基准上实现了词错误率(WER)和字符错误率(CER)的持续下降;同时,在树编辑距离相似性(TEDS)与图表提取得分(CharTeX)等多模态任务中也取得提升。SynthDocs为多语言文档分析研究提供了可扩展、视觉真实的高质量资源。

原文摘要 · Abstract (English)

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 million of samples, including 1.5 million textual data, 270K fully annotated tables, and hundred thousands of real data based charts. Our pipeline leverages authentic scanned backgrounds, bilingual layouts, and diacritic aware fonts to capture the typographic and structural complexity of Arabic documents. In addition to text, the corpus includes variety of rendered styles for charts and tables. Finetuning Qwen-2.5-VL on SynthDocs yields consistent improvements in Word Error Rate (WER) and Character Error Rate (CER) in terms of OCR across multiple public Arabic benchmarks, Tree-Edit Distance Similarity (TEDS) and Chart Extraction Score (CharTeX) improved as well in other modalities. SynthDocs provides a scalable, visually realistic resource for advancing research in multilingual document analysis.

文档理解OCR多语言合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。