用自动生成的文档图像提升表格检测模型性能,减少人工标注
Synthetic Data Augmentation for Table Detection: Re-evaluating TableNet's Performance with Automatically Generated Document Images
- 基于LaTeX构建合成数据管道,生成视觉多样且带真实标注的双栏文档
- 在合成测试集上表检像素误差低至4.04%,真实数据集上达9.18%
- 适合需要降低标注成本的文档智能与表格识别研究者
通过智能手机或扫描仪获取的文档页面常包含表格,但人工提取效率低且易出错。本文提出一种基于LaTeX的自动化流程,合成具有视觉多样性、双栏布局及对齐真实标注掩码的文档图像。该合成数据集扩充了真实世界Marmot基准,并支持对TableNet的系统性分辨率分析。在256×256输入分辨率下,用合成数据训练的TableNet在合成测试集上像素级异或误差为4.04%;在1024×1024分辨率下为4.33%。在Marmot基准上最佳表现达9.18%(256×256),同时显著减少人工标注工作量。
原文摘要 · Abstract (English)
Document pages captured by smartphones or scanners often contain tables, yet manual extraction is slow and error-prone. We introduce an automated LaTeX-based pipeline that synthesizes realistic two-column pages with visually diverse table layouts and aligned ground-truth masks. The generated corpus augments the real-world Marmot benchmark and enables a systematic resolution study of TableNet. Training TableNet on our synthetic data achieves a pixel-wise XOR error of 4.04% on our synthetic test set with a 256x256 input resolution, and 4.33% with 1024x1024. The best performance on the Marmot benchmark is 9.18% (at 256x256), while cutting manual annotation effort through automation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。