通过模拟手写文档风格增强数据,提升模型对关键信息的定位能力。
Enhancing Document Key Information Localization Through Data Augmentation
- 用合成手写风格扩充数字文档数据集
- 在竞赛中实现高精度关键信息定位
- 适合需要泛化到手写文档的场景
Visually Rich Form Document Intelligence and Understanding (VRDIU) Track B 的目标是定位文档图像中的关键信息。要求仅使用数字文档进行训练,但能有效识别数字和手写文档中的对象。本文提出一种简单有效的流程,包含文档增强阶段和目标检测阶段。具体而言,通过模拟手写文档外观来扩充数字文档训练集。实验表明,该方法显著提升了模型的泛化能力,在竞赛中取得优异表现。
原文摘要 · Abstract (English)
The Visually Rich Form Document Intelligence and Understanding (VRDIU) Track B focuses on the localization of key information in document images. The goal is to develop a method capable of localizing objects in both digital and handwritten documents, using only digital documents for training. This paper presents a simple yet effective approach that includes a document augmentation phase and an object detection phase. Specifically, we augment the training set of digital documents by mimicking the appearance of handwritten documents. Our experiments demonstrate that this pipeline enhances the models' generalization ability and achieves high performance in the competition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。