arXiv:2503.04065cs.CVcs.AI2025-03被引 3

PP-DocBee通过数据合成与训练技巧,提升文档图文理解性能。

PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks

  • 构建文档场景专用数据合成策略,增强模型泛化能力
  • 在中英文文档理解任务上均达当前最优效果
  • 适合需要高效文档解析的工业级应用

随着数字化快速发展,各类文档图像在生产与日常生活中广泛应用,对文档图像内容的快速准确解析需求日益迫切。为此,本文提出PP-DocBee,一种面向端到端文档图像理解的新型多模态大语言模型。首先,针对文档场景设计数据合成策略,构建多样化数据集以提升模型泛化性;其次,采用动态比例采样、数据预处理及OCR后处理等训练技术。大量实验表明,PP-DocBee在英文文档理解基准上表现卓越,且在中文文档理解任务中超越现有开源与商业模型。源代码与预训练模型已公开于https://github.com/PaddlePaddle/PaddleMIX。

原文摘要 · Abstract (English)

With the rapid advancement of digitalization, various document images are being applied more extensively in production and daily life, and there is an increasingly urgent need for fast and accurate parsing of the content in document images. Therefore, this report presents PP-DocBee, a novel multimodal large language model designed for end-to-end document image understanding. First, we develop a data synthesis strategy tailored to document scenarios in which we build a diverse dataset to improve the model generalization. Then, we apply a few training techniques, including dynamic proportional sampling, data preprocessing, and OCR postprocessing strategies. Extensive evaluations demonstrate the superior performance of PP-DocBee, achieving state-of-the-art results on English document understanding benchmarks and even outperforming existing open source and commercial models in Chinese document understanding. The source code and pre-trained models are publicly available at \href{https://github.com/PaddlePaddle/PaddleMIX}{https://github.com/PaddlePaddle/PaddleMIX}.

文档理解多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。