arXiv:2506.18023cs.CVcs.AI2025-06

PP-DocBee2提升文档理解性能,推理速度更快。

PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding

  • 用大模型评估并筛选高质量合成数据,优化训练集
  • 在中文业务文档上提升11.4%性能,推理延迟降低73%
  • 适合需要高效多模态文档理解的工业应用

本文介绍PP-DocBee2,是PP-DocBee的升级版,旨在提升多模态文档理解能力。基于大型多模态模型架构,通过改进合成数据质量、优化视觉特征融合策略及推理方法,实现对中文业务文档内部基准测试性能提升11.4%,推理延迟相比原始版本降低73.0%。核心创新在于针对多模态文档任务的数据质量优化策略:利用大规模多模态预训练模型评估数据,并采用新型统计准则过滤异常样本,确保训练数据高质量。受对多模态模型中未充分利用中间特征的启发,通过分层分解视觉变换器(ViT)并引入新型特征融合机制,增强其表征能力,提升复杂推理效果。源代码与预训练模型已开源于https://github.com/PaddlePaddle/PaddleMIX。

原文摘要 · Abstract (English)

This report introduces PP-DocBee2, an advanced version of the PP-DocBee, designed to enhance multimodal document understanding. Built on a large multimodal model architecture, PP-DocBee2 addresses the limitations of its predecessor through key technological improvements, including enhanced synthetic data quality, improved visual feature fusion strategy, and optimized inference methodologies. These enhancements yield an $11.4\%$ performance boost on internal benchmarks for Chinese business documents, and reduce inference latency by $73.0\%$ to the vanilla version. A key innovation of our work is a data quality optimization strategy for multimodal document tasks. By employing a large-scale multimodal pre-trained model to evaluate data, we apply a novel statistical criterion to filter outliers, ensuring high-quality training data. Inspired by insights into underutilized intermediate features in multimodal models, we enhance the ViT representational capacity by decomposing it into layers and applying a novel feature fusion strategy to improve complex reasoning. The source code and pre-trained model are available at \href{https://github.com/PaddlePaddle/PaddleMIX}{https://github.com/PaddlePaddle/PaddleMIX}.

文档理解多模态高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。