MinerU 用精准规则提取多类型文档内容,开源可用。
MinerU: An Open-Source Solution for Precise Document Content Extraction

- 结合PDF-Extract-Kit模型与精细预处理后处理
- 跨多种文档类型保持高精度提取效果
- 适合需要稳定文档解析的开发者或研究者
文档内容分析是计算机视觉中的关键研究方向。尽管光学字符识别(OCR)、版面检测和公式识别等方法已取得显著进展,现有开源方案仍难以在多样文档类型中持续提供高质量的内容提取。为此,我们提出MinerU,一个用于高精度文档内容提取的开源解决方案。MinerU利用先进的PDF-Extract-Kit模型高效提取各类文档内容,并通过精心调优的预处理与后处理规则保障最终结果的准确性。实验表明,MinerU在多种文档类型上均表现出一致的高性能,显著提升了内容提取的质量与一致性。MinerU开源项目已发布于https://github.com/opendatalab/MinerU。
原文摘要 · Abstract (English)
Document content analysis has been a crucial research area in computer vision. Despite significant advancements in methods such as OCR, layout detection, and formula recognition, existing open-source solutions struggle to consistently deliver high-quality content extraction due to the diversity in document types and content. To address these challenges, we present MinerU, an open-source solution for high-precision document content extraction. MinerU leverages the sophisticated PDF-Extract-Kit models to extract content from diverse documents effectively and employs finely-tuned preprocessing and postprocessing rules to ensure the accuracy of the final results. Experimental results demonstrate that MinerU consistently achieves high performance across various document types, significantly enhancing the quality and consistency of content extraction. The MinerU open-source project is available at https://github.com/opendatalab/MinerU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。