arXiv:2512.15098cs.CV2025-12被引 3

高效解析科技文献与专利,支持多模态信息精准提取。

Uni-Parser Technical Report

  • 采用模块化松耦合架构,保持文本、公式、图表等跨模态对齐。
  • 8块RTX 4090D显卡下每秒处理20页PDF,可支撑百亿页规模处理。
  • 适合科研数据挖掘、化学结构提取及大模型训练等场景使用。

本文介绍Uni-Parser,一款面向科学文献与专利的工业级文档解析引擎,具备高吞吐、强鲁棒性与低成本优势。与传统流水线式解析方法不同,Uni-Parser采用模块化、松耦合的多专家架构,在保持文本、公式、表格、图表及化学结构等细粒度跨模态对齐的同时,支持新兴模态的灵活扩展。系统集成自适应GPU负载均衡、分布式推理、动态模块编排及可配置模式,既支持整体解析也支持模态专项处理。优化于大规模云部署,基于8×NVIDIA RTX 4090D GPU实现最高每秒20页的处理速度,可实现百亿页级文档的成本效益推理。该能力广泛支撑文献检索、摘要生成、化学结构与反应路径提取、生物活性数据抽取,以及用于训练下一代大语言模型和AI4Science模型的海量语料库构建。

原文摘要 · Abstract (English)

This technical report introduces Uni-Parser, an industrial-grade document parsing engine tailored for scientific literature and patents, delivering high throughput, robust accuracy, and cost efficiency. Unlike pipeline-based document parsing methods, Uni-Parser employs a modular, loosely coupled multi-expert architecture that preserves fine-grained cross-modal alignments across text, equations, tables, figures, and chemical structures, while remaining easily extensible to emerging modalities. The system incorporates adaptive GPU load balancing, distributed inference, dynamic module orchestration, and configurable modes that support either holistic or modality-specific parsing. Optimized for large-scale cloud deployment, Uni-Parser achieves a processing rate of up to 20 PDF pages per second on 8 x NVIDIA RTX 4090D GPUs, enabling cost-efficient inference across billions of pages. This level of scalability facilitates a broad spectrum of downstream applications, ranging from literature retrieval and summarization to the extraction of chemical structures, reaction schemes, and bioactivity data, as well as the curation of large-scale corpora for training next-generation large language models and AI4Science models.

文档解析多模态AI4Science大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。