用0.28亿参数模型统一处理文档解析,提升效率与精度
DocFusion: A Unified Framework for Document Parsing Tasks
- 轻量级生成模型统一多种文档解析任务
- 通过协同训练使识别与检测性能显著提升
- 适合需要高效多任务文档处理的场景
文档解析对分析复杂文档结构和提取细粒度信息至关重要,支撑众多下游应用。然而现有方法通常需集成多个独立模型处理不同解析任务,导致复杂度高、维护成本大。为此,我们提出DocFusion,一个仅含0.28B参数的轻量级生成模型。该模型通过改进的目标函数统一任务表示并实现协同训练,实验揭示了识别任务间的相互促进机制,并证明融合识别数据可显著提升检测性能。最终结果表明,DocFusion在四项关键任务上均达到当前最优(SOTA)水平。
原文摘要 · Abstract (English)
Document parsing is essential for analyzing complex document structures and extracting fine-grained information, supporting numerous downstream applications. However, existing methods often require integrating multiple independent models to handle various parsing tasks, leading to high complexity and maintenance overhead. To address this, we propose DocFusion, a lightweight generative model with only 0.28B parameters. It unifies task representations and achieves collaborative training through an improved objective function. Experiments reveal and leverage the mutually beneficial interaction among recognition tasks, and integrating recognition data significantly enhances detection performance. The final results demonstrate that DocFusion achieves state-of-the-art (SOTA) performance across four key tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。