首个文档包分割评测数据集,助力大模型处理复杂文档拆分任务
DocSplit: A Comprehensive Benchmark Dataset and Evaluation Approach for Document Packet Recognition and Splitting
- 构建多类型、多模态文档包数据集,涵盖不同布局与复杂场景
- 提出新评估指标,量化模型识别边界、分类类型与排序能力
- 适合法律、金融、医疗等需处理混合文档的领域研究者使用
现实应用中,文档理解常需处理由多页异构文档拼接而成的文档包。尽管视觉文档理解取得进展,文档包分割——即将文档包拆分为独立单元——仍缺乏系统研究。本文提出首个综合性基准数据集DocSplit及新型评估指标,用于评估大语言模型在文档包分割上的表现。DocSplit包含五个不同复杂度的数据集,覆盖多种文档类型、版式和多模态场景。我们正式定义了文档包分割任务,要求模型识别文档边界、分类文档类型并保持正确页序。该基准解决真实世界挑战,如页面错序、文档交错及无明显分隔。我们在多个数据集上对多模态大模型进行广泛实验,揭示当前模型在处理复杂文档拆分任务时存在显著性能差距。所发布的数据集与评估指标为法律、金融、医疗等文档密集型领域提升文档理解能力提供了系统框架。
原文摘要 · Abstract (English)
Document understanding in real-world applications often requires processing heterogeneous, multi-page document packets containing multiple documents stitched together. Despite recent advances in visual document understanding, the fundamental task of document packet splitting, which involves separating a document packet into individual units, remains largely unaddressed. We present the first comprehensive benchmark dataset, DocSplit, along with novel evaluation metrics for assessing the document packet splitting capabilities of large language models. DocSplit comprises five datasets of varying complexity, covering diverse document types, layouts, and multimodal settings. We formalize the DocSplit task, which requires models to identify document boundaries, classify document types, and maintain correct page ordering within a document packet. The benchmark addresses real-world challenges, including out-of-order pages, interleaved documents, and documents lacking clear demarcations. We conduct extensive experiments evaluating multimodal LLMs on our datasets, revealing significant performance gaps in current models' ability to handle complex document splitting tasks. The DocSplit benchmark datasets and proposed novel evaluation metrics provide a systematic framework for advancing document understanding capabilities essential for legal, financial, healthcare, and other document-intensive domains. We release the datasets to facilitate future research in document packet processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。