轻量化模型可高效提取扫描文档中的结构化信息。
Arctic-Extract Technical Report
- 基于轻量设计,仅6.6 GiB大小适配低资源设备。
- 单块A10 GPU可处理最多125页A4文档。
- 适合需要长文档处理的工业级应用。
Arctic-Extract 是一种用于从扫描或数字生成的企业文档中提取结构化数据(如问答、实体、表格)的前沿模型。尽管具备顶尖性能,其模型重量仅为6.6 GiB,可在资源受限硬件上部署,例如配备24 GB内存的A10 GPU。该模型可在这些GPU上处理最多125页的A4文档,适用于长文档处理任务。本文介绍了Arctic-Extract的训练策略与评估结果,验证了其在文档理解任务中的强大表现。
原文摘要 · Abstract (English)
Arctic-Extract is a state-of-the-art model designed for extracting structural data (question answering, entities and tables) from scanned or digital-born business documents. Despite its SoTA capabilities, the model is deployable on resource-constrained hardware, weighting only 6.6 GiB, making it suitable for deployment on devices with limited resources, such as A10 GPUs with 24 GB of memory. Arctic-Extract can process up to 125 A4 pages on those GPUs, making suitable for long document processing. This paper highlights Arctic-Extract's training protocols and evaluation results, demonstrating its strong performance in document understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。