KIEval 从工业应用出发,评估文档关键信息提取的结构化能力。
KIEval: Evaluation Metric for Document Key Information Extraction
- 基于实际工业场景,同时评估实体与信息分组的提取效果
- 相比传统指标,更准确反映真实应用场景中的性能表现
- 适合开发或落地文档信息提取系统的研究人员和工程师
文档关键信息提取(KIE)技术可将文档图像中的有价值信息转化为结构化数据,在工业应用中已成为关键技术。然而,现有评估指标无法准确反映其在工业场景中的核心属性。本文提出 KIEval,一种面向应用的新型文档 KIE 评估指标。与以往方法不同,KIEval 不仅评估单个信息项(实体)的提取,还评估信息之间的结构化关系(分组)。对结构化信息的评估更能体现模型在工业场景下从文档中提取关联信息的能力。该指标专为工业应用设计,我们认为它有望成为实际开发或应用文档 KIE 模型的标准评估工具。代码将公开发布。
原文摘要 · Abstract (English)
Document Key Information Extraction (KIE) is a technology that transforms valuable information in document images into structured data, and it has become an essential function in industrial settings. However, current evaluation metrics of this technology do not accurately reflect the critical attributes of its industrial applications. In this paper, we present KIEval, a novel application-centric evaluation metric for Document KIE models. Unlike prior metrics, KIEval assesses Document KIE models not just on the extraction of individual information (entity) but also of the structured information (grouping). Evaluation of structured information provides assessment of Document KIE models that are more reflective of extracting grouped information from documents in industrial settings. Designed with industrial application in mind, we believe that KIEval can become a standard evaluation metric for developing or applying Document KIE models in practice. The code will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。