arXiv:2507.02009cs.IR2025-07被引 2

用不确定性量化提升科学表格数据提取准确率,减少人工校验量。

Uncertainty-Aware Complex Scientific Table Data Extraction

  • 基于置信区间预测,融合表格结构与文字识别的不确定性评分。
  • 仅需人工校验47%结果,即可使数据质量提升30%。
  • 适合需要高精度数据的科研人员与自动化数据处理系统。

表格结构识别(TSR)和光学字符识别(OCR)在从科学文档中提取结构化数据方面至关重要。然而,现有基于这些方法的提取框架往往无法量化结果的不确定性。为获得高精度的科学数据,所有提取结果通常需人工验证,耗时且费力。本文提出一种面向复杂科学表格的不确定性感知数据提取框架,采用模型无关的不确定性量化(UQ)方法——置信区间预测。我们探索了多种不确定性评分方法,以聚合由TSR和OCR引入的不确定性。通过标准基准和包含六个科学领域复杂表格的自建数据集进行严格评估,结果表明使用UQ可有效检测提取错误。仅手动验证47%的结果,数据质量即可提升30%。本工作定量证明了不确定性量化在提升人机协作效率、获取科学可用数据方面的潜力。所有代码与数据已公开于GitHub。

原文摘要 · Abstract (English)

Table structure recognition (TSR) and optical character recognition (OCR) play crucial roles in extracting structured data from tables in scientific documents. However, existing extraction frameworks built on top of TSR and OCR methods often fail to quantify the uncertainties of extracted results. To obtain highly accurate data for scientific domains, all extracted data must be manually verified, which can be time-consuming and labor-intensive. We propose a framework that performs uncertainty-aware data extraction for complex scientific tables, built on conformal prediction, a model-agnostic method for uncertainty quantification (UQ). We explored various uncertainty scoring methods to aggregate the uncertainties introduced by TSR and OCR. We rigorously evaluated the framework using a standard benchmark and an in-house dataset consisting of complex scientific tables in six scientific domains. The results demonstrate the effectiveness of using UQ for extraction error detection, and by manually verifying only 47% of extraction results, the data quality can be improved by 30%. Our work quantitatively demonstrates the role of UQ with the potential of improving the efficiency in the human-machine cooperation process to obtain scientifically usable data from complex tables in scientific documents. All code and data are available on GitHub at https://github.com/lamps-lab/TSR-OCR-UQ/tree/main.

表格提取不确定性量化科学数据人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。