用大模型从印度财政文件中提取结构化数据,实现高精度验证。
Information Extraction From Fiscal Documents Using LLMs
- 分阶段处理,结合领域知识与文档层级结构
- 利用表格层级关系自动验证数值准确性
- 适合需要处理多页财政文档的研究者
大型语言模型(LLMs)在文本理解方面表现卓越,但对复杂层次化表格数据的处理仍待探索。本文提出一种新方法,利用基于LLM的技术从多页政府财政文件中提取结构化数据。以印度卡纳塔克邦(200+页)年度财政文件为对象,通过融合领域知识、序列上下文与算法验证的多阶段流程,实现了高精度提取。传统OCR难以验证数字准确性,而财政表格本身具有层级总和结构,可提供强内部校验。我们利用这种层级关系构建多级验证机制,证明了LLMs不仅能读取表格,还能处理文档特定的层级结构,为将基于PDF的财政披露转化为可研究数据库提供了可扩展方案,对发展中国家场景具广泛潜力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities in text comprehension, but their ability to process complex, hierarchical tabular data remains underexplored. We present a novel approach to extracting structured data from multi-page government fiscal documents using LLM-based techniques. Applied to annual fiscal documents from the State of Karnataka in India (200+ pages), our method achieves high accuracy through a multi-stage pipeline that leverages domain knowledge, sequential context, and algorithmic validation. A large challenge with traditional OCR methods is the inability to verify the accurate extraction of numbers. When applied to fiscal data, the inherent structure of fiscal tables, with totals at each level of the hierarchy, allows for robust internal validation of the extracted data. We use these hierarchical relationships to create multi-level validation checks. We demonstrate that LLMs can read tables and also process document-specific structural hierarchies, offering a scalable process for converting PDF-based fiscal disclosures into research-ready databases. Our implementation shows promise for broader applications across developing country contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。