用符号规则增强大模型,让文档信息提取更准且无需大量标注。
Neurosymbolic Information Extraction from Transactional Documents
- 结合语言模型与符号验证,生成候选信息并过滤错误结果。
- 在交易文档上提升F1分数和准确率,零样本表现显著改善。
- 适合需要高精度、少标注的金融或法律文档处理场景。
本文提出一种神经符号框架,用于从交易类文档中提取信息。通过基于模式的方法,整合符号验证机制,实现更有效的零样本输出与知识蒸馏。该方法利用语言模型生成候选抽取结果,并通过句法、任务和领域级验证,确保符合特定领域的算术约束。贡献包括一个全面的交易文档模式、重新标注的数据集,以及生成高质量标签以支持知识蒸馏的方法。实验结果表明,在F1分数和准确率方面均有显著提升,证明了神经符号验证在交易文档处理中的有效性。
原文摘要 · Abstract (English)
This paper presents a neurosymbolic framework for information extraction from documents, evaluated on transactional documents. We introduce a schema-based approach that integrates symbolic validation methods to enable more effective zero-shot output and knowledge distillation. The methodology uses language models to generate candidate extractions, which are then filtered through syntactic-, task-, and domain-level validation to ensure adherence to domain-specific arithmetic constraints. Our contributions include a comprehensive schema for transactional documents, relabeled datasets, and an approach for generating high-quality labels for knowledge distillation. Experimental results demonstrate significant improvements in $F_1$-scores and accuracy, highlighting the effectiveness of neurosymbolic validation in transactional document processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。