arXiv:2505.13504cs.IRcs.AI2025-05被引 3

用智能体+强化学习实现自动纠错的表单文档信息提取

An agentic system with reinforcement-learned subsystem improvements for parsing form-like documents

  • 设计多智能体系统,通过强化学习优化提示词生成
  • 在SOIRE和CORD数据集上实现高精度、自适应的信息抽取
  • 适合需要持续改进的自动化文档处理场景

从发票、采购单等表单类文档中提取文本数据通常依赖视觉(OCR)与学习算法,或结构单一的流水线,难以实现系统性优化。本文提出一种基于大语言模型(LLM)的智能体系统,结合强化学习驱动的元智能体,实现面对LLM推理不确定性的自动化、可自我改进的数据提取。该系统突破传统单体式方法局限,采用模块化多智能体架构,通过任务特定提示词与奖励惩罚策略,指导元提示智能体从历史错误中学习,持续优化各任务型执行智能体的提示词。系统可处理多样文档、文件格式、版式及不同大模型,旨在实现无需人工干预的高精度信息提取。在SOIRE与CORD两个基准数据集上的实验结果表明,该智能体框架具有显著优势。

原文摘要 · Abstract (English)

Extracting alphanumeric data from form-like documents such as invoices, purchase orders, bills, and financial documents is often performed via vision (OCR) and learning algorithms or monolithic pipelines with limited potential for systemic improvements. We propose an agentic AI system that leverages Large Language Model (LLM) agents and a reinforcement learning (RL) driver agent to automate consistent, self-improving extraction under LLM inference uncertainty. Our work highlights the limitations of monolithic LLM-based extraction and introduces a modular, multi-agent framework with task-specific prompts and an RL policy of rewards and penalties to guide a meta-prompting agent to learn from past errors and improve prompt-based actor agents. This self-corrective adaptive system handles diverse documents, file formats, layouts, and LLMs, aiming to automate accurate information extraction without the need for human intervention. Results as reported on two benchmark datasets of SOIRE, and CORD, are promising for the agentic AI framework.

智能体系统文档提取强化学习LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。