arXiv:2412.19840cs.CVcs.HC2024-12被引 18

ERPA融合OCR与大模型,9.94秒完成身份证信息提取,提速超94%。

ERPA: Efficient RPA Model Integrating OCR and LLMs for Intelligent Document Processing

  • 用大模型提升OCR识别准确率,解决模糊字符与复杂结构问题
  • 在移民文档处理中,9.94秒完成数据提取,比主流工具快94%
  • 适合需要高速高精度文档自动化的企业与政务场景

本文提出ERPA,一种创新的机器人流程自动化(RPA)模型,用于增强移民工作流中的身份证信息提取与优化光学字符识别(OCR)任务。传统RPA在处理大量文档时常因性能瓶颈导致效率低下。ERPA通过集成大语言模型(LLMs),提升提取文本的准确性和清晰度,有效应对模糊字符与复杂结构。与UiPath、Automation Anywhere等领先平台的基准对比显示,ERPA将处理时间缩短高达94%,仅需9.94秒即可完成身份证数据提取。这些结果表明ERPA有潜力彻底改变文档自动化,提供比现有方案更快更可靠的替代选择。

原文摘要 · Abstract (English)

This paper presents ERPA, an innovative Robotic Process Automation (RPA) model designed to enhance ID data extraction and optimize Optical Character Recognition (OCR) tasks within immigration workflows. Traditional RPA solutions often face performance limitations when processing large volumes of documents, leading to inefficiencies. ERPA addresses these challenges by incorporating Large Language Models (LLMs) to improve the accuracy and clarity of extracted text, effectively handling ambiguous characters and complex structures. Benchmark comparisons with leading platforms like UiPath and Automation Anywhere demonstrate that ERPA significantly reduces processing times by up to 94 percent, completing ID data extraction in just 9.94 seconds. These findings highlight ERPA's potential to revolutionize document automation, offering a faster and more reliable alternative to current RPA solutions.

RPAOCR大模型文档处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。