arXiv:2412.18063cs.ROcs.DL2024-12被引 13

用大模型提升OCR自动化效率,速度最快快52%

LMRPA: Large Language Model-Driven Efficient Robotic Process Automation for OCR

  • 引入大语言模型增强文本识别准确率与可读性
  • 在Tesseract和DocTR上均实现超50%提速,最快仅9.8秒
  • 适合需要高精度文本提取的自动化场景

本文提出LMRPA,一种基于大语言模型的新型机器人流程自动化(RPA)模型,旨在显著提升光学字符识别(OCR)任务的效率与速度。传统RPA平台在处理高负载重复性任务如OCR时常遇性能瓶颈,导致效率低下。LMRPA通过集成大语言模型(LLMs),有效提升提取文本的准确性与可读性,解决模糊字符与复杂文本结构带来的挑战。在多个基准测试中,对比UiPath与Automation Anywhere等主流RPA平台,使用Tesseract与DocTR OCR引擎,结果表明LMRPA表现更优:在Tesseract的Batch 2任务中,仅需9.8秒完成,而UiPath耗时18.1秒,Automation Anywhere为18.7秒;使用DocTR时,LMRPA仅需12.7秒,竞争对手均超过20秒。这些结果凸显了LMRPA在驱动自动化流程中的变革潜力,为现有最先进的RPA模型提供更高效、可靠的替代方案。

原文摘要 · Abstract (English)

This paper introduces LMRPA, a novel Large Model-Driven Robotic Process Automation (RPA) model designed to greatly improve the efficiency and speed of Optical Character Recognition (OCR) tasks. Traditional RPA platforms often suffer from performance bottlenecks when handling high-volume repetitive processes like OCR, leading to a less efficient and more time-consuming process. LMRPA allows the integration of Large Language Models (LLMs) to improve the accuracy and readability of extracted text, overcoming the challenges posed by ambiguous characters and complex text structures.Extensive benchmarks were conducted comparing LMRPA to leading RPA platforms, including UiPath and Automation Anywhere, using OCR engines like Tesseract and DocTR. The results are that LMRPA achieves superior performance, cutting the processing times by up to 52\%. For instance, in Batch 2 of the Tesseract OCR task, LMRPA completed the process in 9.8 seconds, where UiPath finished in 18.1 seconds and Automation Anywhere finished in 18.7 seconds. Similar improvements were observed with DocTR, where LMRPA outperformed other automation tools conducting the same process by completing tasks in 12.7 seconds, while competitors took over 20 seconds to do the same. These findings highlight the potential of LMRPA to revolutionize OCR-driven automation processes, offering a more efficient and effective alternative solution to the existing state-of-the-art RPA models.

OCR自动化大模型RPA效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。