arXiv:2511.05547cs.CVcs.AI2025-11被引 1

用AI融合OCR与大模型,自动精准提取发票信息

Automated Invoice Data Extraction: Using LLM and OCR

  • 结合OCR、深度学习与大模型,实现端到端发票信息提取
  • 在多种复杂布局和手写文本下,准确率显著高于传统方法
  • 适合需要自动化票据处理的企业与系统集成开发者

传统OCR系统因模板依赖性强,难以应对多变的发票版式、手写文字和低质量扫描图像。新兴方案采用卷积神经网络(CNN)与Transformer等深度学习模型,以及领域专用模型,提升不同文档类型中的版面分析与识别精度。大型语言模型(LLMs)通过强大的实体识别与语义理解能力,无需编程即可实现复杂上下文关系建模,显著增强视觉命名实体识别(Visual NER)性能。当前行业最佳实践采用混合架构,融合OCR与LLM以实现高可扩展性与低人工干预。本文提出一个整合OCR、深度学习、大模型与图分析的综合AI平台,在多种发票类型中实现了前所未有的提取质量与一致性。

原文摘要 · Abstract (English)

Conventional Optical Character Recognition (OCR) systems are challenged by variant invoice layouts, handwritten text, and low-quality scans, which are often caused by strong template dependencies that restrict their flexibility across different document structures and layouts. Newer solutions utilize advanced deep learning models such as Convolutional Neural Networks (CNN) as well as Transformers, and domain-specific models for better layout analysis and accuracy across various sections over varied document types. Large Language Models (LLMs) have revolutionized extraction pipelines at their core with sophisticated entity recognition and semantic comprehension to support complex contextual relationship mapping without direct programming specification. Visual Named Entity Recognition (NER) capabilities permit extraction from invoice images with greater contextual sensitivity and much higher accuracy rates than older approaches. Existing industry best practices utilize hybrid architectures that blend OCR technology and LLM for maximum scalability and minimal human intervention. This work introduces a holistic Artificial Intelligence (AI) platform combining OCR, deep learning, LLMs, and graph analytics to achieve unprecedented extraction quality and consistency.

发票提取大模型OCR智能文档处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。