用合成数据训练模型,系统修正文档识别中的错字和结构错误。
Revise: A Framework for Revising OCRed text in Practical Information Systems with Data Contamination Strategy
- 构建错误分类体系,用合成数据模拟真实识别错误。
- 在真实数据上纠正率提升,使文档更易管理与检索。
- 适合需要高精度文档处理的工业场景使用。
大型语言模型(LLMs)在文档智能领域取得显著进展,尤其在问答任务中表现优异。然而,现有方法多聚焦于特定任务,缺乏对文档信息的结构化组织与管理能力。为此,我们提出 Revise 框架,从字符、单词到结构层面系统修正 OCR 引入的错误。Revise 采用全面的 OCR 错误分层分类体系,并设计一种合成数据生成策略,真实模拟各类错误以训练高效纠错模型。实验表明,Revise 能有效纠正 OCR 输出,实现文档内容的更结构化表达与系统化管理,显著提升文档检索与问答任务的下游性能,展现出突破现有文档智能框架结构管理局限的潜力。
原文摘要 · Abstract (English)
Recent advances in Large Language Models (LLMs) have significantly improved the field of Document AI, demonstrating remarkable performance on document understanding tasks such as question answering. However, existing approaches primarily focus on solving specific tasks, lacking the capability to structurally organize and manage document information. To address this limitation, we propose Revise, a framework that systematically corrects errors introduced by OCR at the character, word, and structural levels. Specifically, Revise employs a comprehensive hierarchical taxonomy of common OCR errors and a synthetic data generation strategy that realistically simulates such errors to train an effective correction model. Experimental results demonstrate that Revise effectively corrects OCR outputs, enabling more structured representation and systematic management of document contents. Consequently, our method significantly enhances downstream performance in document retrieval and question answering tasks, highlighting the potential to overcome the structural management limitations of existing Document AI frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。