用AI修复破损文档文本,保持字体格式一致且语义连贯。
DocRevive: A Unified Pipeline for Document Text Restoration

- 融合OCR、图像分析与扩散模型,统一重建受损文本。
- 在3万张合成退化文档上达到高保真恢复效果。
- 适合档案修复与数字保存领域研究者使用。
在文档理解中,恢复破损、遮挡或不完整文本仍是关键但未充分探索的问题。本文提出一种统一的文本恢复流水线,结合先进的光学字符识别(OCR)、图像分析、掩码语言建模及基于扩散的模型,在保留视觉完整性的前提下重建文本。我们构建了一个包含30,078张退化文档图像的合成数据集,涵盖多样退化场景,为该任务设立基准。该流程可检测并识别文本,通过遮挡检测器定位退化区域,并利用内补模型实现语义连贯的重建。一个基于扩散的模块能无缝重置文本,匹配原始字体、字号和对齐方式。为评估恢复质量,我们提出统一上下文相似性度量(UCSM),综合编辑、语义和长度相似性,并引入上下文可预测性惩罚项,当正确文本在上下文中明显时,对偏差进行强化惩罚。本工作推动了文档恢复技术发展,有助于档案研究与数字保存。数据集与代码已开源:Hugging Face与Github。
原文摘要 · Abstract (English)
In Document Understanding, the challenge of reconstructing damaged, occluded, or incomplete text remains a critical yet unexplored problem. Subsequent document understanding tasks can benefit from a document reconstruction process. In response, this paper presents a novel unified pipeline combining state-of-the-art Optical Character Recognition (OCR), advanced image analysis, masked language modeling, and diffusion-based models to restore and reconstruct text while preserving visual integrity. We create a synthetic dataset of 30{,}078 degraded document images that simulates diverse document degradation scenarios, setting a benchmark for restoration tasks. Our pipeline detects and recognizes text, identifies degradation with an occlusion detector, and uses an inpainting model for semantically coherent reconstruction. A diffusion-based module seamlessly reintegrates text, matching font, size, and alignment. To evaluate restoration quality, we propose a Unified Context Similarity Metric (UCSM), incorporating edit, semantic, and length similarities with a contextual predictability measure that penalizes deviations when the correct text is contextually obvious. Our work advances document restoration, benefiting archival research and digital preservation while setting a new standard for text reconstruction. The OPRB dataset and code are available at \href{https://huggingface.co/datasets/kpurkayastha/OPRB}{Hugging Face} and \href{https://github.com/kunalpurkayastha/DocRevive}{Github} respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。