用分层扩散模型统一修复各类高分辨率文档图像。
TextDoctor: Unified Document Image Inpainting via Patch Pyramid Diffusion Models
- 分层补丁扩散架构,按文本尺度逐级修复。
- 在7个数据集上优于现有方法,显著提升修复质量。
- 适合需要处理多风格、高分辨率文档的场景。
真实世界文本文档的数字化版本常因原始文件腐蚀、扫描质量差或人为干扰而受损。现有文档修复与补丁方法通常难以泛化到未见文档风格,且难处理高分辨率图像。为此,我们提出TextDoctor,一种新型统一文档图像补丁修复方法。受人类阅读行为启发,TextDoctor从图像块中恢复基础文本元素,并使用扩散模型对整幅文档图像进行修复,而非针对特定文档类型训练模型。为应对不同文本尺寸并避免显存溢出,我们提出结构金字塔预测与分层补丁扩散模型,利用多尺度输入和金字塔块,在全局与局部层面提升修复质量。在七个公开数据集上的大量定性与定量实验表明,TextDoctor在多种高分辨率文档图像修复任务中均优于现有最优方法。
原文摘要 · Abstract (English)
Digital versions of real-world text documents often suffer from issues like environmental corrosion of the original document, low-quality scanning, or human interference. Existing document restoration and inpainting methods typically struggle with generalizing to unseen document styles and handling high-resolution images. To address these challenges, we introduce TextDoctor, a novel unified document image inpainting method. Inspired by human reading behavior, TextDoctor restores fundamental text elements from patches and then applies diffusion models to entire document images instead of training models on specific document types. To handle varying text sizes and avoid out-of-memory issues, common in high-resolution documents, we propose using structure pyramid prediction and patch pyramid diffusion models. These techniques leverage multiscale inputs and pyramid patches to enhance the quality of inpainting both globally and locally. Extensive qualitative and quantitative experiments on seven public datasets validated that TextDoctor outperforms state-of-the-art methods in restoring various types of high-resolution document images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。