提出动态渐进式文档去畸变方法,提升真实场景下弯曲文档的还原效果。
TADoc: Robust Time-Aware Document Image Dewarping
- 将去畸变建模为一系列中间状态的动态过程,更贴近真实拍摄物理运动。
- 在多个基准上优于现有方法,对高变形文档仍保持稳定性能。
- 设计新评估指标DLS,更准确衡量文档布局还原质量,适合实际应用。
便携设备拍摄的弯曲、褶皱和旋转文档图像去畸变,在数字经济与在线办公兴起背景下日益重要。尽管已有多种方法,但在复杂文档结构和高变形场景下仍表现不佳。本文核心洞察是:与去模糊等任务不同,真实场景中的去畸变是一个渐进过程而非一步变换。为此,我们首次将该任务建模为包含一系列中间状态的动态过程,并设计轻量级框架TADoc(Time-Aware Document Dewarping Network)以应对几何失真。此外,由于传统OCR指标对稀疏文本文档评估不充分,我们提出新指标DLS(Document Layout Similarity),用于评估下游任务中去畸变的有效性。大量实验与深入评估表明,本模型具备强鲁棒性,在多种文档类型与不同程度变形的多个基准上均取得领先性能。
原文摘要 · Abstract (English)
Flattening curved, wrinkled, and rotated document images captured by portable photographing devices, termed document image dewarping, has become an increasingly important task with the rise of digital economy and online working. Although many methods have been proposed recently, they often struggle to achieve satisfactory results when confronted with intricate document structures and higher degrees of deformation in real-world scenarios. Our main insight is that, unlike other document restoration tasks (e.g., deblurring), dewarping in real physical scenes is a progressive motion rather than a one-step transformation. Based on this, we have undertaken two key initiatives. Firstly, we reformulate this task, modeling it for the first time as a dynamic process that encompasses a series of intermediate states. Secondly, we design a lightweight framework called TADoc (Time-Aware Document Dewarping Network) to address the geometric distortion of document images. In addition, due to the inadequacy of OCR metrics for document images containing sparse text, the comprehensiveness of evaluation is insufficient. To address this shortcoming, we propose a new metric -- DLS (Document Layout Similarity) -- to evaluate the effectiveness of document dewarping in downstream tasks. Extensive experiments and in-depth evaluations have been conducted and the results indicate that our model possesses strong robustness, achieving superiority on several benchmarks with different document types and degrees of distortion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。