同时考虑文档水平与垂直方向形变,提升图像矫正精度。
D2Dewarp: Dual Dimensions Geometric Representation Learning Based Document Image Dewarping
- 提出双维度形变感知模型,兼顾水平与垂直线细节
- 在中英文基准上均优于现有方法,显著改善矫正效果
- 自动生成大规模标注数据集,解决真实标注缺失问题
文档图像去扭曲在深度学习时代仍是挑战性任务。现有方法虽借助文本行感知有所提升,但通常仅关注单一水平维度。本文提出细粒度形变感知模型D2Dewarp,聚焦文档水平-垂直线的双维度形变,可捕捉不同方向上的扭曲趋势。为融合横纵方向特征,设计基于X/Y坐标的高效融合模块,促进两维特征交互与互补。由于现有公开去扭曲数据集缺乏标注线特征,本文提出一种利用公开文档纹理图像与自动渲染引擎的自动化细粒度标注方法,构建新大型扭曲训练数据集DocDewarpHV。在三个中英文公开基准上,定量与定性结果均表明,本方法优于当前最先进方法。代码与数据集已开源于https://github.com/xiaomore/D2Dewarp。
原文摘要 · Abstract (English)
Document image dewarping remains a challenging task in the deep learning era. While existing methods have improved by leveraging text line awareness, they typically focus only on a single horizontal dimension. In this paper, we propose a fine-grained deformation perception model that focuses on Dual Dimensions of document horizontal-vertical-lines to improve document Dewarping called D2Dewarp. It can perceive distortion trends in different directions across document details. To combine the horizontal and vertical granularity features, an effective fusion module based on X and Y coordinate is designed to facilitate interaction and constraint between the two dimensions for feature complementarity. Due to the lack of annotated line features in current public dewarping datasets, we also propose an automatic fine-grained annotation method using public document texture images and automatic rendering engine to build a new large-scale distortion training dataset named DocDewarpHV. On three public Chinese and English benchmarks, both quantitative and qualitative results show that our method achieves better rectification results compared with the state-of-the-art methods. The code and dataset are available at https://github.com/xiaomore/D2Dewarp.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。