自适应融合多任务,提升文档图像扭曲矫正精度
Document Image Rectification Bases on Self-Adaptive Multitask Fusion
- 设计自适应多任务融合模块,增强任务间特征互补性
- 在三个基准上显著提升矫正效果,最高达9.6%性能增益
- 适合需要高精度文档图像预处理的场景,如OCR系统
扭曲文档图像矫正对实际文档理解任务(如版面分析和文本识别)至关重要。然而,现有多任务方法(如背景去除、3D坐标预测、文本行分割)常忽视任务间的互补特征及其相互作用。为此,我们提出一种自适应可学习多任务融合矫正网络SalmRec。该网络引入任务间特征聚合模块,自适应增强几何失真感知能力,提升特征互补性并降低负干扰。同时,设计门控机制有效平衡全局任务与局部任务间的特征融合。在两个英文基准(DIR300、DocUNet)和一个中文基准(DocReal)上的实验表明,本方法显著提升矫正性能。消融实验证明不同任务对去扭曲有正向贡献,且所提模块有效性得到验证。
原文摘要 · Abstract (English)
Deformed document image rectification is essential for real-world document understanding tasks, such as layout analysis and text recognition. However, current multi-task methods -- such as background removal, 3D coordinate prediction, and text line segmentation -- often overlook the complementary features between tasks and their interactions. To address this gap, we propose a self-adaptive learnable multi-task fusion rectification network named SalmRec. This network incorporates an inter-task feature aggregation module that adaptively improves the perception of geometric distortions, enhances feature complementarity, and reduces negative interference. We also introduce a gating mechanism to balance features both within global tasks and between local tasks effectively. Experimental results on two English benchmarks (DIR300 and DocUNet) and one Chinese benchmark (DocReal) demonstrate that our method significantly improves rectification performance. Ablation studies further highlight the positive impact of different tasks on dewarping and the effectiveness of our proposed module.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。