arXiv:2606.17355cs.CV2026-06

用增强方法提升低资源页面布局分类准确率

Complex Layout Classification in the Wild: A Low-Resource Approach with Layout-Preserving Augmentations

论文配图:Complex Layout Classification in the Wild: A Low-Resource Approach with Layout-Preserving Augmentations
图 1 · 摘自论文原文
  • 设计保留版式结构的增强策略,聚焦全局几何特征
  • 在标注稀缺下达到85.2%准确率,显著优于基线
  • 适合文档分析、古籍数字化等低资源场景

许多数字文献因标注稀少、扫描噪声大、分辨率低或版式结构复杂,导致自动转录质量差。为应对低资源语言的分类难题,我们构建了一个包含八类版式类型的复杂布局数据集,并提出一种基于CNN的新型训练策略。该方法采用领域感知的强增强技术:通过窄长各向异性高斯掩码抑制局部文字细节,保留关键分隔区域,使模型聚焦全局版式结构;同时引入反射引发的标签变换,在保持非对称类别标签一致性的同时丰富训练分布。实验表明,针对版式的特定增强可显著提升极端标注稀缺条件下的页面级版式分类性能。

原文摘要 · Abstract (English)

Many digitized corpora suffer from low resources because annotations may be scarce, page scans are noisy and of poor resolution, or layouts are structurally complex in ways that negatively affect the quality of automatic transcription. Developing robust classification models for low-resource languages is inhibited by the lack of large-scale annotated data and by the frequent semantic complexity of page layouts. To this end, we have curated a complex-layout dataset, manually classified into eight distinct layout types based on their separator regions. To overcome data scarcity, we propose a novel training strategy in the form of a CNN-based classifier that employs strong, domain-aware augmentations to improve generalization. We utilize narrow anisotropic Gaussian masking to suppress incidental textual details while preserving essential separations, compelling the model to learn global geometric arrangements. Additionally, we implement reflection-induced label transformations to enrich the training distribution while maintaining label consistency across asymmetric categories. The results demonstrate that layout-specific augmentations can substantially improve page-level layout classification under severe annotation scarcity.

版式分类低资源数据增强CNN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。