构建跨时空文档布局分析数据集,支持历史文档重建与模型训练。
Diachronic Document Dataset for Semantic Layout Analysis
- 基于TEI标准构建7254页多时期多类型文档标注数据
- 1280像素输入使YOLO在该数据集上表现最优
- 模块化设计适合领域定制,通用模型融合子集更优
我们提出一个开源的新颖数据集,用于语义布局分析,旨在支持文档重建流程,并通过与文本编码倡议(TEI)标准映射实现。该数据集包含7,254个标注页面,覆盖1600至2024年间的数字化与原生数字材料,涵盖杂志、文理论文、博士论文、专著、剧本、行政报告等多种文档类型,并按模块化子集划分。通过整合不同时期与体裁内容,该数据集覆盖了多样的布局复杂性及历史结构变迁。模块化设计支持领域特定配置。我们在该数据集上评估目标检测模型,考察输入尺寸和子集训练的影响。结果表明,1280像素输入对YOLO模型最有效,且将子集融入通用模型优于微调预训练权重。
原文摘要 · Abstract (English)
We present a novel, open-access dataset designed for semantic layout analysis, built to support document recreation workflows through mapping with the Text Encoding Initiative (TEI) standard. This dataset includes 7,254 annotated pages spanning a large temporal range (1600-2024) of digitised and born-digital materials across diverse document types (magazines, papers from sciences and humanities, PhD theses, monographs, plays, administrative reports, etc.) sorted into modular subsets. By incorporating content from different periods and genres, it addresses varying layout complexities and historical changes in document structure. The modular design allows domain-specific configurations. We evaluate object detection models on this dataset, examining the impact of input size and subset-based training. Results show that a 1280-pixel input size for YOLO is optimal and that training on subsets generally benefits from incorporating them into a generic model rather than fine-tuning pre-trained weights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。