收录7550页历史文档,标注25类非文本元素,助力版面分析研究。
AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization
- 基于捷克图像文档处理法,人工标注25类非文本元素边界框。
- 包含1485至现代的捷克德语历史文献,覆盖晚19世纪至20世纪。
- 适合文档分析、版面理解及历史数字人文研究者使用。
我们介绍AnnoPage数据集,这是一个包含7,550页历史文档的新数据集,主要来自捷克和德国,时间跨度从1485年至今,重点覆盖19世纪末至20世纪初。该数据集旨在支持文档版面分析与目标检测研究。每页均以轴对齐边界框(AABB)标注25类非文本元素,如图片、地图、装饰性元素或图表,遵循捷克图像文档处理方法。标注由专业图书管理员完成,确保准确性和一致性。数据集还整合了多个主要的历史文档数据集,以增强多样性并保持连续性。数据集分为开发集与测试集,测试集按类别分布精心选取。我们提供了基于YOLO和DETR的目标检测基线结果,为未来研究提供参考。AnnoPage数据集已公开于Zenodo(https://doi.org/10.5281/zenodo.12788419),附带以YOLO格式提供的真实标签。
原文摘要 · Abstract (English)
We introduce the AnnoPage Dataset, a novel collection of 7,550 pages from historical documents, primarily in Czech and German, spanning from 1485 to the present, focusing on the late 19th and early 20th centuries. The dataset is designed to support research in document layout analysis and object detection. Each page is annotated with axis-aligned bounding boxes (AABB) representing elements of 25 categories of non-textual elements, such as images, maps, decorative elements, or charts, following the Czech Methodology of image document processing. The annotations were created by expert librarians to ensure accuracy and consistency. The dataset also incorporates pages from multiple, mainly historical, document datasets to enhance variability and maintain continuity. The dataset is divided into development and test subsets, with the test set carefully selected to maintain the category distribution. We provide baseline results using YOLO and DETR object detectors, offering a reference point for future research. The AnnoPage Dataset is publicly available on Zenodo (https://doi.org/10.5281/zenodo.12788419), along with ground-truth annotations in YOLO format.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。