arXiv:2503.16664cs.CV2025-03

构建历史捷克文档图像数据集,支持无OCR的逻辑分页分割研究。

TextBite: A Historical Czech Document Dataset for Logical Page Segmentation

  • 基于图像域直接分割文本逻辑块,无需依赖OCR或几何信息。
  • 包含8449张页面、78863个语义连贯文本段,覆盖18至20世纪多种文献。
  • 仅用前景文字像素评估,避免背景几何变化干扰,适合文档理解研究者。

逻辑页面分割是文档分析中的关键步骤,有助于提升语义表征、信息检索与文本理解。现有方法通常依赖文本或几何对象,需依赖OCR或精确几何信息。为避免对OCR的依赖,本文将任务定义为纯图像域的分割。为确保评估不受非文本相关几何变化影响,提出仅使用前景文字像素进行评估,忽略所有背景像素。为此,我们构建了TextBite数据集,涵盖18至20世纪的历史捷克文档,包括报纸、词典和手稿等多种布局,共含8,449张页面图像与78,863个逻辑与主题一致的文本段标注。我们提出了结合文本区域检测与关系预测的基线方法。数据集、基线模型与评估框架已公开于https://github.com/DCGM/textbite-dataset。

原文摘要 · Abstract (English)

Logical page segmentation is an important step in document analysis, enabling better semantic representations, information retrieval, and text understanding. Previous approaches define logical segmentation either through text or geometric objects, relying on OCR or precise geometry. To avoid the need for OCR, we define the task purely as segmentation in the image domain. Furthermore, to ensure the evaluation remains unaffected by geometrical variations that do not impact text segmentation, we propose to use only foreground text pixels in the evaluation metric and disregard all background pixels. To support research in logical document segmentation, we introduce TextBite, a dataset of historical Czech documents spanning the 18th to 20th centuries, featuring diverse layouts from newspapers, dictionaries, and handwritten records. The dataset comprises 8,449 page images with 78,863 annotated segments of logically and thematically coherent text. We propose a set of baseline methods combining text region detection and relation prediction. The dataset, baselines and evaluation framework can be accessed at https://github.com/DCGM/textbite-dataset.

文档分析图像分割历史文献无OCR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。