arXiv:2507.21114cs.IRcs.AI2025-07被引 1

自动分类古籍页面内容,提升数字化处理效率

Page image classification for content-specific data processing

  • 基于AI构建历史文档页面分类系统
  • 支持手写、印刷、图表等多类型内容识别
  • 适合需要分类型处理的数字人文项目

人文领域的数字化项目常产生海量历史文献页面图像,带来人工整理与分析的巨大挑战。这些档案包含多种内容,如手写、打字、印刷文字,以及图画、地图、照片等图形元素和纯文本、表格、表单等版式。高效处理这类异构数据需自动化方法对页面内容进行分类,以支持针对性的下游分析流程。本研究开发并评估了一种专为历史文献页面设计的图像分类系统,利用人工智能与机器学习技术,所选类别旨在支持内容特定的处理工作流,区分需不同分析手段的页面(如文本用OCR,图形用图像分析)。

原文摘要 · Abstract (English)

Digitization projects in humanities often generate vast quantities of page images from historical documents, presenting significant challenges for manual sorting and analysis. These archives contain diverse content, including various text types (handwritten, typed, printed), graphical elements (drawings, maps, photos), and layouts (plain text, tables, forms). Efficiently processing this heterogeneous data requires automated methods to categorize pages based on their content, enabling tailored downstream analysis pipelines. This project addresses this need by developing and evaluating an image classification system specifically designed for historical document pages, leveraging advancements in artificial intelligence and machine learning. The set of categories was chosen to facilitate content-specific processing workflows, separating pages requiring different analysis techniques (e.g., OCR for text, image analysis for graphics)

图像分类古籍数字化内容识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。