arXiv:2603.00147cs.CVcs.IR2026-03

用AI技术提升古籍图文分割与标注效率,助力历史文献数字化

Leveraging GenAI for Segmenting and Labeling Centuries-old Technical Documents

  • 结合SAM2、Florence2和ChatGPT实现古籍图像自动分割与标签生成
  • 在16-17世纪造船文献上验证效果,初步提升古籍编目与检索能力
  • 引入专业航海术语本体库,增强标签准确性,适合历史文献数字化研究者

图像分割与识别是图像处理领域的成熟技术,分别用于定位图像区域和识别其中对象。现代图像因训练数据丰富,已达到高精度;但对数百年历史文献的图像处理仍具挑战,主要因训练数据稀缺且领域高度专业化。然而,实现此类文献的自动分割与识别对知识的自动化整理、编目与传播至关重要,可使珍贵典籍更易被学者与公众获取。本文报告了针对16至17世纪造船文献(即大航海时代)的初步研究:采用SAM2进行图像分割,使用Florence2与ChatGPT进行标签生成,并结合专用本体ontoShip与术语表glosShip以增强标注质量。初步结果表明,融合这些技术能有效提升珍贵历史文献的编目与检索效率。同时讨论了当前方法的局限性及未来改进方向。

原文摘要 · Abstract (English)

Image segmentation and image recognition are well established computational techniques in the broader discipline of image processing. Segmentation allows to locate areas in an image, while recognition identifies specific objects within an image. These techniques have shown remarkable accuracy with modern images, mainly because the amount of training data is vast. Achieving similar accuracy in digitized images of centuries-old documents is more challenging. This difficulty is due to two main reasons: first, the lack of sufficient training data, and second, because the degree of specialization in a given domain. Despite these limitations, the ability to segment and recognize objects in these collections is important for automating the curation, cataloging, and dissemination of knowledge, making the contents of priceless collections accessible to scholars and the general public. In this paper, we report on our ongoing work in segmenting and labeling images pertaining to shipbuilding treatises from the XVI and XVII centuries, a historical period known as the Age of Exploration. To this end, we leverage SAM2 for image segmentation; Florence2 and ChatGPT for labeling; and a specialized ontology ontoShip and glossary glosShip of nautical architecture for enhancing the labeling process. Preliminary results demonstrate the potential of marrying these technologies for improving curation and retrieval of priceless historical documents. We also discuss the challenges and limitations encountered in this approach and ideas on how to overcome them in the future.

古籍数字化图像分割GenAI应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。