arXiv:2412.20088cs.CVcs.AI2024-12被引 1

用大模型自动收集散落的考古文物目录,提升数据整合效率。

An archaeological Catalog Collection Method Based on Large Vision-Language Models

  • 分三步:定位文档、理解区块内容、匹配图文信息
  • 在大宝沟和庙子沟陶器目录上验证,准确率显著优于现有方法
  • 适合考古数据数字化、文化遗产研究者使用

考古目录包含文物图像、形制描述和发掘信息等关键要素,对研究文物演化与文化传承至关重要。这些数据广泛分散于各类出版物中,亟需自动化收集方法。然而,现有大视觉语言模型及其衍生采集方法在处理考古目录时,面临图像精准检测与多模态匹配困难的问题,导致自动化采集难以实现。为此,我们提出一种基于大视觉语言模型的新型考古目录采集方法,采用三个模块协同的流程:文档定位、区块理解与区块匹配。通过在大宝沟与庙子沟陶器目录上的实际数据采集及对比实验,验证了该方法的有效性,为考古目录的自动化采集提供了可靠解决方案。

原文摘要 · Abstract (English)

Archaeological catalogs, containing key elements such as artifact images, morphological descriptions, and excavation information, are essential for studying artifact evolution and cultural inheritance. These data are widely scattered across publications, requiring automated collection methods. However, existing Large Vision-Language Models (VLMs) and their derivative data collection methods face challenges in accurate image detection and modal matching when processing archaeological catalogs, making automated collection difficult. To address these issues, we propose a novel archaeological catalog collection method based on Large Vision-Language Models that follows an approach comprising three modules: document localization, block comprehension and block matching. Through practical data collection from the Dabagou and Miaozigou pottery catalogs and comparison experiments, we demonstrate the effectiveness of our approach, providing a reliable solution for automated collection of archaeological catalogs.

考古数据大模型图像理解数据采集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。