用AI和知识图谱提升古籍数字资源的可搜性和关联性
Knowledge Graphs for Digitized Manuscripts in Jagiellonian Digital Library Application
- 结合计算机视觉与语义网技术,自动提取古籍内容信息
- 构建知识图谱实现跨馆藏古籍的智能关联
- 适合数字人文、图书馆学与AI交叉研究者
数字化文化遗产对历史文物保存和公众获取至关重要。美术馆、图书馆、档案馆与博物馆(GLAM机构)正积极数字化馆藏,形成大规模数字资源。这些资源虽配有元数据描述物品信息,但通常不包含具体内容。雅盖隆数字图书馆作为典型代表,通过OAI-PMH等协议提供数据访问。然而,元数据的完整性和标准化仍是重大挑战,限制了资源的可检索性与跨馆藏关联。为此,本文探索融合计算机视觉(CV)、人工智能(AI)与语义网技术的综合方法,以丰富元数据并为数字化手稿与早期印刷品构建知识图谱。
原文摘要 · Abstract (English)
Digitizing cultural heritage collections has become crucial for preservation of historical artifacts and enhancing their availability to the wider public. Galleries, libraries, archives and museums (GLAM institutions) are actively digitizing their holdings and creates extensive digital collections. Those collections are often enriched with metadata describing items but not exactly their contents. The Jagiellonian Digital Library, standing as a good example of such an effort, offers datasets accessible through protocols like OAI-PMH. Despite these improvements, metadata completeness and standardization continue to pose substantial obstacles, limiting the searchability and potential connections between collections. To deal with these challenges, we explore an integrated methodology of computer vision (CV), artificial intelligence (AI), and semantic web technologies to enrich metadata and construct knowledge graphs for digitized manuscripts and incunabula.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。