用生成式AI与大模型协作,解耦元数据模型的复杂纠缠。
A Generative AI-driven Metadata Modelling Approach
- 将元数据建模视为五个功能联动的表示层级,由本体驱动。
- 揭示各层级隐含的表示多样性,导致模型概念纠缠。
- 提出生成式AI+LLM协同方法,实现元数据模型解耦。
多年来,元数据建模始终是学术图书馆运作的核心。随着生成式人工智能驱动的信息活动和服务日益普及,其重要性进一步提升。然而,元数据建模过程面临诸多挑战,影响其可重用性、跨映射及与其他模型的互操作性。本文认为,问题根源在于一种假设:只需少数核心元数据模型即可满足任何信息服务需求,无论领域内或跨领域的异构性如何。本文提出相反观点,并通过三个关键步骤论证:首先,将图书馆元数据模型重新定义为从感知到意向性定义的五级功能互联表示体系,由本体驱动;其次,揭示每一层级内在的表示多样性,这些多样性累积导致概念纠缠;最后,提出一种基于生成式AI与人类-大语言模型(LLM)协作的元数据建模方法,以解耦每个表示层级中的纠缠,生成概念清晰的元数据模型。全文通过处理癌症信息的代表性图书馆案例进行说明。
原文摘要 · Abstract (English)
Since decades, the modelling of metadata has been core to the functioning of any academic library. Its importance has only enhanced with the increasing pervasiveness of Generative Artificial Intelligence (AI)-driven information activities and services which constitute a library's outreach. However, with the rising importance of metadata, there arose several outstanding problems with the process of designing a library metadata model impacting its reusability, crosswalk and interoperability with other metadata models. This paper posits that the above problems stem from an underlying thesis that there should only be a few core metadata models which would be necessary and sufficient for any information service using them, irrespective of the heterogeneity of intra-domain or inter-domain settings. To that end, this paper advances a contrary view of the above thesis and substantiates its argument in three key steps. First, it introduces a novel way of thinking about a library metadata model as an ontology-driven composition of five functionally interlinked representation levels from perception to its intensional definition via properties. Second, it introduces the representational manifoldness implicit in each of the five levels which cumulatively contributes to a conceptually entangled library metadata model. Finally, and most importantly, it proposes a Generative AI-driven Human-Large Language Model (LLM) collaboration based metadata modelling approach to disentangle the entanglement inherent in each representation level leading to the generation of a conceptually disentangled metadata model. Throughout the paper, the arguments are exemplified by motivating scenarios and examples from representative libraries handling cancer information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。