用大模型自动标注医疗数据元信息,解决异构数据难用问题
H2: A Dual Hybrid Semantic Data Lake Architecture for Medical Data Harmonization with Human-In-the-Loop verified, LLM Driven Metadata Annotation System

- 构建双层语义数据湖,结合知识图谱与大模型实现灵活元数据标注
- 通过人类验证的大模型标注,使未标记数据可适配特定机器学习方法
- 适合医疗数据治理、临床研究团队快速构建可用数据资产
医疗数据在模态(如图像、文本、时间序列)、表结构和非结构化文本等方面具有高度异质性。数据湖虽能集中存储各类数据而无需预设模式,但易陷入‘数据沼泽’困境。为此,本文提出一种双层语义数据湖架构,利用知识图谱动态建模关系,并引入大语言模型(LLM)驱动的生成式元数据标注系统,对未标记数据进行自动化标注。该系统通过人机协同验证,识别数据类型与适用的机器学习操作之间的匹配性,提升数据可用性。该架构支持跨模态、跨机构数据的语义对齐与可操作性评估,为临床研究和数据驱动决策提供可靠基础。
原文摘要 · Abstract (English)
Medical data, by its nature, exhibit a high degree of heterogeneity on multiple levels ranging from (a) different modalities like images, text and time series, (b) diverse tabular schemata introduced by institutions and (c) completely unstructured textual information data provided by healthcare professionals. Data lakes are often used in medical data storage to consolidate all heterogeneous diverse data in a single, central location, where it can be saved "as is", without the need to impose a schema like a data warehouse does. Despite their flexibility, though, data lakes are notorious for the "data swamp" failure. Thus, providing a reliable data harmonization mechanism through metadata, without compromising integrity or flexibility, is a real challenge. To this end, knowledge graphs have attracted attention since they provide a dynamic way to depict relationships without a rigid schema-on-write approach. Additionally, another rigorous task relies on the interoperability of data: application of appropriate ML techniques on such a diverse nature of data is not an easy task, as a domain expert must decide the efficacy of a method to a specific data type or dataset. Metadata annotation can aid by tagging applicable operations, however this requires manual intervention, not to mention the plethora of existing datasets which lack such information. To tackle both challenges, in this paper, we propose a semantic data lake architecture that promotes data harmonization and incorporates a generative annotation process (i.e. LLMs) of non-labeled metadata collections to support the application of meaningful ML techniques. Building on top of this approach, we create a higher level of knowledge, identifying suitability of data with respect to applicable ML operations based on their data nature...
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。