arXiv:2506.02160cs.IRcs.CL2025-06

用AI自动归类生物医学数据元素,提升数据互通效率。

A Dynamic Framework for Semantic Grouping of Common Data Elements (CDE) Using Embeddings and Clustering

  • 用大模型生成语义向量,再用聚类算法分组相似数据项。
  • 在2.4万条数据中识别出118个有意义的类别,分类准确率达90.46%。
  • 适合需要整合多源生物医学数据的研究者使用。

本研究旨在构建一个动态可扩展的框架,以解决异构生物医学数据集间常见数据元素(CDEs)因语义差异、结构变化和上下文依赖带来的整合难题,从而促进数据融合、提升互操作性并加速科学发现。方法基于大语言模型(LLMs)生成上下文感知的文本嵌入,将CDE转化为捕捉语义关系的稠密向量;采用层次化密度聚类(HDBSCAN)对嵌入进行无监督聚类;通过LLM摘要实现自动标签生成;最后利用有监督学习训练分类器,为新或未聚类的CDE分配到已有类别。在包含超24,000个CDE的美国国立卫生研究院国家医学图书馆CDE库上评估,系统在最小聚类大小设为20时识别出118个有意义的聚类。分类器整体准确率达90.46%,在大类别中表现更优。外部验证与重力项目社会决定因素健康领域对比显示高度一致性(调整兰德指数0.52,标准化互信息0.78),表明嵌入能有效表征聚类特征。该灵活可扩展的方法为CDE标准化提供实用解决方案,显著提升数据选择效率,支持持续的数据互操作性。

原文摘要 · Abstract (English)

This research aims to develop a dynamic and scalable framework to facilitate harmonization of Common Data Elements (CDEs) across heterogeneous biomedical datasets by addressing challenges such as semantic heterogeneity, structural variability, and context dependence to streamline integration, enhance interoperability, and accelerate scientific discovery. Our methodology leverages Large Language Models (LLMs) for context-aware text embeddings that convert CDEs into dense vectors capturing semantic relationships and patterns. These embeddings are clustered using Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) to group semantically similar CDEs. The framework incorporates four key steps: (1) LLM-based text embedding to mathematically represent semantic context, (2) unsupervised clustering of embeddings via HDBSCAN, (3) automated labeling using LLM summarization, and (4) supervised learning to train a classifier assigning new or unclustered CDEs to labeled clusters. Evaluated on the NIH NLM CDE Repository with over 24,000 CDEs, the system identified 118 meaningful clusters at an optimized minimum cluster size of 20. The classifier achieved 90.46 percent overall accuracy, performing best in larger categories. External validation against Gravity Projects Social Determinants of Health domains showed strong agreement (Adjusted Rand Index 0.52, Normalized Mutual Information 0.78), indicating that embeddings effectively capture cluster characteristics. This adaptable and scalable approach offers a practical solution to CDE harmonization, improving selection efficiency and supporting ongoing data interoperability.

数据整合语义聚类大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。