用知识图谱增强表格医疗数据,自动补全缺失医学特征。
Cross-Domain Feature Expansion for Tabular Medical Data via Knowledge Graphs Injection

- 通过双注意力机制直接处理原始表格数据
- 生成数据在跨领域场景中保持高保真度和真实性
- 适合医疗数据稀缺时的特征补全与生成任务
获取全面的跨领域生物医学特征通常成本高昂且耗时,导致医学研究中严重的数据稀缺。为应对这一挑战,我们提出MedKGTab,一种专为表格医疗数据跨域特征扩展设计的知识注入框架。MedKGTab通过利用已知特征间的统计依赖性与既有的医学关联,推断未采集的生物医学特征。该方法采用行-列双注意力机制,在原始结构化表格数据上直接运行,避免了分词带来的结构损失,精准捕捉数值分布。关键在于,MedKGTab将数据驱动的统计先验与SPOKE生物医学知识图谱融合,实现数据通道与知识通道的最优协同。在此协同中,数据通道的表征受注入的医学知识调制,确保生成数据基于实证医学研究。实验表明,MedKGTab在跨域特征扩展中实现了高数据保真度与真实表示,优于当前最先进的医疗大模型(如Baichuan M3-plus)及专用于医疗数据生成的表格模型。此外,无论是在同一数据集内推断缺失特征,还是在不同医疗队列间泛化,其性能均持续领先。
原文摘要 · Abstract (English)
Acquiring comprehensive cross-domain biomedical profiles is often costly and time-consuming, resulting in severe data scarcity in medical research. To address this challenge, we propose MedKGTab, a knowledge-injected framework specifically engineered for cross-domain feature expansion in tabular medical data. MedKGTab seeks to infer uncollected biomedical features from available ones by exploiting their inherent statistical dependencies and established medical correlations. By employing a row-column dual-attention mechanism, MedKGTab operates directly on raw structured tabular data, inherently capturing exact numerical distributions without the structural loss caused by tokenization. Crucially, MedKGTab integrates data-driven statistical priors with the SPOKE biomedical knowledge graph, achieving an optimal synergy between the data and knowledge channels. Within this synergy, the representations derived from the data channel are modulated by the injected biomedical knowledge, ensuring the final generated data are grounded in empirical medical research. Experimental results demonstrate that MedKGTab achieves high data fidelity and realistic data representation in cross-domain feature expansion. It outperforms both SOTA medical large models (e.g., Baichuan M3-plus) and specialized tabular models designed for medical data generation. Furthermore, MedKGTab consistently delivers superior performance across various data generation scenarios, whether inferring missing features within the same dataset or generalizing across different medical cohorts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。