用双曲空间嵌入医学术语,提升肺癌预测准确率
Clinical semantics for lung cancer prediction
- 将SNOMED术语映射到双曲空间生成嵌入表示
- ResNet模型使用10维双曲嵌入后校准性能提升
- 适合关注临床知识融合的医疗AI研究者
现有临床预测模型常忽略临床概念间的语义关系。本研究通过Poincaré嵌入将SNOMED医学术语层次结构映射到低维双曲空间,以改进肺癌发病预测。基于Optum电子病历数据集的回顾性队列,构建了临床知识图谱,并使用黎曼随机梯度下降生成Poincaré嵌入。这些嵌入被引入ResNet和Transformer两种深度学习架构中。评估指标包括区分度(ROC曲线下面积)和校准度(观测与预测概率的平均绝对差)。结果表明,与随机初始化的欧氏嵌入基线相比,预训练的Poincaré嵌入在区分度上实现小幅但稳定的提升;使用10维嵌入的ResNet模型校准性能显著改善,而Transformer模型在各类配置下保持稳定校准。结论:将临床知识图谱嵌入双曲空间并融入深度学习模型,可有效保留临床术语的层级结构,从而提升肺癌发病预测能力,为数据驱动特征提取与既定临床知识的结合提供可行路径。
原文摘要 · Abstract (English)
Background: Existing clinical prediction models often represent patient data using features that ignore the semantic relationships between clinical concepts. This study integrates domain-specific semantic information by mapping the SNOMED medical term hierarchy into a low-dimensional hyperbolic space using Poincaré embeddings, with the aim of improving lung cancer onset prediction. Methods: Using a retrospective cohort from the Optum EHR dataset, we derived a clinical knowledge graph from the SNOMED taxonomy and generated Poincaré embeddings via Riemannian stochastic gradient descent. These embeddings were then incorporated into two deep learning architectures, a ResNet and a Transformer model. Models were evaluated for discrimination (area under the receiver operating characteristic curve) and calibration (average absolute difference between observed and predicted probabilities) performance. Results: Incorporating pre-trained Poincaré embeddings resulted in modest and consistent improvements in discrimination performance compared to baseline models using randomly initialized Euclidean embeddings. ResNet models, particularly those using a 10-dimensional Poincaré embedding, showed enhanced calibration, whereas Transformer models maintained stable calibration across configurations. Discussion: Embedding clinical knowledge graphs into hyperbolic space and integrating these representations into deep learning models can improve lung cancer onset prediction by preserving the hierarchical structure of clinical terminologies used for prediction. This approach demonstrates a feasible method for combining data-driven feature extraction with established clinical knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。