arXiv:2509.03662cs.CL2025-09被引 1

用临床数据挖掘术语共现与语义关系,发现嵌入模型能补全未记录的临床概念。

Semantic Analysis of SNOMED CT Concept Co-occurrences in Clinical Documentation using MIMIC-IV

  • 结合共现统计与临床嵌入模型分析术语关联
  • 嵌入模型预测的术语常被后续文档记录,准确率高
  • 可辅助临床注释、疾病分型和决策支持

临床笔记包含丰富的叙事信息,但其非结构化格式给大规模分析带来挑战。标准化术语如SNOMED CT虽提升互操作性,但概念间的共现模式与语义相似性关系仍不明确。本研究基于MIMIC-IV数据库,利用归一化点互信息(NPMI)和预训练嵌入模型(如ClinicalBERT、BioBERT),探究概念共现频率与语义相似性的关联,评估嵌入模型是否能识别潜在缺失概念,并分析其在时间与专科间的演变规律。结果表明,共现与语义相似性相关性较弱,但嵌入模型能捕捉临床意义明确的关联,且其建议常与后续文档记录一致。概念嵌入聚类生成了症状、检验、诊断、心血管病等清晰临床主题,对应患者表型与诊疗模式。此外,共现模式与死亡率、再入院等结局相关,凸显该方法在提升文档完整性、揭示隐含临床关系及支持决策与表型分析中的应用价值。

原文摘要 · Abstract (English)

Clinical notes contain rich clinical narratives but their unstructured format poses challenges for large-scale analysis. Standardized terminologies such as SNOMED CT improve interoperability, yet understanding how concepts relate through co-occurrence and semantic similarity remains underexplored. In this study, we leverage the MIMIC-IV database to investigate the relationship between SNOMED CT concept co-occurrence patterns and embedding-based semantic similarity. Using Normalized Pointwise Mutual Information (NPMI) and pretrained embeddings (e.g., ClinicalBERT, BioBERT), we examine whether frequently co-occurring concepts are also semantically close, whether embeddings can suggest missing concepts, and how these relationships evolve temporally and across specialties. Our analyses reveal that while co-occurrence and semantic similarity are weakly correlated, embeddings capture clinically meaningful associations not always reflected in documentation frequency. Embedding-based suggestions frequently matched concepts later documented, supporting their utility for augmenting clinical annotations. Clustering of concept embeddings yielded coherent clinical themes (symptoms, labs, diagnoses, cardiovascular conditions) that map to patient phenotypes and care patterns. Finally, co-occurrence patterns linked to outcomes such as mortality and readmission demonstrate the practical utility of this approach. Collectively, our findings highlight the complementary value of co-occurrence statistics and semantic embeddings in improving documentation completeness, uncovering latent clinical relationships, and informing decision support and phenotyping applications.

临床文本语义嵌入共现分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。