arXiv:2606.28623cs.LGstat.ML2026-06

用Mamba架构学习病历表示,提升纵向患者分型准确率

Improving Patient Subtyping on Longitudinal Data using Representations from Mamba-based Architecture

论文配图:Improving Patient Subtyping on Longitudinal Data using Representations from Mamba-based Architecture
图 1 · 摘自论文原文
  • 基于自监督Mamba架构提取电子病历时序特征
  • 在真实数据集上优于主流基线模型的分类与聚类表现
  • 适合从事精准医疗与病程建模的研究者参考

利用电子健康记录(EHR)数据对患者进行有效分型(也称分组或聚类),可显著推动精准医学发展。然而,由于EHR数据存在复杂性和不规则性,对其进行时序分型仍具挑战。本文提出一种基于自监督Mamba的模型,用于学习有效的EHR表示,并提升患者分型性能。我们在公开和私有真实世界EHR数据集上评估该模型,基于已有标签进行分类,并利用模型学习到的表示进行患者分型。通过大量实验,我们证明所提模型的设计选择在预测任务上优于多种竞争性基线模型。此外,我们测试了多种聚类方法,结果表明该模型为基于时序EHR数据的患者分型提供了有价值的新见解。

原文摘要 · Abstract (English)

Effective sub-typing (also known as grouping or clustering) of patients using their electronic health record (EHR) data can greatly inform precision medicine efforts. However, subtyping temporal EHR datasets is known to be challenging due to inherent EHR issues, including complexity and irregularity. In this study, we propose a self-supervised Mamba-based model that learns effective EHR representations and enables enhanced patient subtyping. We evaluate the proposed model on public and private real-world EHR datasets to classify the data based on the available labels and subtype patients based on the representations learned from the model. Through an extensive set of experiments, we demonstrate that our model's design choices lead to better performance compared to competitive baseline models for prediction. Moreover, we evaluate several clustering techniques to demonstrate that our findings offer valuable insights into subtyping patients based on temporal records from EHR models\footnote{Our implementations are available at https://github.com/healthylaife/triplet_mamba.

患者分型时序建模MambaEHR分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。