用临床笔记预测糖尿病风险,兼顾时间细节与隐私安全。
Early Risk Prediction with Temporally and Contextually Grounded Clinical Language Processing
- 构建分层时序图网络,捕捉病历中的时间事件结构和医疗知识
- 在近期内预测准确率最高,且对真实病例敏感度提升23%
- 适合关注慢病早筛、注重可解释性与隐私保护的医疗研究者
电子健康记录中的临床笔记包含丰富的事件时间信息、医生推理和生活方式因素,常被结构化数据忽略。利用这些文本进行预测建模,有助于慢性病的早期识别。然而,其存在长文本、事件分布不均、复杂时间依赖、隐私限制和资源约束等核心自然语言处理挑战。本文提出两种互补方法:首先,设计HiTGNN——一种分层时序图神经网络,整合病历内的时间事件结构、就诊间动态变化及医学知识,实现细粒度的时间轨迹建模;其次,提出ReVeAL——一种轻量级测试时框架,将大模型的推理能力蒸馏至小型验证模型。在基于私有与公开医院语料构建的时序真实队列上,用于2型糖尿病(T2D)的机会性筛查,HiTGNN在近期内预测表现最优,同时保障隐私并减少对大型专有模型的依赖;ReVeAL提升了对真实T2D病例的敏感度,且保留可解释性。消融实验证明了时间结构与知识增强的价值,公平性分析显示HiTGNN在各亚组间表现更均衡。
原文摘要 · Abstract (English)
Clinical notes in Electronic Health Records (EHRs) capture rich temporal information on events, clinician reasoning, and lifestyle factors often missing from structured data. Leveraging them for predictive modeling can be impactful for timely identification of chronic diseases. However, they present core natural language processing (NLP) challenges: long text, irregular event distribution, complex temporal dependencies, privacy constraints, and resource limitations. We present two complementary methods for temporally and contextually grounded risk prediction from longitudinal notes. First, we introduce HiTGNN, a hierarchical temporal graph neural network that integrates intra-note temporal event structures, inter-visit dynamics, and medical knowledge to model patient trajectories with fine-grained temporal granularity. Second, we propose ReVeAL, a lightweight test-time framework that distills LLMs' reasoning into smaller verifier models. Applied to opportunistic screening for Type 2 Diabetes (T2D) using temporally realistic cohorts curated from private and public hospital corpora, HiTGNN achieves the highest predictive accuracy, especially for near-term risk, while preserving privacy and limiting reliance on large proprietary models. ReVeAL enhances sensitivity to true T2D cases and retains explanatory reasoning. Our ablations confirm the value of temporal structure and knowledge augmentation, and fairness analysis shows HiTGNN performs more equitably across subgroups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。