arXiv:2602.22228cs.LG2026-02

用患者自述症状构建图模型,实现高精度早期中风风险预警

Patient-Centered, Graph-Augmented Artificial Intelligence-Enabled Passive Surveillance for Early Stroke Risk Detection in High-Risk Individuals

  • 基于患者语言构建症状图谱,结合异构图神经网络与统计模型识别风险模式
  • 90天窗口内灵敏度达0.72,特异性与阳性预测值均为1.00,误报极低
  • 适合糖尿病等高危人群的低负担、高精度被动监测,助力临床早期干预

中风每年影响数百万人,但症状识别不佳常导致就医延迟。为弥补风险识别缺口,我们开发了一套基于糖尿病患者自述症状的被动监测系统,用于早期中风风险检测。通过构建基于患者自身语言的症状分类体系,并采用双机器学习流程(异构图神经网络与EN/LASSO),识别出与后续中风相关的症状模式。将结果转化为融合症状相关性与时序临近性的混合风险筛查系统,在电子健康记录(EHR)模拟中评估了3-90天窗口。在保守阈值下,系统实现高特异性(1.00)和预设患病率调整后的阳性预测值(1.00),灵敏度为0.72,该表现是精度优先的合理权衡,且在90天窗口表现最佳。仅依赖患者自述语言即可实现高精度、低负担的早期中风风险检测,为高危人群提供宝贵的临床评估与干预时间窗口。

原文摘要 · Abstract (English)

Stroke affected millions annually, yet poor symptom recognition often delayed care-seeking. To address risk recognition gap, we developed a passive surveillance system for early stroke risk detection using patient-reported symptoms among individuals with diabetes. Constructing a symptom taxonomy grounded in patients own language and a dual machine learning pipeline (heterogeneous GNN and EN/LASSO), we identified symptom patterns associated with subsequent stroke. We translated findings into a hybrid risk screening system integrating symptom relevance and temporal proximity, evaluated across 3-90 day windows through EHR-based simulations. Under conservative thresholds, intentionally designed to minimize false alerts, the screening system achieved high specificity (1.00) and prevalence-adjusted positive predictive value (1.00), with good sensitivity (0.72), an expected trade-off prioritizing precision, that was highest in 90-day window. Patient-reported language alone supported high-precision, low-burden early stroke risk detection, that could offer a valuable time window for clinical evaluation and intervention for high-risk individuals.

中风预警被动监测图神经网络糖尿病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。