arXiv:2607.15721cs.LG2026-07中稿 · ACM ICCA 2026

跨数据源预测糖尿病、高血压和心血管病,兼顾准确性与可靠性。

CardioMeta: Calibrated Multi-Task Prediction of Diabetes, Hypertension, and Cardiovascular Disease Across Population and EHR Data

论文配图:CardioMeta: Calibrated Multi-Task Prediction of Diabetes, Hypertension, and Cardiovascular Disease Across Population and EHR Data
图 1 · 摘自论文原文
  • 多任务联合建模,共享编码器+疾病特异性门控头,避免标签泄露。
  • 在NHANES上实现0.839宏AUROC、0.024期望校准误差,表现稳定。
  • 强调可复现性与校准概率,适合医疗数据融合与模型可信评估场景。

心血管代谢性疾病因糖尿病、高血压和心血管疾病常共现且共享代谢、血管、人口统计与行为因素,仍是导致可预防发病率的主要原因。现有机器学习研究多聚焦单一数据集的判别性能,却忽视标签泄露、校准性、时间鲁棒性、外部可迁移性与子群体可靠性问题。本文提出CardioMeta,一种跨人群调查与电子健康记录(EHR)数据的校准型多任务预测框架,用于联合预测三类疾病。研究使用NHANES进行群体水平建模与时间验证,MIMIC-IV用于应对显著分布偏移的EHR领域评估。为减少循环标签重构,主分析中排除疾病定义变量用于对应预测头,全临床特征设置仅作为敏感性分析保留。CardioMeta结合共享心血管代谢编码器与疾病特异的门控头,并引入后处理概率校准。在无标签泄露的时间验证中,模型取得0.839宏AUROC、0.536宏AUPRC、0.614宏F1与0.024期望校准误差,较强梯度提升与神经网络基线有小幅但一致提升。在MIMIC-IV上的外部评估显示明显性能下降,有限微调部分恢复性能。结果表明,多任务心血管代谢建模的核心价值不在于虚高准确率,而在于可复现的标签泄露控制、校准的概率输出与跨异构医疗数据源的透明可靠性报告。

原文摘要 · Abstract (English)

Cardiometabolic diseases remain among the most persistent drivers of preventable morbidity because diabetes, hypertension, and cardiovascular disease frequently co-occur and share metabolic, vascular, demographic, and behavioral determinants. Existing machine learning studies for chronic disease prediction often emphasize discrimination on a single dataset, while underreporting label leakage, calibration, temporal robustness, external transportability, and subgroup reliability. This paper presents CardioMeta, a calibrated multi-task framework for joint prediction of diabetes, hypertension, and cardiovascular disease across population survey and electronic health record (EHR) data. The study uses NHANES for population-level model development and temporal validation, and MIMIC-IV for EHR-domain evaluation under substantial distribution shift. To reduce circular label reconstruction, the primary analysis excludes disease-defining variables from the corresponding prediction heads, while a full-clinical feature setting is retained only as sensitivity analysis. CardioMeta combines a shared cardiometabolic encoder with disease-specific gated heads and post-hoc probability calibration. In the leakage-reduced temporal validation setting, the model achieved a macro-AUROC of 0.839, macro-AUPRC of 0.536, macro-F1 of 0.614, and expected calibration error of 0.024, with modest but consistent improvements over strong gradient-boosting and neural tabular baselines. External evaluation on MIMIC-IV showed clear degradation under domain shift, while limited fine-tuning partially recovered performance. The findings indicate that the principal value of multi-task cardiometabolic modeling lies not in inflated accuracy, but in reproducible leakage control, calibrated probabilities, and transparent reliability reporting across heterogeneous healthcare data sources.

多任务学习疾病预测医疗数据模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。