arXiv:2607.14191cs.LG2026-07

用儿童病历数据训练小型模型,提前多年预测罕见病风险。

TEDDY: A Pediatric Foundation Model for Risk Forewarning from ICD-Coded Diagnostic Histories

论文配图:TEDDY: A Pediatric Foundation Model for Risk Forewarning from ICD-Coded Diagnostic Histories
图 1 · 摘自论文原文
  • 基于160万儿童病历,用184万参数模型捕捉诊断时间序列规律。
  • 对797种疾病预测中位AUC达72.0%,罕见病90%超过随机水平。
  • 无需超大模型或全人群数据,适合临床早期预警场景。

儿科电子健康记录蕴含发育相关的临床轨迹,但其在生成式医疗基础模型中的潜力尚未被充分探索。本文提出TEDDY(青少年疾病时序事件解码器),一个184万参数的解码器变压器,基于单机构约7300万条ICD-10诊断记录、160万儿童的数据进行训练。该模型可建模纵向诊断轨迹与就诊时间。预测在首次诊断代码揭示前完成,仅考虑首次发生,并与同性别、同年龄对照组比较。在覆盖16个ICD-10章节的797个疾病初发预测任务中,中位AUC达72.0%,显著优于相同数据下的DenseNet(50.0%)、CNN(57.2%)、RNN(60.1%)和LSTM(62.7%)基线,在96%-99%的任务上表现更优。性能在不同性别和年龄组间保持稳定,尤其在低流行率疾病中表现突出:225种最罕见疾病中有202种(90%)的95%置信区间高于随机水平。预测信号可在首次记录诊断前两年以上仍被检测到,自由分析下中位AUC为59.7%,固定队列敏感性分析中达64.4%。在哮喘与注意力缺陷多动障碍基准测试中,AUC分别为79.3%和84.7%,远超最强对比模型(71.7%),后者为规模大三阶的通用语言模型。就诊时间预测的平均绝对受限生存时间误差为3.0天(365天周期内),但中位与长尾返回间隔仍存在校准偏差。结果表明,儿科诊断历史可作为紧凑生成模型的基础,实现广泛、罕见病及长周期风险预测,无需大规模人口数据或十亿参数模型。

原文摘要 · Abstract (English)

Pediatric electronic health records capture developmentally structured clinical trajectories, yet their potential for generative healthcare foundation models remains largely unexplored. Here we present TEDDY (Temporal Event Decoder for Disease in Youth), a 1.84-million-parameter decoder transformer trained on approximately 73 million ICD-10 diagnoses from 1.6 million children at a single pediatric institution. TEDDY models longitudinal diagnosis trajectories and visit timing. Predictions were made before visit codes were revealed, limited to first occurrences, and evaluated against sex- and age-matched controls. Across 797 disease-onset prediction tasks spanning 16 ICD-10 chapters, TEDDY achieved a median AUC of 72.0%, outperforming same-data DenseNet (50.0%), CNN (57.2%), RNN (60.1%), and LSTM (62.7%) baselines on 96-99% of tasks. Performance held across sex and age and was strongest among lower-prevalence diagnoses; 202 of the 225 rarest conditions (90%) had 95% confidence intervals above chance. Predictive signal remained detectable more than two years before first recorded diagnosis, with median AUCs of 59.7% in the unrestricted analysis and 64.4% in a fixed-cohort sensitivity analysis. In asthma and attention-deficit/hyperactivity disorder benchmarks, AUCs were 79.3% and 84.7%, compared with 62.7% and 71.7% for the strongest comparators, including a general-purpose language model three orders of magnitude larger. Visit-timing predictions had a 3.0-day mean absolute restricted mean survival-time error over 365 days, although median and long-tail return intervals remained miscalibrated. Together, these results establish pediatric diagnostic histories as a substrate for compact generative models supporting broad, rare-disease, and long-horizon risk forecasting without population-scale data or billion-parameter models.

医疗AI风险预测小模型罕见病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。