arXiv:2602.03730stat.MLcs.LG2026-02被引 2

提出两种新方法,让医疗模型预测更准更快,尤其擅长罕见病风险评估。

Efficient Generative Prediction for EHR Foundation Models: The SCOPE and REACH Estimators

  • 利用未来病历片段的概率分布,设计新估算器提升效率
  • 在11个临床任务中,用更少的计算量达到与100次采样相当的精度
  • 适合需要高精度、低资源消耗的医疗风险预测场景

基于分词电子健康记录(EHR)时间线训练的生成式基础模型,可通过蒙特卡洛采样模拟未来病程进行临床结局预测。然而该方法存在三个相互关联的局限:估计分布稀疏、计算成本极高、采样方差大。本文提出两种新估算器——条件结果概率之和估算器(SCOPE)与预期条件风险的风险估算器(REACH),充分利用了传统蒙特卡洛方法忽视的下一词概率分布。理论上证明二者均为无偏估计,且REACH在任何模型和结局下均能保证比蒙特卡洛更低的方差;同时,REACH是任意保留非结局词分布的重要性采样方案的Rao-Blackwell化。实验表明,在MIMIC-IV和芝加哥大学医疗系统中的11个临床重要结局上,SCOPE与REACH可实现与100样本蒙特卡洛相当的准确率,平均减少2.5至3.4倍的生成令牌数,对最罕见结局的减少超过80倍,且校准性始终如一。由于SCOPE可复用单次采样池应对任意多个结局而无需额外开销,而REACH提供每个任务的方差保障,二者互补,显著降低生成式EHR基础模型的推理预算,尤其适用于罕见但高影响的医疗结局预测。

原文摘要 · Abstract (English)

Generative foundation models trained on tokenized electronic health record (EHR) timelines show promise for clinical outcome prediction via Monte Carlo sampling of simulated future trajectories. However, this approach suffers from three coupled limitations: sparse estimate distributions that poorly differentiate patient risk levels, extreme computational cost, and high sampling variance. We propose two new estimators that leverage next-token probability distributions underutilized by standard Monte Carlo: the Sum of Conditional Outcome Probability Estimator (SCOPE) and Risk Estimation from Anticipated Conditional Hazards (REACH). We prove both are unbiased, that REACH guarantees variance reduction over Monte Carlo for any model and outcome, and that REACH is a Rao-Blackwellization of any naive importance sampling scheme that preserves the non-outcome token distribution. Empirically, across $11$ clinically important outcomes in MIMIC-IV and the UChicago health system, SCOPE and REACH match $100$-sample Monte Carlo accuracy with median token reductions of $2.5\times$ to $3.4\times$ and reductions exceeding $80\times$ for the rarest outcomes, with calibration preserved throughout. Because SCOPE reuses a single sampled pool across an arbitrary number of outcomes at no marginal generation cost while REACH provides a per-task variance guarantee, the two estimators are complementary in deployment and together meaningfully reduce the inference budget required for generative EHR foundation models, particularly for rare, high-impact outcomes in healthcare.

医疗预测生成模型高效推理风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。