提出新基准CAREBench,评估医疗时序数据中风险预测的稳定性与准确性。
Stable Prediction of Adverse Events in Medical Time-Series Data
- 构建多模态输入基准,融合电子病历、心电波形与临床文本。
- 发现大模型在高精度下召回率差,难以兼顾准确与稳定。
- 适合关注临床可信赖连续监测系统的研究者使用。
早期事件预测(EEP)系统持续评估患者即时风险以支持临床决策。为确保床旁可信度,风险轨迹必须准确且时间稳定,仅随新证据变化。然而,现有基准(a)忽略风险评分的稳定性,(b)主要评估表格式输入,未测试轨迹行为。为此,我们提出CAREBench,一个基于多模态输入——表格式电子健康记录(EHR)、心电波形和临床文本——的EEP基准,评估部署可行性,并同时衡量时间稳定性与预测准确性。我们提出一种稳定性度量,量化个体患者风险的短期波动,并基于局部Lipschitz常数惩罚突变震荡。CAREBench涵盖六项预测任务,如脓毒症发作,比较传统学习器、深度序列模型与零样本大语言模型(LLM)。跨任务分析显示,现有方法尤其是LLM,在同时优化准确性和稳定性方面表现不佳,尤其在高精度操作点上召回率显著下降。结果凸显了在连续监测场景中,需开发与证据对齐、轨迹稳定的模型以赢得临床信任。(代码:https://github.com/SeewonChoi/CAREBench。)
原文摘要 · Abstract (English)
Early event prediction (EEP) systems continuously estimate a patient's imminent risk to support clinical decision-making. For bedside trust, risk trajectories must be accurate and temporally stable, shifting only with new, relevant evidence. However, current benchmarks (a) ignore stability of risk scores and (b) evaluate mainly on tabular inputs, leaving trajectory behavior untested. To address this gap, we introduce CAREBench, an EEP benchmark that evaluates deployability using multi-modal inputs-tabular EHR, ECG waveforms, and clinical text-and assesses temporal stability alongside predictive accuracy. We propose a stability metric that quantifies short-term variability in per-patient risk and penalizes abrupt oscillations based on local-Lipschitz constants. CAREBench spans six prediction tasks such as sepsis onset and compares classical learners, deep sequence models, and zero-shot LLMs. Across tasks, existing methods, especially LLMs, struggle to jointly optimize accuracy and stability, with notably poor recall at high-precision operating points. These results highlight the need for models that produce evidence-aligned, stable trajectories to earn clinician trust in continuous monitoring settings. (Code: https://github.com/SeewonChoi/CAREBench.)
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。