arXiv:2604.01841cs.AI2026-04被引 3

针对电子病历临床风险预测,提出对齐任务的检索框架,提升低数据下的预测鲁棒性。

Retrieval-aligned Tabular Foundation Models Enable Robust Clinical Risk Prediction in Electronic Health Records Under Real-world Constraints

  • 设计任务对齐检索机制,用监督嵌入与轻量适配器增强检索质量。
  • 在极端类别不平衡下,AUPRC提升最高达12.2%,数据越复杂增益越大。
  • 适合医疗数据少、样本不均衡场景的临床预测模型开发者参考。

从结构化电子健康记录(EHR)中进行临床预测面临高维、异质性、类别不平衡和分布偏移等挑战。尽管表格上下文学习(TICL)与检索增强方法在通用基准上表现良好,其在真实临床环境中的行为仍不明确。我们构建了一个多队列EHR基准,对比经典模型、深度表格模型及TICL模型在不同数据规模、特征维度、结局罕见性和跨队列泛化能力下的表现。基于PFN的TICL模型在低数据情况下具有样本效率,但随着异质性和不平衡性的增加,简单距离检索会显著退化。为此,我们提出AWARE——一种使用监督嵌入学习与轻量适配器的任务对齐检索框架。AWARE在极端不平衡条件下使AUPRC提升最高达12.2%,且随着数据复杂度增加,性能增益持续扩大。结果揭示:检索质量与检索-推理对齐是部署表格上下文学习于临床预测的关键瓶颈。

原文摘要 · Abstract (English)

Clinical prediction from structured electronic health records (EHRs) is challenging due to high dimensionality, heterogeneity, class imbalance, and distribution shift. While tabular in-context learning (TICL) and retrieval-augmented methods perform well on generic benchmarks, their behavior in clinical settings remains unclear. We present a multi-cohort EHR benchmark comparing classical, deep tabular, and TICL models across varying data scale, feature dimensionality, outcome rarity, and cross-cohort generalization. PFN-based TICL models are sample-efficient in low-data regimes but degrade under naive distance-based retrieval as heterogeneity and imbalance increase. We propose AWARE, a task-aligned retrieval framework using supervised embedding learning and lightweight adapters. AWARE improves AUPRC by up to 12.2% under extreme imbalance, with gains increasing with data complexity. Our results identify retrieval quality and retrieval-inference alignment as key bottlenecks for deploying tabular in-context learning in clinical prediction.

临床预测表格学习检索增强电子病历

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。