arXiv:2603.13231cs.LGcs.CL2026-03

图注意力模型提升电子病历预测,但评估仍存盲区。

Translational Gaps in Graph Transformers for Longitudinal EHR Prediction: A Critical Appraisal of GT-BEHRT

  • 用图结构建模就诊数据,捕捉临床事件间关系。
  • 心衰预测365天内AUROC达94.37,表现优异。
  • 适合关注医疗AI可解释性与落地的科研人员。

基于Transformer的模型通过大规模自监督预训练提升了纵向电子健康记录(EHR)的预测能力。然而,多数EHR Transformer将每次就诊视为无序代码集合,难以捕捉就诊内的有意义关系。图-变压器方法旨在通过建模就诊级结构来弥补这一不足,同时保留学习长期时间模式的能力。本文对GT-BEHRT——一种在MIMIC-IV重症结局和All of Us研究计划中用于心衰预测的图变压器架构进行了批判性评估。我们考察其报告的性能提升是否真正源于架构优势,以及评估方法能否支持其稳健性和临床相关性。从表示设计、预训练策略、队列构建透明度、超越区分能力的评估、公平性分析、可复现性及部署可行性七个维度展开分析。尽管GT-BEHRT在365天内心衰预测中表现出色,AUROC为94.37±0.20,AUPRC为73.96±0.83,F1为64.70±0.85,但仍存在显著缺陷:缺乏校准分析,公平性评估不完整,对队列选择敏感,跨表型和预测时长的分析有限,且对实际部署考虑不足。总体而言,GT-BEHRT在EHR表示学习方面具有重要架构进展,但在校准、公平性和部署方面需更严格评估,方能可靠支持临床决策。

原文摘要 · Abstract (English)

Transformer-based models have improved predictive modeling on longitudinal electronic health records through large-scale self-supervised pretraining. However, most EHR transformer architectures treat each clinical encounter as an unordered collection of codes, which limits their ability to capture meaningful relationships within a visit. Graph-transformer approaches aim to address this limitation by modeling visit-level structure while retaining the ability to learn long-term temporal patterns. This paper provides a critical review of GT-BEHRT, a graph-transformer architecture evaluated on MIMIC-IV intensive care outcomes and heart failure prediction in the All of Us Research Program. We examine whether the reported performance gains reflect genuine architectural benefits and whether the evaluation methodology supports claims of robustness and clinical relevance. We analyze GT-BEHRT across seven dimensions relevant to modern machine learning systems, including representation design, pretraining strategy, cohort construction transparency, evaluation beyond discrimination, fairness assessment, reproducibility, and deployment feasibility. GT-BEHRT reports strong discrimination for heart failure prediction within 365 days, with AUROC 94.37 +/- 0.20, AUPRC 73.96 +/- 0.83, and F1 64.70 +/- 0.85. Despite these results, we identify several important gaps, including the lack of calibration analysis, incomplete fairness evaluation, sensitivity to cohort selection, limited analysis across phenotypes and prediction horizons, and limited discussion of practical deployment considerations. Overall, GT-BEHRT represents a meaningful architectural advance in EHR representation learning, but more rigorous evaluation focused on calibration, fairness, and deployment is needed before such models can reliably support clinical decision-making.

图神经网络医疗预测EHR建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。