arXiv:2607.29090cs.LG2026-07综述

梳理手术风险预测的机器学习方法,发现数据和评估标准严重缺失。

What Is Missing in Surgical Risk Stratification and Outcome Prediction: A Scoping Review of End-to-End Machine Learning Approaches

  • 系统分析190项研究的全流程,从数据预处理到模型评估
  • 仅三分之一研究用可解释性方法,多数依赖单一中心数据
  • 缺乏公开数据集和统一评测标准,难复现且临床应用受限

术后不良事件(包括死亡和并发症)仍是全球重大负担,许多可通过早期识别高危患者并实施针对性围术期干预预防。准确的风险分层至关重要。随着大规模电子健康记录(EHR)的可用性增加,机器学习(ML)为建模复杂临床模式提供了数据驱动方法。然而,现有研究在设计上差异巨大,方法学实践分散。本系统性述评对使用EHR数据进行手术风险分层与结局预测的端到端机器学习流程进行了刻画。我们审查了190项研究,涵盖数据预处理、算法选择、模型评估和可解释性等环节。大多数研究依赖单中心私有数据集,数据模态有限;公开手术数据集稀缺,制约了可复现性和泛化能力。关键预处理步骤(如缺失值处理、特征选择、类别不平衡)报告常不完整。传统机器学习模型和简单神经网络占主导,深度学习与多模态方法仍少见。基准数据集和标准化评估协议基本缺失,阻碍跨研究比较。仅有约三分之一的研究纳入可解释性方法。本综述揭示了限制临床稳健性术后机器学习工具的方法学缺口,并提供结构化参考,以支持更严谨、可复现且具临床意义的围术期机器学习开发。

原文摘要 · Abstract (English)

Postoperative adverse events, including mortality and morbidity, remain a major global burden, many of which are preventable through early identification of high-risk patients and targeted perioperative care. Accurate risk stratification is therefore essential. With the growing availability of large-scale electronic health records (EHRs), machine learning (ML) provides a data-driven approach to model complex clinical patterns. However, existing studies vary widely in design, and methodological practices remain fragmented. This scoping review characterizes ML pipelines for surgical risk stratification and outcome prediction using EHR data. We reviewed 190 studies covering the ML workflow, including data preprocessing, algorithm selection, model evaluation, and explainability. Most studies relied on single-center private datasets with limited data modalities, while the scarcity of open-access surgical datasets constrained reproducibility and generalizability. Reporting of key preprocessing steps, including missing data handling, feature selection, and class imbalance, was often incomplete. Conventional ML models and simple neural networks predominated, whereas deep learning and multimodal approaches remained uncommon. Benchmark datasets and standardized evaluation protocols were largely absent, hindering cross-study comparisons. Only about one-third of studies incorporated explainability methods. This review identifies methodological gaps limiting clinically robust postoperative ML tools and provides a structured reference to support more rigorous, reproducible, and clinically meaningful ML development for perioperative care.

风险预测机器学习医疗数据可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。