arXiv:2511.08994stat.APcs.LG2025-11

用通用术前数据预测手术时长,跨医院跨时间都准。

Generalisable prediction model of surgical case duration: multicentre development and temporal validation

  • 用四种机器学习模型融合预测,仅依赖常见术前信息。
  • 2024年外部验证中误差小,校准度良好(斜率0.921)。
  • 适合想提升手术室排班效率的医院管理者使用。

准确预测手术时长对手术室调度至关重要,但现有模型多依赖特定医院或医生输入,且极少进行外部验证,限制了通用性。本研究基于日本两家综合医院的回顾性多中心数据(开发集:2021年1月1日至2023年12月31日;时间外验证集:2024年1月1日至12月31日),纳入择期工作日手术且ASA分级为1-4的病例。预设术前预测变量包括手术背景(年、月、星期、预计时长、全身麻醉标志、体位)和患者因素(性别、年龄、体重指数、过敏史、感染史、合并症、ASA)。缺失数据通过链式方程多重插补处理。四种学习器(弹性网、广义加性模型、随机森林、梯度提升树)在内部-外部交叉验证(留一聚类法,按中心-年分组)中调优,并通过堆叠泛化组合预测对数转换后的手术时长。共分析63,206例手术(开发集45,647例,时间外验证集17,559例)。内部-外部交叉验证显示各中心与各年份性能稳定。2024年时间外验证中,模型校准良好(截距0.423,95%CI 0.372–0.474;斜率0.921,95%CI 0.911–0.932)。结论:仅使用广泛可得的术前变量,堆叠机器学习模型即可实现高精度、良好校准的预测,在时间上具有强外推能力,支持跨机构与跨时间应用。此类通用工具可无需依赖特殊输入,提升手术室调度效率。

原文摘要 · Abstract (English)

Background: Accurate prediction of surgical case duration underpins operating room (OR) scheduling, yet existing models often depend on site- or surgeon-specific inputs and rarely undergo external validation, limiting generalisability. Methods: We undertook a retrospective multicentre study using routinely collected perioperative data from two general hospitals in Japan (development: 1 January 2021-31 December 2023; temporal test: 1 January-31 December 2024). Elective weekday procedures with American Society of Anesthesiologists (ASA) Physical Status 1-4 were included. Pre-specified preoperative predictors comprised surgical context (year, month, weekday, scheduled duration, general anaesthesia indicator, body position) and patient factors (sex, age, body mass index, allergy, infection, comorbidity, ASA). Missing data were addressed by multiple imputation by chained equations. Four learners (elastic-net, generalised additive models, random forest, gradient-boosted trees) were tuned within internal-external cross-validation (IECV; leave-one-cluster-out by centre-year) and combined by stacked generalisation to predict log-transformed duration. Results: We analysed 63,206 procedures (development 45,647; temporal test 17,559). Cluster-specific and pooled errors and calibrations from IECV are provided with consistent performance across centres and years. In the 2024 temporal test cohort, calibration was good (intercept 0.423, 95%CI 0.372 to 0.474; slope 0.921, 95%CI 0.911 to 0.932). Conclusions: A stacked machine-learning model using only widely available preoperative variables achieved accurate, well-calibrated predictions in temporal external validation, supporting transportability across sites and over time. Such general-purpose tools may improve OR scheduling without relying on idiosyncratic inputs.

手术预测机器学习排班优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。