用机器学习提前预测肠杆菌科耐药风险,帮医生少开碳青霉烯类抗生素。
Machine Learning for Pre-Culture ESBL Risk Stratification to Guide Empiric Antibiotic Selection: A 12-Hospital Study of Enterobacteriaceae Cultures

- 基于电子病历45个变量训练XGBoost模型,提前预测耐药表型。
- 90%敏感性下阴性预测值达95.8%,每1000例可省307次不必要的广谱用药。
- 优先感染史是关键因素,模型可公平部署且不受社会经济因素影响。
针对疑似产超广谱β-内酰胺酶(ESBL)肠杆菌科感染的经验性治疗,需在48-72小时无培养结果时做出决策,临床常面临低估耐药风险或过度使用碳青霉烯类导致耐药加剧的两难。本研究构建了成本敏感的XGBoost模型,在12家医院132,955份肠杆菌科培养样本(来自72,217名患者,其中14.41%为ESBL阳性)中,利用45个入院前电子病历特征预测培养时的ESBL表型。按患者分层划分数据。在90%敏感性条件下,模型达到95.8%阴性预测值(NPV),使培养后ESBL概率降至4.2%,该阈值可能支持非重症监护病房安全地避免碳青霉烯类使用;每1000例中可减少307例不必要的广谱抗生素暴露,代价是漏诊14例ESBL病例。SHAP分析显示既往感染史是主导预测因子,优于既往菌负荷和社区贫困程度;移除贫困相关特征仅导致AUROC下降0.020,不影响模型公平性部署。在严格遵循IDSA ESBL-E定义下,模型表现稳定(AUROC 0.766),加入标本类型作为变量后为0.764,未做类别不平衡修正时仍达0.762,各菌种亚组间性能在0.71至0.78之间波动。
原文摘要 · Abstract (English)
Empiric antibiotic therapy for suspected ESBL-producing Enterobacteriaceae must be selected 48-72 hours before culture results, forcing clinicians to choose between undertreating resistant infections and overusing carbapenems that drive further resistance. We developed a cost-sensitive XGBoost model predicting an ESBL phenotype (resistance to ceftriaxone, ceftazidime, cefepime or piperacillin-tazobactam) at culture ordering using 45 pre-culture EHR features across 132,955 cultures from 72,217 patients at 12 hospitals (14.41% with the ESBL phenotype). Cultures were partitioned at the patient level. At 90% sensitivity, the model achieved 95.8% NPV, reducing post-test ESBL probability to 4.2%, a threshold that may support safe carbapenem-sparing in non-ICU settings, while sparing 307 of every 1,000 cultures an unnecessary broad-spectrum course at the cost of 14 missed ESBL cases per 1,000. SHAP analysis identified prior ESBL colonization as the dominant predictor, ahead of prior organism burden and neighborhood deprivation; removing deprivation features caused minimal performance loss ($\Delta\text{AUROC} = -0.020$), enabling equitable bedside deployment. Discrimination was unchanged under a strict IDSA ESBL-E definition (AUROC 0.766), with specimen type added as a predictor (0.764) and without any class-imbalance correction (0.762), and ranged from 0.71 to 0.78 across organism strata.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。