arXiv:2512.21602cs.LGcs.CV2025-12被引 1

比较六类模型在急诊重症数据中的抗偏倚与扩展性,发现基础模型表现逼近传统方法且更省资源。

An Empirical Study of Machine Learning Robustness and Scalability for Imbalanced Tabular Clinical Data in Emergency and Critical Care

  • 对比决策树、随机森林、XGBoost等六类模型在不平衡临床数据上的表现
  • TabPFN和TabICL在MIMIC-IV-ED上平均宏F1得分最高,XGBoost在eICU持续领先
  • 基础模型计算成本低,适合资源有限的临床场景

每年数百万患者经过急诊科和重症监护室,医生需在时间紧迫与不确定性中做出高风险决策。机器学习可辅助预测恶化、分诊及罕见危重结果,但临床数据常严重失衡,导致模型偏向多数类,降低预测性能。本研究在MIMIC-IV-ED与eICU数据库的不平衡表格式临床数据上评估了六类模型:决策树、随机森林、XGBoost、TabNet、TabICL和TabPFN v2.6。可训练模型采用贝叶斯超参数调优,基础模型以预训练推理模式评估,不进行任务特定重加权。通过宏F1分数、对不平衡加剧的鲁棒性以及跨七项临床预测任务的计算可扩展性进行评估。结果因数据集而异:在MIMIC-IV-ED上,TabPFN v2.6和TabICL平均宏F1排名最强,XGBoost仍具竞争力;在eICU上,XGBoost始终最优,其次为其他树基方法,基础模型表现居中。跨两数据集,TabNet在不平衡加剧时性能下降最显著,且计算开销最高。训练时间分析显示,树基方法随数据规模增长扩展性最好,基础模型则具备低每任务适配成本。结果表明,无单一模型家族在所有临床环境中占优。但表格式基础模型正缩小与强经典基线的性能差距,同时提供独特的效率-性能权衡,可能惠及资源受限的临床环境。

原文摘要 · Abstract (English)

Every year, millions of patients pass through emergency departments and intensive care units, where clinicians must make high-stakes decisions under time pressure and uncertainty. Machine learning could support prediction of deterioration, triage, and rare critical outcomes, but clinical data are often severely imbalanced, biasing models toward majority classes and reducing predictive performance. Developing robust and efficient models for imbalanced clinical tabular data therefore remains an important challenge. We evaluated six model families on imbalanced tabular data from the MIMIC-IV-ED and eICU databases: Decision Tree, Random Forest, XGBoost, TabNet, TabICL, and TabPFN v2.6. Trainable models were optimized using Bayesian hyperparameter tuning, while foundation models were evaluated in their pretrained inference regime without task-specific reweighting. Models were assessed using Macro F1-score, robustness to increasing imbalance, and computational scalability across seven clinical prediction tasks. Results differed across datasets. On MIMIC-IV-ED, TabPFN v2.6 and TabICL achieved the strongest average Macro F1 ranks, with XGBoost remaining competitive. On eICU, XGBoost consistently performed best, followed by other tree-based methods, while foundation models achieved intermediate performance. Across both datasets, TabNet showed the largest degradation under increasing imbalance and the highest computational cost. Training-time analysis showed that tree-based methods scaled most favorably with dataset size, while foundation models offered low per-task adaptation cost. These findings suggest that no single model family dominates across all clinical settings. However, tabular foundation models are narrowing the performance gap with strong classical baselines while offering a distinct efficiency-performance trade-off that may benefit resource-constrained clinical environments.

临床预测不平衡数据模型效率表格式模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。