arXiv:2602.11360cs.LGcs.AI2026-02

用自举法训练深度模型,让临床预测更稳定可靠。

Bootstrapping-based Regularisation for Reducing Individual Prediction Instability in Clinical Risk Prediction Models

  • 将自举采样过程嵌入训练,约束不同数据子集上的预测差异。
  • 在三个临床数据集上,预测偏差平均降低60%以上,稳定性显著提升。
  • 保持模型可解释性,适合医疗场景中对可靠性要求高的应用。

临床预测模型广泛用于支持患者诊疗,但许多基于深度学习的方法存在预测不稳定性问题——同一人群的不同采样训练结果差异大,影响可靠性与临床采纳。本文提出一种基于自举法的正则化框架,将自举过程直接嵌入深度神经网络训练中,使模型在不同重采样数据上保持一致预测,形成具有内在稳定性的单一模型。我们在模拟数据及三个临床数据集(GUSTO-I、Framingham、SUPPORT)上评估该方法,结果表明:相比传统模型和集成模型,新方法在所有数据集上均实现更高预测稳定性,如在GUSTO-I中平均绝对差从0.059降至0.019,在Framingham中从0.088降至0.057;显著减少异常偏离预测。同时,模型判别性能和特征重要性一致性得以保持,各模型间SHAP值相关性高(如GUSTO-I为0.894,Framingham为0.965)。尽管集成模型稳定性更高,但牺牲了可解释性——各子模型使用预测因子方式不一。本方法通过约束预测与自举分布对齐,使模型兼具鲁棒性与可解释性,为数据受限的医疗环境提供更可信的深度学习路径。

原文摘要 · Abstract (English)

Clinical prediction models are increasingly used to support patient care, yet many deep learning-based approaches remain unstable, as their predictions can vary substantially when trained on different samples from the same population. Such instability undermines reliability and limits clinical adoption. In this study, we propose a novel bootstrapping-based regularisation framework that embeds the bootstrapping process directly into the training of deep neural networks. This approach constrains prediction variability across resampled datasets, producing a single model with inherent stability properties. We evaluated models constructed using the proposed regularisation approach against conventional and ensemble models using simulated data and three clinical datasets: GUSTO-I, Framingham, and SUPPORT. Across all datasets, our model exhibited improved prediction stability, with lower mean absolute differences (e.g., 0.019 vs. 0.059 in GUSTO-I; 0.057 vs. 0.088 in Framingham) and markedly fewer significantly deviating predictions. Importantly, discriminative performance and feature importance consistency were maintained, with high SHAP correlations between models (e.g., 0.894 for GUSTO-I; 0.965 for Framingham). While ensemble models achieved greater stability, we show that this came at the expense of interpretability, as each constituent model used predictors in different ways. By regularising predictions to align with bootstrapped distributions, our approach allows prediction models to be developed that achieve greater robustness and reproducibility without sacrificing interpretability. This method provides a practical route toward more reliable and clinically trustworthy deep learning models, particularly valuable in data-limited healthcare settings.

临床预测模型稳定自举法可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。