用单一模型+多次随机试验,让个体化医疗模型更稳定可解释。
Stabilizing Machine Learning for Reproducible and Explainable Results: A Novel Validation Approach to Subject-Specific Insights
- 用同一随机森林模型反复试验(最多400次),通过种子变化捕捉个体特征
- 在9个数据集上验证,显著提升个体与群体层面特征重要性的一致性
- 适合关注临床可解释性、资源有限的医疗机器学习研究者
机器学习正推动医学研究发展,但通用模型难以应对人类生物多样性。为此,我们提出一种新型验证方法:仅用一个随机森林模型,在9个不同领域、样本量和人口特征的数据集上,通过每名受试者最多400次随机种子试验,生成400组特征集,识别出个体特异性关键特征,并汇总得出群体特征重要性。该方法相比传统方式,在保持高准确率的同时,显著提升了特征重要性的一致性。本方法以低成本实现稳定、可解释的个体化预测,为临床研究提供实用替代方案。
原文摘要 · Abstract (English)
Machine Learning is transforming medical research by improving diagnostic accuracy and personalizing treatments. General ML models trained on large datasets identify broad patterns across populations, but their effectiveness is often limited by the diversity of human biology. This has led to interest in subject-specific models that use individual data for more precise predictions. However, these models are costly and challenging to develop. To address this, we propose a novel validation approach that uses a general ML model to ensure reproducible performance and robust feature importance analysis at both group and subject-specific levels. We tested a single Random Forest (RF) model on nine datasets varying in domain, sample size, and demographics. Different validation techniques were applied to evaluate accuracy and feature importance consistency. To introduce variability, we performed up to 400 trials per subject, randomly seeding the ML algorithm for each trial. This generated 400 feature sets per subject, from which we identified top subject-specific features. A group-specific feature importance set was then derived from all subject-specific results. We compared our approach to conventional validation methods in terms of performance and feature importance consistency. Our repeated trials approach, with random seed variation, consistently identified key features at the subject level and improved group-level feature importance analysis using a single general model. Subject-specific models address biological variability but are resource-intensive. Our novel validation technique provides consistent feature importance and improved accuracy within a general ML model, offering a practical and explainable alternative for clinical research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。