17种因果机器学习方法在两大临床试验中均未验证其个性化治疗效果的可靠性。
Causal Machine Learning Methods for Estimating Personalised Treatment Effects -- Insights on validity from two large trials
- 用两大随机对照试验数据评估17种因果机器学习方法的泛化能力
- 训练集与测试集间预测效果差异显著,无分布偏移也未能保持一致性
- 提示当前因果模型在精准医疗中应用需更严格验证
因果机器学习方法有望推动精准医疗发展,通过估计个性化治疗效果。然而,其在真实场景中的可靠性尚未充分验证。本研究基于两项大型随机对照试验——国际卒中试验(N=19,435)和中国急性卒中试验(N=21,106)的数据,评估了17种主流因果异质性机器学习方法(包括metalearners、树模型和深度学习方法)的内部与外部有效性。结果显示,所有方法均未在训练与测试数据间保持一致表现,即使在无分布偏移的情况下,训练所得个体化治疗效应也无法推广至测试数据。这表明当前因果机器学习模型在精准医疗中的实际应用仍存疑,亟需更稳健的验证机制以保障泛化性能。
原文摘要 · Abstract (English)
Causal machine learning (ML) methods hold great promise for advancing precision medicine by estimating personalized treatment effects. However, their reliability remains largely unvalidated in empirical settings. In this study, we assessed the internal and external validity of 17 mainstream causal heterogeneity ML methods -- including metalearners, tree-based methods, and deep learning methods -- using data from two large randomized controlled trials: the International Stroke Trial (N=19,435) and the Chinese Acute Stroke Trial (N=21,106). Our findings reveal that none of the ML methods reliably validated their performance, neither internal nor external, showing significant discrepancies between training and test data on the proposed evaluation metrics. The individualized treatment effects estimated from training data failed to generalize to the test data, even in the absence of distribution shifts. These results raise concerns about the current applicability of causal ML models in precision medicine, and highlight the need for more robust validation techniques to ensure generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。