arXiv:2505.00571stat.MLcs.LG2025-05被引 1

融合贝叶斯与树模型,实现流行病数据中非线性效应的可解释推断

Discovery and inference beyond linearity for epidemiological data by integrating Bayesian regression, tree ensembles and Shapley values

  • 用贝叶斯稀疏回归+规则生成器构建可解释框架
  • 首次在个体层面提供特征影响的不确定性量化
  • 适合需要可解释推断的流行病学研究者使用

机器学习在流行病学中用于无假设发现风险与保护因素日益流行。尽管其擅长捕捉非线性关系和交互作用,但缺乏可靠的统计推断能力。虽然Shapley值能提供局部特征贡献度,但通常缺乏有效的不确定性量化,阻碍了统计推断。本文提出RuleSHAP框架,结合专用贝叶斯稀疏回归模型、改进的基于树的规则生成器与Shapley值归因,实现非线性与交互效应的检测,并在个体层面提供不确定性量化。我们推导出该框架内边际Shapley值的高效计算公式。将RuleSHAP应用于流行病队列数据,成功识别高胆固醇与血压相关的非线性交互效应,涉及年龄、性别、种族、体重指数和血糖水平等特征。最后,在模拟数据上验证了框架的有效性。

原文摘要 · Abstract (English)

Machine Learning (ML) is gaining popularity in epidemiology and healthcare studies for hypothesis-free discovery of risk and protective factors. ML is strong at discovering nonlinearities and interactions, but this power is compromised by a lack of reliable inference. Although Shapley values provide local measures of features' effects, valid uncertainty quantification for these effects is typically lacking, thus precluding statistical inference. We propose RuleSHAP, a framework that addresses this limitation by combining a dedicated Bayesian sparse regression model with an improved tree-based rule generator and Shapley value attribution. RuleSHAP provides detection of nonlinear and interaction effects, with uncertainty quantification at the individual level as a key contribution. We derive an efficient formula for computing marginal Shapley values within this framework. We apply RuleSHAP to data from an epidemiological cohort to detect and infer several effects for high cholesterol and blood pressure, such as nonlinear interaction effects between features like age, sex, ethnicity, BMI and glucose level. To conclude, we demonstrate the validity of our framework on simulated data.

可解释性贝叶斯方法流行病学特征归因

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。