检验深度模型在多组学数据中特征归因的一致性,发现结果受架构和初始化影响大。
Consistency of Feature Attribution in Deep Learning Architectures for Multi-Omics
- 用SHAP方法评估多视图深度模型的特征重要性
- 不同模型架构和随机初始化导致特征排序差异显著
- 提出简单方法验证关键生物分子识别的可靠性
过去十年中,机器学习与深度学习在生物研究中日益流行,但仍面临模型可解释性挑战。本文研究了在多组学数据上应用多视图深度学习模型时,使用Shapley加性解释(SHAP)方法识别关键生物分子的有效性。通过比较不同网络架构下特征的重要性排序,评估该方法的一致性。我们进行了多项计算实验,以检验SHAP的鲁棒性,并探索提升特征重要性识别可靠性的建模策略与诊断方法。以基于重要特征子集训练的随机森林模型准确率,以及仅使用这些特征的聚类质量作为有效性指标。结果表明,SHAP生成的特征排名对模型架构和权重初始化非常敏感,提示在多组学深度学习模型中使用归因方法需谨慎。本文还提出一种简便方法,用于评估关键生物分子识别的稳健性。
原文摘要 · Abstract (English)
Machine and deep learning have grown in popularity and use in biological research over the last decade but still present challenges in interpretability of the fitted model. The development and use of metrics to determine features driving predictions and increase model interpretability continues to be an open area of research. We investigate the use of Shapley Additive Explanations (SHAP) on a multi-view deep learning model applied to multi-omics data for the purposes of identifying biomolecules of interest. Rankings of features via these attribution methods are compared across various architectures to evaluate consistency of the method. We perform multiple computational experiments to assess the robustness of SHAP and investigate modeling approaches and diagnostics to increase and measure the reliability of the identification of important features. Accuracy of a random-forest model fit on subsets of features selected as being most influential as well as clustering quality using only these features are used as a measure of effectiveness of the attribution method. Our findings indicate that the rankings of features resulting from SHAP are sensitive to the choice of architecture as well as different random initializations of weights, suggesting caution when using attribution methods on multi-view deep learning models applied to multi-omics data. We present an alternative, simple method to assess the robustness of identification of important biomolecules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。