用图模型和优化方法,让随机森林的特征交互更可解释。
Surrogate Interpretable Graph for Random Decision Forests
- 构建代理可解释图,通过图结构和整数规划分析特征交互。
- 可视化每个决策路径上的特征使用情况与主导交互关系。
- 适合医疗等高风险领域,提升模型可信度与合规性。
健康信息学领域受随机森林模型的深刻影响,这类模型在特征交互可解释性方面取得了显著进展。其对过拟合的鲁棒性和并行化能力使其在该领域尤为适用。然而,随着特征和估计器数量增加,领域专家难以准确理解全局特征交互,进而影响信任度与监管合规性。为此,提出一种称为代理可解释图的方法,利用图结构与混合整数线性规划分析并可视化特征交互。该方法通过决策-特征-交互表展示特征使用情况,揭示预测中主导的层次化特征交互关系。代理可解释图的实现显著提升了全局可解释性,这对高风险领域至关重要。
原文摘要 · Abstract (English)
The field of health informatics has been profoundly influenced by the development of random forest models, which have led to significant advances in the interpretability of feature interactions. These models are characterized by their robustness to overfitting and parallelization, making them particularly useful in this domain. However, the increasing number of features and estimators in random forests can prevent domain experts from accurately interpreting global feature interactions, thereby compromising trust and regulatory compliance. A method called the surrogate interpretability graph has been developed to address this issue. It uses graphs and mixed-integer linear programming to analyze and visualize feature interactions. This improves their interpretability by visualizing the feature usage per decision-feature-interaction table and the most dominant hierarchical decision feature interactions for predictions. The implementation of a surrogate interpretable graph enhances global interpretability, which is critical for such a high-stakes domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。