用多准则决策方法融合多种解释指标,提升AI解释的可靠性。
Multi-criteria Rank-based Aggregation for Explainable AI
- 基于排名的加权聚合,同时平衡多个解释质量指标
- 在多个数据集上验证了方法在复杂度、忠实度和稳定性上的鲁棒性
- 适合需要可解释性且关注多维度评估的研究者与工程师
可解释性对提升黑箱机器学习模型的透明度至关重要。随着LIME、SHAP等解释方法的发展,各类XAI性能指标被提出用于评估解释质量。然而,不同解释器对同一预测可能给出相互矛盾的解释,导致在多个质量指标间存在权衡。尽管现有聚合方法提升了鲁棒性、降低了解释差异,但极少研究采用多准则决策方法。为此,本文提出一种基于排名的多准则加权聚合方法,能同时平衡多个质量指标,生成解释模型集成。此外,我们提出了复杂度、忠实度和稳定性三个指标的排名版本,以更准确评估排序型特征重要性解释。在公开数据集上的大量实验表明,该方法在各项指标上均表现出良好鲁棒性。对比多种多准则决策与排名聚合算法,TOPSIS与WSUM表现最佳。
原文摘要 · Abstract (English)
Explainability is crucial for improving the transparency of black-box machine learning models. With the advancement of explanation methods such as LIME and SHAP, various XAI performance metrics have been developed to evaluate the quality of explanations. However, different explainers can provide contrasting explanations for the same prediction, introducing trade-offs across conflicting quality metrics. Although available aggregation approaches improve robustness, reducing explanations' variability, very limited research employed a multi-criteria decision-making approach. To address this gap, this paper introduces a multi-criteria rank-based weighted aggregation method that balances multiple quality metrics simultaneously to produce an ensemble of explanation models. Furthermore, we propose rank-based versions of existing XAI metrics (complexity, faithfulness and stability) to better evaluate ranked feature importance explanations. Extensive experiments on publicly available datasets demonstrate the robustness of the proposed model across these metrics. Comparative analyses of various multi-criteria decision-making and rank aggregation algorithms showed that TOPSIS and WSUM are the best candidates for this use case.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。