提出新型多目标进化算法,生成更可信的反事实解释。
A Novel Multi-Objective Evolutionary Algorithm for Counterfactual Generation
- 用词典优化替代帕累托支配,改进反事实生成思路。
- 在15组实验中,生成的反事实有效性显著提升。
- 适合关注模型可解释性与公平性的研究人员。
机器学习模型在决策中广泛应用,但其黑箱特性使用户难以理解预测结果,尤其当结果为负面时(如贷款被拒)。反事实解释可帮助用户了解如何改变自身特征以获得正面结果,例如:若收入从3.5万英镑增至5万英镑,贷款可能获批。本文提出两项新贡献:(a)基于词典优化的新型多目标进化算法,取代主流的帕累托支配方法;(b)对反事实有效性定义进行扩展,通过测量其对单调性约束违反的鲁棒性来评估,例如收入上升应导致审批概率单调增加。在15个实验设置(3种黑盒模型 × 5个数据集)中,所提算法在性能上媲美现有帕累托方法,且有效性指标显著提升。
原文摘要 · Abstract (English)
Machine learning algorithms that learn black-box predictive models (which cannot be directly interpreted) are increasingly used to make predictions affecting the lives of people. It is important that users understand the predictions of such models, particularly when the model outputs a negative prediction for the user (e.g. denying a loan). Counterfactual explanations provide users with guidance on how to change some of their characteristics to receive a different, positive classification by a predictive model. For example, if a predictive model rejected a loan application from a user, a counterfactual explanation might state: If your salary was £50,000 (rather than your current £35,000), then your loan would be approved. This paper proposes two novel contributions: (a) a novel multi-objective Evolutionary Algorithm (EA) for counterfactual generation based on lexicographic optimisation, rather than the more popular Pareto dominance approach; and (b) an extension to the definition of the objective of validity for a counterfactual, based on measuring the resilience of a counterfactual to violations of monotonicity constraints which are intuitively expected by users; e.g., intuitively, the probability of a loan application to be approved would monotonically increase with an increase in the salary of the applicant. Experiments involving 15 experimental settings (3 types of black box models times 5 datasets) have shown that the proposed lexicographic optimisation-based EA is very competitive with an existing Pareto dominance-based EA; and the proposed extension of the validity objective has led to a substantial increase in the validity of the counterfactuals generated by the proposed EA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。