arXiv:2501.05795cs.LGcs.AI2025-01被引 4

用多目标优化生成更稳健的反事实解释,提升模型决策可信度。

Robust Counterfactual Explanations under Model Multiplicity Using Multi-Objective Optimization

  • 引入帕累托改进视角,通过多目标优化生成抗模型多样性干扰的解释。
  • 在模拟与真实数据上验证,生成的解释在多个相似模型间保持稳定。
  • 适合关注AI可解释性、安全决策和自动化规划的研究者参考。

近年来,机器学习可解释性受到广泛关注。其中,基于示例的反事实解释(Counterfactual Explanation, CE)因其直观性备受瞩目。然而,当存在多个精度相近的机器学习模型时,传统CE方法的稳定性不足,影响其在安全决策中的应用。本文提出一种鲁棒的反事实解释方法,引入帕累托改进(Pareto improvement)的新视角,并采用多目标优化框架实现。通过在模拟数据和真实数据上的实验评估,结果表明该方法在多个相似模型下仍能生成一致且可靠的解释,兼具鲁棒性与实用性。研究强调了将社会福利概念应用于机器学习决策可解释性的潜力,为机器学习可解释性、决策支持及基于模型的行动规划等领域提供了重要基础。

原文摘要 · Abstract (English)

In recent years, explainability in machine learning has gained importance. In this context, counterfactual explanation (CE), which is an explanation method that uses examples, has attracted attention. However, it has been pointed out that CE is not robust when there are multiple machine-learning models with similar accuracy. These problems are important when using machine learning to make safe decisions. In this paper, we propose robust CEs that introduce a new viewpoint -- Pareto improvement -- and a method that uses multi-objective optimization to generate it. To evaluate the proposed method, we conducted experiments using both simulated and real data. The results demonstrate that the proposed method is both robust and practical. This study highlights the potential of ensuring robustness in decision-making by applying the concept of social welfare. We believe that this research can serve as a valuable foundation for various fields, including explainability in machine learning, decision-making, and action planning based on machine learning.

反事实解释可解释AI多目标优化鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。