arXiv:2504.19027cs.AIcs.LG2025-04被引 7

提升机器学习反事实解释的稳定性与可靠性,适配医疗金融等高风险场景。

DiCE-Extended: A Robust Approach to Counterfactual Explanations in Machine Learning

  • 引入多目标优化与新鲁棒性度量,平衡接近性、多样性与抗扰动能力。
  • 在多个数据集和模型上验证,生成解释更稳定且贴近决策边界。
  • 适合需要可信赖解释的医疗、金融等关键领域应用。

可解释人工智能(XAI)在医疗、金融、法律等决策关键领域日益重要。反事实(CF)解释通过建议最小化输入特征修改以改变模型输出,为用户提供可操作洞察。尽管进展显著,现有方法常难以兼顾接近性、多样性和鲁棒性,限制了实际应用。主流框架DiCE强调多样性但缺乏鲁棒性,导致解释对微小扰动敏感。为此,我们提出DiCE-Extended,融合多目标优化技术,在保持可解释性的前提下增强鲁棒性。引入基于Dice-Sørensen系数的新鲁棒性度量,提升小输入变化下的稳定性;通过加权损失项(lambda_p, lambda_d, lambda_r)协调接近性、多样性与鲁棒性。在COMPAS、Lending Club、German Credit、Adult Income等基准数据集上,跨Scikit-learn、PyTorch、TensorFlow多种后端验证,结果表明其生成的反事实解释具有更高有效性、稳定性和决策边界一致性。研究凸显了该框架在高风险场景中生成更可靠解释的潜力。未来可探索自适应优化与领域约束以进一步提升实用性。

原文摘要 · Abstract (English)

Explainable artificial intelligence (XAI) has become increasingly important in decision-critical domains such as healthcare, finance, and law. Counterfactual (CF) explanations, a key approach in XAI, provide users with actionable insights by suggesting minimal modifications to input features that lead to different model outcomes. Despite significant advancements, existing CF generation methods often struggle to balance proximity, diversity, and robustness, limiting their real-world applicability. A widely adopted framework, Diverse Counterfactual Explanations (DiCE), emphasizes diversity but lacks robustness, making CF explanations sensitive to perturbations and domain constraints. To address these challenges, we introduce DiCE-Extended, an enhanced CF explanation framework that integrates multi-objective optimization techniques to improve robustness while maintaining interpretability. Our approach introduces a novel robustness metric based on the Dice-Sørensen coefficient, enabling stability under small input variations. Additionally, we refine CF generation using weighted loss components (lambda_p, lambda_d, lambda_r) to balance proximity, diversity, and robustness. We empirically validate DiCE-Extended on benchmark datasets (COMPAS, Lending Club, German Credit, Adult Income) across multiple ML backends (Scikit-learn, PyTorch, TensorFlow). Results demonstrate improved CF validity, stability, and alignment with decision boundaries compared to standard DiCE-generated explanations. Our findings highlight the potential of DiCE-Extended in generating more reliable and interpretable CFs for high-stakes applications. Future work could explore adaptive optimization techniques and domain-specific constraints to further enhance CF generation in real-world scenarios

反事实解释可解释AI鲁棒性多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。