让机器解释更真实:用符号规则约束生成可行的反事实建议。
PACE: A Neuro-Symbolic Framework for Plausible and Actionable Counterfactual Explanations

- 用神经网络预测+符号逻辑约束,确保建议符合现实规则
- 在成人收入数据集上,满足领域可行性要求的解释比例提升显著
- 适合需要可操作、可信解释的医疗、金融等场景
反事实解释通过找出最小输入变动来说明模型决策。现有方法常生成不切实际的建议,因缺乏显式引入领域知识和干预约束的机制。本文提出PACE框架,将预测与推理分离:神经网络负责分类,符号推理层在生成反事实时强制执行领域特定约束。该框架显式建模可行干预,使解释既符合领域知识又可理解且可操作。方法具备模型无关性,适用于需现实决策支持的领域。以成人收入数据集为例,结合多层感知机与答案集编程(ASP)规则,编码教育、职业、工作时长的可行修改,同时保持不可变属性。结果揭示了反事实有效性与合理性之间的权衡,证明符号约束能显著提升解释的可行性,展示了神经符号方法在可解释人工智能中生成透明、可行反事实解释的潜力。
原文摘要 · Abstract (English)
Counterfactual explanations explain machine learning predictions by identifying minimal input changes that would alter a model's decision. Although many existing methods successfully generate prediction-changing alternatives, they often produce unrealistic or infeasible recommendations due to a lack of explicit mechanisms for incorporating domain knowledge and intervention constraints. Neuro-symbolic AI offers a promising direction by combining data-driven predictive models with symbolic reasoning capable of representing human-understandable rules and feasible actions. This paper presents PACE, a modular neuro-symbolic framework for generating feasibility-aware counterfactual explanations. The framework separates prediction and reasoning into two components: a neural predictive model for classification and a symbolic reasoning layer that enforces domain-specific constraints during counterfactual generation. By explicitly modeling feasible interventions, the framework produces explanations consistent with domain knowledge while remaining interpretable and actionable. The approach is model-agnostic and adaptable to domains requiring realistic decision support. A case study is conducted on the Adult Income dataset, combining a multilayer perceptron classifier with Answer Set Programming (ASP) rules encoding feasible modifications to education, occupation, and working hours while preserving immutable attributes. Results highlight the trade-off between counterfactual validity and plausibility and show that symbolic constraints yield explanations that better satisfy domain-specific feasibility requirements, illustrating the potential of neuro-symbolic methods for transparent, feasibility-aware counterfactual explanation in explainable AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。