提出新方法AXE,精准识别模型解释中的虚假误导。
Evaluating the Ability of Explanations to Disambiguate Models in a Rashomon Set
- 基于三个评估原则设计无真值依赖的解释评价方法
- 在对抗性伪装场景中实现100%的虚假解释检测率
- 可识别模型是否利用敏感属性进行预测,适合可信AI部署
可解释人工智能(XAI)旨在揭示模型内部机制。对于性能相近的模型集合(Rashomon集),解释有助于区分各模型行为,辅助部署选择。然而解释结果受解释器影响,需有效评估。本文提出解释评估的三项原则及新方法AXE,用于评估特征重要性解释的质量。研究表明,依赖理想真值对比的评估指标会掩盖Rashomon集中模型间的实际行为差异。而遵循新原则的评估能凸显这些差异,助力模型筛选。值得注意的是,从Rashomon集中选取的替代模型虽保持相同预测结果,却可能诱导解释器生成虚假解释,并使传统评估方法误判其质量为高。相比之下,本研究提出的AXE方法可在该对抗性场景下实现100%的虚假解释检测成功率。与基于模型敏感性或真值比对的方法不同,AXE还能判断模型是否使用了敏感属性进行预测。
原文摘要 · Abstract (English)
Explainable artificial intelligence (XAI) is concerned with producing explanations indicating the inner workings of models. For a Rashomon set of similarly performing models, explanations provide a way of disambiguating the behavior of individual models, helping select models for deployment. However explanations themselves can vary depending on the explainer used, and need to be evaluated. In the paper "Evaluating Model Explanations without Ground Truth", we proposed three principles of explanation evaluation and a new method "AXE" to evaluate the quality of feature-importance explanations. We go on to illustrate how evaluation metrics that rely on comparing model explanations against ideal ground truth explanations obscure behavioral differences within a Rashomon set. Explanation evaluation aligned with our proposed principles would highlight these differences instead, helping select models from the Rashomon set. The selection of alternate models from the Rashomon set can maintain identical predictions but mislead explainers into generating false explanations, and mislead evaluation methods into considering the false explanations to be of high quality. AXE, our proposed explanation evaluation method, can detect this adversarial fairwashing of explanations with a 100% success rate. Unlike prior explanation evaluation strategies such as those based on model sensitivity or ground truth comparison, AXE can determine when protected attributes are used to make predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。