提出新方法提升解释模型对对抗攻击的鲁棒性
SHLIME: Foiling adversarial attacks fooling SHAP and LIME
- 设计模块化测试框架系统评估解释方法
- 发现某些集成配置可显著提升偏见检测能力
- 适用于高风险场景下模型透明度保障
事后解释方法如LIME和SHAP为黑箱分类器提供可解释洞察,日益用于评估模型偏见与泛化能力。然而,这些方法易受对抗操纵,可能掩盖有害偏见。基于Slack等(2020)的工作,我们研究了LIME和SHAP对有偏模型的脆弱性,并评估增强与集成解释策略的鲁棒性。我们首先复现原始COMPAS实验以验证先前结论并建立基线。随后引入模块化测试框架,实现对不同性能分类器上增强与集成解释方法的系统评估。利用该框架,我们在分布外模型上评估多种LIME/SHAP集成配置,对比其抵抗偏见隐藏的能力。结果识别出能显著提升偏见检测的配置,凸显其在高风险机器学习部署中增强透明性的潜力。
原文摘要 · Abstract (English)
Post hoc explanation methods, such as LIME and SHAP, provide interpretable insights into black-box classifiers and are increasingly used to assess model biases and generalizability. However, these methods are vulnerable to adversarial manipulation, potentially concealing harmful biases. Building on the work of Slack et al. (2020), we investigate the susceptibility of LIME and SHAP to biased models and evaluate strategies for improving robustness. We first replicate the original COMPAS experiment to validate prior findings and establish a baseline. We then introduce a modular testing framework enabling systematic evaluation of augmented and ensemble explanation approaches across classifiers of varying performance. Using this framework, we assess multiple LIME/SHAP ensemble configurations on out-of-distribution models, comparing their resistance to bias concealment against the original methods. Our results identify configurations that substantially improve bias detection, highlighting their potential for enhancing transparency in the deployment of high-stakes machine learning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。