发现安全AI解释性因特征相关性而失效,提出新方法稳定解释结果。
Stabilising Explainability Fragility in Cybersecurity AI: The Impact and Mitigation of Multicollinearity in Public Benchmark Datasets

- 证明特征共线性会放大解释方差,导致重要性不可识别。
- 在UNSW-NB15数据集上验证4类模型,解释不稳定度提升达60%以上。
- 提出脆弱性评分和两种新方法,适配安全领域可解释性需求。
本文研究了入侵检测系统(IDS)中人工智能可解释性所面临的一种未被关注却影响深远的漏洞:多重共线性引发的不稳定性。尽管广泛依赖SHAP、LIME等后验解释工具,其对相关特征的鲁棒性影响尚未评估。我们提出一个形式化定理,表明多重共线性会放大归因方差,导致解释与特征重要性在共线性下不可识别。在代表性基准数据集UNSW-NB15上,通过全面实验验证该定理,涵盖线性、树基、核函数及神经网络四类模型,基于VIF与相关阈值进行全集与剪枝特征集测试。提出新型可解释性脆弱性评分指标,并设计两种缓解方法:CAA-Filtering通过整合已训练模型的归因以稳定解释;SHARP是一种新颖的训练时正则化框架,惩罚归因不稳定性,实现可控且单调的可解释性稳定性提升。通过Kendall's τ量化自助采样解释间的不稳定性,结果支持预测性能的稳定性。本工作对安全关键场景中XAI的可信性与可复现性具有直接意义,推动将多重共线性缓解措施纳入IDS流水线,为从业者提供实用指南。
原文摘要 · Abstract (English)
This paper investigates a unexplored yet impactful vulnerability in AI explainability used in intrusion detection (IDS): multicollinearity-induced instability. Despite extensive reliance on post-hoc explainability tools such as SHAP or LIME, the impact of correlated features on explanation robustness is not evaluated. We introduce a formal theorem stating that multicollinearity inflates attribution variance. This demonstrates that explanations and feature importances are non-identifiable under multicollinearity. A suite of comprehensive experiments validates the theorem on a representative benchmark dataset, UNSW-NB15. Four widely used families of models are evaluated, including linear, tree-based, kernel, and neural, across full and pruned feature sets based on VIF and correlation thresholding. We propose the novel metric of Explanability Fragility Score and two novel methods to mitigate it with variable integration complexity. CAA-Filtering focuses on stabilising explanations by grouping attributions of trained models. SHARP is a novel training-time regularisation framework that penalises attribution instability, enabling controllable and monotonic improvement of explainability stability. The findings support stable predictive performance, using Kendall's τ to quantify instability across bootstrapped explanations. This work has direct implications for the trustworthiness and reproducibility of XAI in security-critical contexts, and motivates incorporating multicollinearity mitigations into the IDS pipelines, providing a set of guidelines for practitioners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。