用同义词相似度改进XAI评估,更真实反映模型抗干扰能力
Quantifying True Robustness: Synonymity-Weighted Similarity for Trustworthy XAI Evaluation
- 引入同义词权重修正评估指标,区分语义相近与无关的词替换
- 实验显示传统方法高估攻击成功率,新方法更准确衡量系统脆弱性
- 适合关注XAI可信性、对抗鲁棒性的研究者与开发者
对抗攻击通过改变解释内容但保持模型输出不变来挑战可解释AI(XAI)的可靠性。当前对文本型XAI攻击效果的评估多依赖标准信息检索指标,但这些指标将所有词语扰动视为等价,忽视了同义性,从而可能错误估计攻击的真实影响。为此,本文提出同义词权重法,通过融合被扰动词语的语义相似度来修正评估指标,使脆弱性分析更精准,有效防止对攻击成功率的高估,为评估AI系统的真正鲁棒性提供可靠工具。
原文摘要 · Abstract (English)
Adversarial attacks challenge the reliability of Explainable AI (XAI) by altering explanations while the model's output remains unchanged. The success of these attacks on text-based XAI is often judged using standard information retrieval metrics. We argue these measures are poorly suited in the evaluation of trustworthiness, as they treat all word perturbations equally while ignoring synonymity, which can misrepresent an attack's true impact. To address this, we apply synonymity weighting, a method that amends these measures by incorporating the semantic similarity of perturbed words. This produces more accurate vulnerability assessments and provides an important tool for assessing the robustness of AI systems. Our approach prevents the overestimation of attack success, leading to a more faithful understanding of an XAI system's true resilience against adversarial manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。