arXiv:2501.02042cs.LGcs.CR2025-01被引 1

改进文本解释的稳定性评估,让模型更可靠地识别对抗攻击。

Towards Robust and Accurate Stability Estimation of Local Surrogate Models in Text-based Explainable AI

  • 引入语义相似性加权机制,提升解释稳定性评估精度
  • 发现多数相似度度量过于敏感,易误判模型脆弱性
  • 适合关注AI可解释性安全性的研究人员与开发者

近期研究关注自然语言处理中可解释AI(XAI)的对抗攻击,重点考察局部代理方法(如Lime)对输入微小扰动的脆弱性。此类攻击在不改变原始输入语义和结构的前提下,篡改生成的解释,尤其在医疗诊断或法律纠纷等高风险场景下令人担忧。尽管已揭示多种XAI方法存在缺陷,但其根源仍缺乏深入探讨。核心问题在于用于衡量解释差异的相似度度量选择不当,可能误导对XAI稳定性和对抗鲁棒性的判断。本文系统评估了文献中用于文本排名列表的多种相似度度量,发现多数方法过于敏感,导致稳定性误判。为此,提出一种融合特征间同义关系的加权方案,显著提升对对抗样本真实脆弱性的估计准确性。

原文摘要 · Abstract (English)

Recent work has investigated the concept of adversarial attacks on explainable AI (XAI) in the NLP domain with a focus on examining the vulnerability of local surrogate methods such as Lime to adversarial perturbations or small changes on the input of a machine learning (ML) model. In such attacks, the generated explanation is manipulated while the meaning and structure of the original input remain similar under the ML model. Such attacks are especially alarming when XAI is used as a basis for decision making (e.g., prescribing drugs based on AI medical predictors) or for legal action (e.g., legal dispute involving AI software). Although weaknesses across many XAI methods have been shown to exist, the reasons behind why remain little explored. Central to this XAI manipulation is the similarity measure used to calculate how one explanation differs from another. A poor choice of similarity measure can lead to erroneous conclusions about the stability or adversarial robustness of an XAI method. Therefore, this work investigates a variety of similarity measures designed for text-based ranked lists referenced in related work to determine their comparative suitability for use. We find that many measures are overly sensitive, resulting in erroneous estimates of stability. We then propose a weighting scheme for text-based data that incorporates the synonymity between the features within an explanation, providing more accurate estimates of the actual weakness of XAI methods to adversarial examples.

可解释AI对抗攻击稳定性评估文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。