arXiv:2409.17774cs.CLcs.AI2024-09中稿 · EMNLP被引 4

提出对抗敏感性评估方法,更真实反映NLP解释的可靠性。

Faithfulness and the Notion of Adversarial Sensitivity in NLP Explanations

  • 用对抗攻击下解释器的响应变化衡量可解释性可靠性。
  • 揭示现有评估方法在捕捉模型真实推理上的不足。
  • 适合关注解释可信度与模型鲁棒性的研究者。

可解释AI的可靠性评估中,忠实性(Faithfulness)是最关键的指标。在自然语言处理领域,当前的忠实性评估方法存在显著偏差与不一致,难以真实反映模型的推理过程。本文提出一种新方法——对抗敏感性(Adversarial Sensitivity),通过观察模型在对抗攻击下解释器的响应变化,来评估解释的忠实性。该方法从一个此前被忽视但至关重要的角度量化了忠实性,弥补了现有评估技术的重大缺陷,为提升NLP解释的可信度提供了新范式。

原文摘要 · Abstract (English)

Faithfulness is arguably the most critical metric to assess the reliability of explainable AI. In NLP, current methods for faithfulness evaluation are fraught with discrepancies and biases, often failing to capture the true reasoning of models. We introduce Adversarial Sensitivity as a novel approach to faithfulness evaluation, focusing on the explainer's response when the model is under adversarial attack. Our method accounts for the faithfulness of explainers by capturing sensitivity to adversarial input changes. This work addresses significant limitations in existing evaluation techniques, and furthermore, quantifies faithfulness from a crucial yet underexplored paradigm.

可解释AI忠实性对抗攻击NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。