arXiv:2504.02911cs.CLcs.AI2025-04被引 3

通过可控噪声扰动,更真实地解释大模型决策依据。

Noiser: Bounded Input Perturbations for Attributing Large Language Models

  • 对输入嵌入施加有界噪声,评估模型鲁棒性以获取归因。
  • 在6个大模型、3项任务上,归因忠实度与可回答性均领先。
  • 适合需要可信解释的AI安全、医疗等高风险场景使用。

特征归因(FA)方法是解释大型语言模型(LLMs)预测行为的常用后处理手段。本文提出Noiser,一种基于扰动的归因方法,在每个输入嵌入上施加有界噪声,通过测量模型对部分噪声输入的鲁棒性来获得输入归因。同时,我们设计了一个可回答性指标,利用指令式判别模型评估高分标记是否足以恢复预测结果。在六个大模型和三个任务上的综合评估表明,Noiser在归因忠实度和可回答性方面均显著优于现有的梯度法、注意力法和扰动法,展现出更强的鲁棒性和有效性,为解释语言模型预测提供了可靠方案。

原文摘要 · Abstract (English)

Feature attribution (FA) methods are common post-hoc approaches that explain how Large Language Models (LLMs) make predictions. Accordingly, generating faithful attributions that reflect the actual inner behavior of the model is crucial. In this paper, we introduce Noiser, a perturbation-based FA method that imposes bounded noise on each input embedding and measures the robustness of the model against partially noised input to obtain the input attributions. Additionally, we propose an answerability metric that employs an instructed judge model to assess the extent to which highly scored tokens suffice to recover the predicted output. Through a comprehensive evaluation across six LLMs and three tasks, we demonstrate that Noiser consistently outperforms existing gradient-based, attention-based, and perturbation-based FA methods in terms of both faithfulness and answerability, making it a robust and effective approach for explaining language model predictions.

模型解释大模型归因方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。