arXiv:2601.03089cs.CLcs.AI2026-01

提出新评估框架,让大模型解释结果更公平可比。

Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information

  • 用固定保留概率控制干扰词数,避免评分被误导。
  • 新方法Grad-ELLM在分类任务中解释更全面且稳定。
  • 适合关注模型可解释性研究的学者与工程师。

大型语言模型(LLM)越来越多地使用输入归因方法进行评估,但比较这些解释仍具挑战性。现有软扰动类可信度指标(如Soft-NC和Soft-NS)会将归因质量与扰动过程中保留的词数混淆:归因得分较高的方法可能因保留更多词而获得虚高评分。为此,我们提出π-Soft-NC和π-Soft-NS,一种在相同期望保留概率下比较归因方法的评估框架,从而控制保留词数。我们还引入Grad-ELLM,一种专为自回归解码器模型设计的基于梯度的归因方法,该方法在每一步解码时结合梯度导出的通道重要性与注意力导出的标记重要性。在Llama和Mistral上的分类与开放生成任务实验表明,Grad-ELLM在π-Soft-NC下表现出强全面性导向的可信度,而在π-Soft-NS下则无主导方法。该评估框架为LLM可解释性方法提供了严谨的比较基础,有助于推动该领域进展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly evaluated with input attribution methods, yet comparing such explanations remains challenging. Existing soft-perturbation faithfulness metrics, such as Soft-NC and Soft-NS, can conflate attribution quality with the number of words retained during perturbation: attribution methods with larger average scores may keep more words and therefore obtain inflated scores. To address this issue, we propose $π$-Soft-NC and $π$-Soft-NS, an evaluation framework that compares attribution methods under the same expected retaining probability, thus controlling the number of retained words. We further introduce Grad-ELLM, a gradient-based attribution method tailored to autoregressive decoder-only LLMs, which combines gradient-derived channel importance with attention-derived token importance at each decoding step. Experiments on classification and open-generation tasks with Llama and Mistral show that Grad-ELLM achieves strong comprehensiveness-oriented faithfulness under $π$-Soft-NC, while there is no dominant method under $π$-Soft-NS. Our evaluation metric serves as a rigorous framework to compare XAI methods for LLMs, which will support progress in the field.

可解释性大模型归因评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。