通过局部扰动与自解释比对,评估大模型回答的可信度。
Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs
- 用局部扰动找出模型正确回答所需的必要信息片段
- 在Natural Questions数据集上验证了方法有效性
- 适合关注大模型决策可信性的研究者使用
本文提出一种新任务,通过局部扰动和自解释来评估大语言模型(LLMs)的可信性。许多LLM在回答某些问题时需要额外上下文才能正确作答。为此,我们提出一种受常用留一法启发的高效可解释性技术:通过该方法识别出模型生成正确答案所必需且充分的信息部分,作为解释内容。我们还提出了一个评估可信性的指标,将这些关键部分与模型自身的解释进行比对。在Natural Questions数据集上验证了该方法的有效性,证明其能有效解释模型决策并评估其可信程度。
原文摘要 · Abstract (English)
This paper introduces a novel task to assess the faithfulness of large language models (LLMs) using local perturbations and self-explanations. Many LLMs often require additional context to answer certain questions correctly. For this purpose, we propose a new efficient alternative explainability technique, inspired by the commonly used leave-one-out approach. Using this approach, we identify the sufficient and necessary parts for the LLM to generate correct answers, serving as explanations. We propose a metric for assessing faithfulness that compares these crucial parts with the self-explanations of the model. Using the Natural Questions dataset, we validate our approach, demonstrating its effectiveness in explaining model decisions and assessing faithfulness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。