提出量化大模型解释可信度的新方法,发现其常隐瞒真实推理依据。
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
- 用辅助模型生成输入的反事实修改,模拟概念影响变化
- 通过贝叶斯层级模型计算概念在样本和数据集层面的因果效应
- 揭示模型解释中隐藏的社会偏见与误导性证据关联
大语言模型(LLMs)能生成看似合理的答案解释,但这些解释可能扭曲模型的真实推理过程,导致过度信任和误用。本文提出一种衡量解释可信度的新方法:首先给出可信度的严格定义——解释中声称有影响的概念集合,与实际起作用的概念集合之间的差异;其次设计新方法,利用辅助模型修改输入中的概念值以生成合理反事实样本,并采用贝叶斯层级模型量化概念在个体样本和整个数据集层面的因果效应。实验表明该方法可有效量化并发现可解释的不可信模式:在社会偏见任务中,发现模型解释掩盖了社会偏见的影响;在医学问答任务中,发现解释误导性地宣称某些证据主导了决策。
原文摘要 · Abstract (English)
Large language models (LLMs) are capable of generating plausible explanations of how they arrived at an answer to a question. However, these explanations can misrepresent the model's "reasoning" process, i.e., they can be unfaithful. This, in turn, can lead to over-trust and misuse. We introduce a new approach for measuring the faithfulness of LLM explanations. First, we provide a rigorous definition of faithfulness. Since LLM explanations mimic human explanations, they often reference high-level concepts in the input question that purportedly influenced the model. We define faithfulness in terms of the difference between the set of concepts that LLM explanations imply are influential and the set that truly are. Second, we present a novel method for estimating faithfulness that is based on: (1) using an auxiliary LLM to modify the values of concepts within model inputs to create realistic counterfactuals, and (2) using a Bayesian hierarchical model to quantify the causal effects of concepts at both the example- and dataset-level. Our experiments show that our method can be used to quantify and discover interpretable patterns of unfaithfulness. On a social bias task, we uncover cases where LLM explanations hide the influence of social bias. On a medical question answering task, we uncover cases where LLM explanations provide misleading claims about which pieces of evidence influenced the model's decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。