大模型能自己解释反事实推理吗?实验发现它们常自相矛盾。
Can LLMs Explain Themselves Counterfactually?
- 让大模型生成自身输出的反事实解释
- 多数模型生成的解释与预测结果不一致
- 适合研究模型可解释性与可信度的学者
解释是理解机器学习模型行为、建立用户信任和满足监管要求的重要工具。近年来,随着大语言模型(LLM)卓越的推理能力,自解释(即通过提示让模型解释自身输出)成为新范式。本文研究一种特定类型的自解释——自生成反事实解释(SCEs)。我们设计了测试来评估不同大模型家族、模型规模、温度设置及数据集下生成SCEs的有效性。分析表明,大模型有时难以生成有效的反事实解释,即使生成,其预测结果也常与自身解释不一致。
原文摘要 · Abstract (English)
Explanations are an important tool for gaining insights into the behavior of ML models, calibrating user trust and ensuring regulatory compliance. Past few years have seen a flurry of post-hoc methods for generating model explanations, many of which involve computing model gradients or solving specially designed optimization problems. However, owing to the remarkable reasoning abilities of Large Language Model (LLMs), self-explanation, that is, prompting the model to explain its outputs has recently emerged as a new paradigm. In this work, we study a specific type of self-explanations, self-generated counterfactual explanations (SCEs). We design tests for measuring the efficacy of LLMs in generating SCEs. Analysis over various LLM families, model sizes, temperature settings, and datasets reveals that LLMs sometimes struggle to generate SCEs. Even when they do, their prediction often does not agree with their own counterfactual reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。