让大模型自解释,能帮人更好预测它对变体问题的回答。
Do LLM Self-Explanations Help Users Predict Model Behavior? Evaluating Counterfactual Simulatability with Pragmatic Perturbations
- 用语义合理的扰动构建反事实问题,测试用户能否模拟模型行为。
- 有自解释时,人类和大模型的预测准确率均提升,最高达27%。
- 自解释对强判断者帮助更大,且效果受扰动方式影响显著。
大型语言模型(LLMs)能生成口头自解释,但先前研究认为这些理由未必真实反映其决策过程。本文探讨这类解释是否有助于用户预测模型行为,以反事实可模拟性为衡量标准。在StrategyQA数据集上,评估人类与大模型裁判在有无模型思维链或事后解释的情况下,预测模型对反事实追问答案的能力。比较了大模型生成的反事实与基于语用学的扰动作为测试用例的有效性。结果表明,自解释能持续提升人类与大模型的模拟准确率,但提升幅度和稳定性强烈依赖于扰动策略和判断者能力。对人类用户自由文本理由的定性分析显示,访问解释有助于形成更准确的预测。
原文摘要 · Abstract (English)
Large Language Models (LLMs) can produce verbalized self-explanations, yet prior studies suggest that such rationales may not reliably reflect the model's true decision process. We ask whether these explanations nevertheless help users predict model behavior, operationalized as counterfactual simulatability. Using StrategyQA, we evaluate how well humans and LLM judges can predict a model's answers to counterfactual follow-up questions, with and without access to the model's chain-of-thought or post-hoc explanations. We compare LLM-generated counterfactuals with pragmatics-based perturbations as alternative ways to construct test cases for assessing the potential usefulness of explanations. Our results show that self-explanations consistently improve simulation accuracy for both LLM judges and humans, but the degree and stability of gains depend strongly on the perturbation strategy and judge strength. We also conduct a qualitative analysis of free-text justifications written by human users when predicting the model's behavior, which provides evidence that access to explanations helps humans form more accurate predictions on the perturbed questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。