arXiv:2605.01164cs.AI2026-05

警告:别再把大模型当决策解释器,它只会编理由。

LLMs Should Not Yet Be Credited with Decision Explanation

  • 区分预测、理由生成和真正解释,防止混淆
  • 现有证据只支持预测与合理化,不支持真正解释
  • 适合研究人类行为建模的学者警惕夸大模型能力

本文主张,当前不应将决策解释的功劳归于大语言模型。近年来,研究者常将准确的行为预测、看似合理的理由以及结果依赖的推理轨迹视为模型解释人类决策的证据,这可能导致对解释性进步的过早定义。本文首先区分三种不同证据要求的主张:决策预测、理由生成与决策解释。接着指出,目前最常见的证据仅支持前两项,有时支持解释性假设生成,但无法区分解释与仅支持预测的合理化。为此提出解释信用的桥梁标准:更强主张需明确解释目标、排除弱替代方案、采用目标适配的过程或干预敏感验证,并限定适用范围。最后将其置于竞争观点与相关文献中,阐明该标准在保留模型作为预测者、叙述者与假设生成者价值的同时,抵制了过早赋予其解释信用。结论提出信用校准原则:模型只能获得其证据所支持的最强主张,且不得更多;若采纳此原则,可使模型从说服力强的决策叙述者转变为更可靠的发现、检验与传播人类行为解释的工具。

原文摘要 · Abstract (English)

This position paper argues that LLMs should not yet be credited with decision explanation. This matters because recent work increasingly treats accurate behavioral prediction, plausible rationales, and outcome-conditioned reasoning traces as evidence that LLMs explain why people decide as they do, risking a premature redefinition of what counts as explanatory progress in human decision modeling. We first distinguish three claims with different evidential burdens: decision prediction, rationale generation, and decision explanation. We then argue that the evidence most commonly offered for LLM-based decision accounts directly supports the first two claims, and sometimes explanatory hypothesis generation, but does not distinguish decision explanation from prediction-supportive rationalization. Next, we propose a bridge standard for decision-explanation credit: stronger claims should specify explanatory targets, discriminate against weaker rationalizer alternatives, use target-appropriate process- or intervention-sensitive validation, and bound their scope. We then situate this standard against competing views and related literatures, clarifying why it preserves the value of LLMs as predictors, narrators, and hypothesis generators while resisting premature explanatory credit. We conclude with a principle of credit calibration: LLMs should be credited for the strongest claim their evidence warrants, and no stronger; if adopted, this principle can help turn LLMs from persuasive narrators of decisions into more reliable instruments for discovering, testing, and communicating explanations of human behavior.

大模型决策解释可信推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。