arXiv:2606.14838cs.AI2026-06

提出兼顾事实与听众认知的解释定义,揭示大模型难解释的原因。

A Definition of Good Explanations and the Challenges Explaining LLM Outputs

  • 基于反事实推理,引入听众先验信念作为解释标准
  • 指出大模型输出难以满足认知适配性要求
  • 适合关注AI可解释性与认知科学交叉的研究者

如何定义优质解释是长期存在的哲学议题,近年来在人工智能输出背景下重获关注。解释能力对AI在诸多场景中的采纳至关重要,但要生成高质量解释,必须首先明确什么是好解释。本文提出一种受反事实解释启发的定义,强调需考虑解释中每个事实所针对的对话者先验信念。我们探讨该定义对AI可解释性的启示,尤其阐明大语言模型输出难以生成优质解释的根本原因。

原文摘要 · Abstract (English)

How to define a good explanation is a long-standing philosophical debate which has found recent renewed interest in the context of AI outputs. Explainability is crucial for AI adoption in many contexts, but in order to produce good explanations of AI systems, we must first have an understanding of what good explanations are. In this paper we propose a definition inspired by the notion of counterfactual explanations, however we argue that one must also take into account the interlocutor's prior beliefs in each fact that could be offered in an explanation. We explore the ramifications of this definition for AI explainability and, in particular, why LLM outputs are difficult to produce good explanations for.

可解释性大模型认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。