arXiv:2607.15957cs.CL2026-07

大模型自解释虽看似合理,但未必真实,却仍可指导决策。

From Plausible to Actionable: A Position on LLM Self-Explanations

  • 分析大模型自解释的合理性与真实性差异
  • 提出评估自解释可信度的实用指南
  • 强调其在实际决策中的可用性价值

大型语言模型(LLMs)能够生成自然语言解释以说明自身决策过程,这一现象被称为自解释。尽管这些解释往往显得合理,但它们是否真实反映模型内部推理过程仍存疑。本文认为,自解释可能高度合理、可信度存疑,但仍具高度可操作性。从传统可解释人工智能(XAI)视角出发,指出现有评估协议的局限性,并提出评估自解释合理性和忠实性的实践指南。此外,主张评估应拓展至可操作性,强调大模型推理能力在支持多方利益相关者进行明智决策和采取适当行动方面的应用价值。

原文摘要 · Abstract (English)

Large Language Models (LLMs) can generate natural language explanations that rationalize their own decisions, a phenomenon commonly referred to as self-explanations.Such explanations have emerged as a promising direction for explainable artificial intelligence (XAI), particularly for interpreting LLM behavior.However, while self-explanations often appear plausible, whether they faithfully reflect a model's underlying reasoning process remains an open question. In this opinion paper, we argue that self-explanations can be highly plausible, questionably faithful, and yet highly actionable. From a traditional XAI perspective, we identify the limitations of standard evaluation protocols for LLM-generated self-explanations and propose practical guidelines for assessing their plausibility and faithfulness. Moreover, we argue that evaluation should extend beyond these criteria to actionability, highlighting applications of LLM rationalization capabilities that support informed decision-making and appropriate action across diverse stakeholders.

可解释AI大模型决策支持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。