arXiv:2602.18905cs.LGcs.AI2026-02

提出可验证的统一解释框架,让大模型推理过程更可信、可分析。

TRUE: A Trustworthy Unified Explanation Framework for Large Language Model Reasoning

  • 将推理路径转为可执行规范,通过盲执行验证有效性
  • 构建可行区域有向图,揭示局部输入稳定性与可执行范围
  • 识别共性失败模式并量化其因果影响,适合调试与优化模型

大语言模型在复杂推理任务中表现强劲,但其决策过程难以解释。现有解释方法缺乏可信的结构洞察,且仅限单个实例分析,无法揭示推理稳定性与系统性失效机制。为此,我们提出可信统一解释框架TRUE,融合可执行推理验证、可行区域有向图建模与因果失效模式分析。在实例层面,将推理轨迹重定义为可执行过程规范,并引入盲执行验证评估运行有效性。在局部结构层面,通过结构一致扰动构建可行区域有向图,显式刻画推理稳定性及局部输入空间中的可执行区域。在类别层面,提出因果失效模式分析方法,识别重复出现的结构性失效模式,并使用Shapley值量化其因果影响。多个推理基准上的实验表明,该框架提供了多层次可验证的解释:个体实例的可执行推理结构、邻近输入的可行区域表示,以及类级别可解释且带重要性量化值的失效模式。这些结果建立了一个统一、严谨的提升大模型推理可解释性与可靠性的范式。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated strong capabilities in complex reasoning tasks, yet their decision-making processes remain difficult to interpret. Existing explanation methods often lack trustworthy structural insight and are limited to single-instance analysis, failing to reveal reasoning stability and systematic failure mechanisms. To address these limitations, we propose the Trustworthy Unified Explanation Framework (TRUE), which integrates executable reasoning verification, feasible-region directed acyclic graph (DAG) modeling, and causal failure mode analysis. At the instance level, we redefine reasoning traces as executable process specifications and introduce blind execution verification to assess operational validity. At the local structural level, we construct feasible-region DAGs via structure-consistent perturbations, enabling explicit characterization of reasoning stability and the executable region in the local input space. At the class level, we introduce a causal failure mode analysis method that identifies recurring structural failure patterns and quantifies their causal influence using Shapley values. Extensive experiments across multiple reasoning benchmarks demonstrate that the proposed framework provides multi-level, verifiable explanations, including executable reasoning structures for individual instances, feasible-region representations for neighboring inputs, and interpretable failure modes with quantified importance at the class level. These results establish a unified and principled paradigm for improving the interpretability and reliability of LLM reasoning systems.

大模型解释推理可解释可信推理因果分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。