arXiv:2608.16747cs.LGcs.AI2026-08

用反事实实验检验大模型解释的有效性,发现现有方法大多无效。

Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

论文配图:Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
图 1 · 摘自论文原文
  • 构建自动反事实分析管道CHIVE,挖掘真实场景中的模型异常行为。
  • 测试多种解释方法后发现均无法提升对反事实行为的预测能力。
  • 基于反事实数据训练的模型能泛化到分布外场景,效果显著。

许多AI研究领域,如大语言模型可解释性和思维链忠实性,旨在解释模型行为。但什么是‘好’的解释?本文从反事实可模拟性角度评估解释——即解释能否帮助预测模型在相关反事实输入下的行为。为此,我们提出CHIVE(反事实假设调查通过编辑),一个新型代理式流程,用于识别现实场景中意外的模型行为,并通过反事实提示编辑进行探究。该方法生成数千条高质量解释及其支持的反事实证据。我们以两种方式应用CHIVE:第一,评估常见可解释性技术是否提升代理预测反事实行为的能力,结果令人意外——所有技术均未带来提升;第二,利用CHIVE生成训练数据,发现训练模型预测CHIVE生成的反事实实验结果,能泛化至多种分布外设置。总体而言,CHIVE可自动发现自然发生的大模型行为解释,为评估和改进解释方法提供了新路径。

原文摘要 · Abstract (English)

Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But what constitutes a "good" explanation? In this work, we evaluate explanations through the lens of counterfactual simulatability-whether the explanation is useful for predicting model behaviors on related counterfactual inputs. To this end, we introduce CHIVE (Counterfactual Hypothesis Investigation Via Edits), a novel agentic pipeline that identifies unexpected model behaviors in the wild and investigates them with counterfactual prompt edits. This yields thousands of high-quality explanations for naturally-occurring model behaviors along with supporting counterfactual evidence. We apply CHIVE in two ways. First, we evaluate whether common LLM interpretability techniques improve an agent's ability to predict counterfactual model behaviors. Surprisingly, we find no uplift from any of the interpretability techniques studied. Second, we use CHIVE to generate training data. We find that training models to predict outcomes of CHIVE-generated counterfactual experiments generalizes to various out-of-distribution settings. Overall, CHIVE automatically discovers explanations of naturally-occurring LLM behaviors, enabling us to evaluate and improve methods for explaining LLM behaviors.

可解释性反事实大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。