arXiv:2509.22297cs.AI2025-09中稿 · KR 2026被引 4

提出新方法让大模型生成更符合其真实行为的反事实结果。

Large Language Models as Nondeterministic Causal Models

  • 将大模型视为非确定性因果模型,避免对采样过程的假设
  • 无需修改模型即可适用于任意黑箱大模型
  • 为不同用途设计特定反事实生成方法提供理论基础

Chatzi 等人和Ravfogel 等人首次提出了生成概率大语言模型反事实的方法,可回答‘若输入从𝐱变为𝐱*,输出可能为何’。该能力对解释、评估和改进大模型行为至关重要。但现有方法存在双重误解:既未忠实于大模型的实现机制,又将非确定性模型强行转为确定性因果模型。本文提出更简洁的新方法,基于大模型的本意,将其建模为非确定性因果模型。该方法不依赖任何实现细节,可直接应用于任意黑箱大模型。相较而言,原方法虽能生成特定类型反事实,但适用范围受限。本文通过建立基于意图语义的反事实推理理论框架,厘清二者关系,并为面向特定应用的反事实生成方法奠定基础。

原文摘要 · Abstract (English)

Recent work by Chatzi et al. and Ravfogel et al. has developed, for the first time, a method for generating counterfactuals of probabilistic Large Language Models. Such counterfactuals tell us what would - or might - have been the output of an LLM if some factual prompt ${\bf x}$ had been ${\bf x}^*$ instead. The ability to generate such counterfactuals is an important necessary step towards explaining, evaluating, and eventually improving, the behavior of LLMs. I argue, however, that the existing method rests on an ambiguous interpretation of LLMs: it does not interpret LLMs literally, for the method involves the assumption that one can change the implementation of an LLM's sampling process without changing the LLM itself, nor does it interpret LLMs as intended, for the method involves explicitly representing a nondeterministic LLM as a deterministic causal model. I here present a much simpler method for generating counterfactuals that is based on an LLM's intended interpretation by representing it as a nondeterministic causal model instead. The advantage of my simpler method is that it is directly applicable to any black-box LLM without modification, as it is agnostic to any implementation details. The advantage of the existing method, on the other hand, is that it directly implements the generation of a specific type of counterfactuals that is useful for certain purposes, but not for others. I clarify how both methods relate by offering a theoretical foundation for reasoning about counterfactuals in LLMs based on their intended semantics, thereby laying the groundwork for novel application-specific methods for generating counterfactuals.

大模型解释反事实推理因果模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。