arXiv:2506.02946cs.LG2025-06NeurIPS被引 1

为大模型智能体设计了更合理的反事实推理方法。

Abstract Counterfactuals for Language Model Agents

  • 用高层动作特征替代细粒度词元,避免语义漂移
  • 在文本游戏和生成任务中实现一致且有意义的反事实结果
  • 适合关注智能体行为可解释性的研究者

反事实推理是分析和评估自主智能体的强大工具,但其在语言模型(LM)智能体中的应用仍具挑战。现有工作主要聚焦于词元级反事实,因语言模型智能体具有开放式的行动空间,此类方法常不适用。传统智能体的行动空间固定明确,而语言模型智能体的行动通常隐含在其输出字符串中,难以定义与解读。此外,单个词元的含义会随上下文变化,使词元级推理复杂化,可能导致偏差或无意义的反事实。本文提出「抽象反事实」框架,强调环境内动作与交互的高层特征,实现面向用户相关特性的反事实推理。实验表明,该方法生成的反事实结果一致且有意义,显著减少了词元级方法的副作用。我们在文本游戏和反事实文本生成任务上进行了测试,同时考察了词元级与潜在空间干预的效果。

原文摘要 · Abstract (English)

Counterfactual inference is a powerful tool for analysing and evaluating autonomous agents, but its application to language model (LM) agents remains challenging. Existing work on counterfactuals in LMs has primarily focused on token-level counterfactuals, which are often inadequate for LM agents due to their open-ended action spaces. Unlike traditional agents with fixed, clearly defined action spaces, the actions of LM agents are often implicit in the strings they output, making their action spaces difficult to define and interpret. Furthermore, the meanings of individual tokens can shift depending on the context, adding complexity to token-level reasoning and sometimes leading to biased or meaningless counterfactuals. We introduce \emph{Abstract Counterfactuals}, a framework that emphasises high-level characteristics of actions and interactions within an environment, enabling counterfactual reasoning tailored to user-relevant features. Our experiments demonstrate that the approach produces consistent and meaningful counterfactuals while minimising the undesired side effects of token-level methods. We conduct experiments on text-based games and counterfactual text generation, while considering both token-level and latent-space interventions.

反事实推理大模型智能体可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。