arXiv:2412.11867cs.LGcs.AI2024-12被引 16

Transformer在迷宫任务中自发构建因果世界模型,能泛化推理复杂结构。

Transformers Use Causal World Models in Maze-Solving Tasks

  • 用稀疏自编码器与注意力模式分析发现模型具结构化世界表征。
  • 激活特定特征比抑制更易实现,且模型可推理训练外的复杂迷宫。
  • 位置编码方式影响世界模型在残差流中的组织结构,适合可解释性研究。

近期可解释性研究揭示,训练于不同任务的Transformer模型常自发形成高度结构化的内部表征。当这些表征全面反映任务域结构时,被称为“世界模型”(World Models, WMs)。本文在训练于迷宫求解任务的Transformer中识别出此类模型。通过稀疏自编码器(SAEs)与注意力模式分析,我们考察了世界模型的构建过程,并证实了基于SAE特征与电路分析的一致性。进一步对孤立特征进行干预以验证其因果作用,发现激活特征比抑制更容易。此外,模型能够推理包含比训练中更多活跃特征的迷宫;但当相同复杂度的迷宫以输入标记形式给出时,模型则失效。最后,我们发现位置编码方案会影响世界模型在模型残差流中的结构组织方式。

原文摘要 · Abstract (English)

Recent studies in interpretability have explored the inner workings of transformer models trained on tasks across various domains, often discovering that these networks naturally develop highly structured representations. When such representations comprehensively reflect the task domain's structure, they are commonly referred to as "World Models" (WMs). In this work, we identify WMs in transformers trained on maze-solving tasks. By using Sparse Autoencoders (SAEs) and analyzing attention patterns, we examine the construction of WMs and demonstrate consistency between SAE feature-based and circuit-based analyses. By subsequently intervening on isolated features to confirm their causal role, we find that it is easier to activate features than to suppress them. Furthermore, we find that models can reason about mazes involving more simultaneously active features than they encountered during training; however, when these same mazes (with greater numbers of connections) are provided to models via input tokens instead, the models fail. Finally, we demonstrate that positional encoding schemes appear to influence how World Models are structured within the model's residual stream.

世界模型可解释性Transformer因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。