arXiv:2502.15801cs.LGcs.AI2025-02被引 9

揭示了小型Transformer中组合归纳的可解释电路,实现行为精准调控。

An explainable transformer circuit for compositional generalization

  • 通过因果消融定位并解析组合归纳的内部电路机制。
  • 发现该电路能精确预测并控制模型在新结构上的表现。
  • 为大模型行为干预提供可解释的路径,适合研究可解释AI者参考。

组合泛化——将已知组件系统性地组合成新结构——仍是认知科学与机器学习的核心挑战。尽管基于Transformer的大语言模型在某些组合任务上表现出色,但其背后驱动能力的机制仍不清晰,影响了模型的可解释性。本文在小型Transformer中识别出负责组合归纳的电路,并通过因果消融验证其作用,以类似程序的描述形式形式化其运作方式。进一步表明,这种机制理解可实现对模型行为的精准激活编辑,从而可预测地调控模型表现。研究深化了对Transformer复杂行为的理解,并表明此类洞察可直接用于模型控制。

原文摘要 · Abstract (English)

Compositional generalization-the systematic combination of known components into novel structures-remains a core challenge in cognitive science and machine learning. Although transformer-based large language models can exhibit strong performance on certain compositional tasks, the underlying mechanisms driving these abilities remain opaque, calling into question their interpretability. In this work, we identify and mechanistically interpret the circuit responsible for compositional induction in a compact transformer. Using causal ablations, we validate the circuit and formalize its operation using a program-like description. We further demonstrate that this mechanistic understanding enables precise activation edits to steer the model's behavior predictably. Our findings advance the understanding of complex behaviors in transformers and highlight such insights can provide a direct pathway for model control.

可解释性Transformer组合泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。