arXiv:2507.08802cs.LG2025-07NeurIPS被引 28

打破线性假设后,因果抽象变得无意义,无法真正解释模型机制。

The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?

  • 放宽映射函数的线性限制,使任意模型可对应任意算法
  • 随机初始化语言模型也能实现100%干预准确率,但根本不会解任务
  • 揭示非线性表示困境:复杂度与精度难平衡,需新假设支撑

因果抽象被用于解析机器学习模型的黑箱决策过程:若存在函数可在模型与高层算法间双向映射,则该模型可被视为该算法的抽象。当前多数工作将映射设为线性,源于线性表征假说。然而,定义并不要求线性。本文证明,在合理假设下,任意神经网络均可映射至任意算法,导致无约束的因果抽象变得平凡且无信息量。实证显示,即使随机初始化的语言模型,其对间接宾语识别任务的干预-交换准确率仍可达100%。这引发非线性表征困境:放弃线性约束后,缺乏原则性方法权衡映射的复杂度与精度。结果表明,仅靠因果抽象不足以实现机制可解释性,必须依赖关于信息编码方式的额外假设。未来研究应探索该假设与因果抽象的关系。

原文摘要 · Abstract (English)

The concept of causal abstraction got recently popularised to demystify the opaque decision-making processes of machine learning models; in short, a neural network can be abstracted as a higher-level algorithm if there exists a function which allows us to map between them. Notably, most interpretability papers implement these maps as linear functions, motivated by the linear representation hypothesis: the idea that features are encoded linearly in a model's representations. However, this linearity constraint is not required by the definition of causal abstraction. In this work, we critically examine the concept of causal abstraction by considering arbitrarily powerful alignment maps. In particular, we prove that under reasonable assumptions, any neural network can be mapped to any algorithm, rendering this unrestricted notion of causal abstraction trivial and uninformative. We complement these theoretical findings with empirical evidence, demonstrating that it is possible to perfectly map models to algorithms even when these models are incapable of solving the actual task; e.g., on an experiment using randomly initialised language models, our alignment maps reach 100\% interchange-intervention accuracy on the indirect object identification task. This raises the non-linear representation dilemma: if we lift the linearity constraint imposed to alignment maps in causal abstraction analyses, we are left with no principled way to balance the inherent trade-off between these maps' complexity and accuracy. Together, these results suggest an answer to our title's question: causal abstraction is not enough for mechanistic interpretability, as it becomes vacuous without assumptions about how models encode information. Studying the connection between this information-encoding assumption and causal abstraction should lead to exciting future work.

因果抽象可解释性非线性表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。