首次揭示变压器符号回归模型内部工作机制,发现28个关键电路。
Explaining the Explainer: Understanding the Inner Workings of Transformer-based Symbolic Regression Models
- 提出进化算法PATCHES,自动挖掘符号回归模型中的核心计算电路。
- 验证28个电路功能正确,均通过忠实性、完备性与最小性三重测试。
- 证明性能导向的平均修补法优于传统归因方法,更适用于因果发现。
随着在多个领域的成功应用,变压器在符号回归(SR)中也表现出色;然而其生成数学算子的内部机制仍不明确。尽管机制可解释性已在语言和视觉模型中成功识别出电路,但尚未应用于符号回归。本文提出PATCHES——一种进化电路发现算法,用于识别符号回归中的紧凑且正确的电路。利用该方法,我们分离出28个电路,首次实现了对符号回归变压器的电路级表征。通过基于忠实性、完备性和最小性的稳健因果评估框架验证这些发现。分析表明,基于性能的平均修补法最可靠地识别出功能正确的电路;而直接的logit归因和探针分类器主要捕捉相关特征而非因果特征,限制了其在电路发现中的应用。总体而言,这些结果确立了符号回归作为机制可解释性的重要应用领域,并提出了一种原则性的电路发现方法。
原文摘要 · Abstract (English)
Following their success across many domains, transformers have also proven effective for symbolic regression (SR); however, the internal mechanisms underlying their generation of mathematical operators remain largely unexplored. Although mechanistic interpretability has successfully identified circuits in language and vision models, it has not yet been applied to SR. In this article, we introduce PATCHES, an evolutionary circuit discovery algorithm that identifies compact and correct circuits for SR. Using PATCHES, we isolate 28 circuits, providing the first circuit-level characterisation of an SR transformer. We validate these findings through a robust causal evaluation framework based on key notions such as faithfulness, completeness, and minimality. Our analysis shows that mean patching with performance-based evaluation most reliably isolates functionally correct circuits. In contrast, we demonstrate that direct logit attribution and probing classifiers primarily capture correlational features rather than causal ones, limiting their utility for circuit discovery. Overall, these results establish SR as a high-potential application domain for mechanistic interpretability and propose a principled methodology for circuit discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。