揭示神经网络在群运算任务中的新结构,提升解释一致性与准确性。
Towards a unified and verified understanding of group-operation networks
- 发现模型对输入变量的等变性近似,揭示内部计算机制。
- 解释可覆盖45%模型,准确率超95%,推理速度比暴力法快3倍。
- 适用于对群运算建模的神经网络,适合关注可解释性的研究者。
近期机械可解释性研究聚焦于逆向解析训练于有限群二元运算的神经网络的计算过程。本文研究单隐藏层神经网络在此任务中的内部结构,揭示了此前未被识别的规律,推动了对已有工作(Chughtai et al., 2023;Stander et al., 2024)解释的统一。显著发现:这些模型在每个输入变量上近似保持等变性。我们通过构建紧凑证明,验证该解释适用于大量此类网络,并定量评估其对模型内部机制的忠实与简洁程度。以对称群S5为例,该解释能提供比暴力计算快3倍的准确率保证,在45%的训练模型中达到≥95%的准确率边界。而仅依赖先前解释无法获得非平凡、非平凡的准确率界。
原文摘要 · Abstract (English)
A recent line of work in mechanistic interpretability has focused on reverse-engineering the computation performed by neural networks trained on the binary operation of finite groups. We investigate the internals of one-hidden-layer neural networks trained on this task, revealing previously unidentified structure and producing a more complete description of such models in a step towards unifying the explanations of previous works (Chughtai et al., 2023; Stander et al., 2024). Notably, these models approximate equivariance in each input argument. We verify that our explanation applies to a large fraction of networks trained on this task by translating it into a compact proof of model performance, a quantitative evaluation of the extent to which we faithfully and concisely explain model internals. In the main text, we focus on the symmetric group S5. For models trained on this group, our explanation yields a guarantee of model accuracy that runs 3x faster than brute force and gives a >=95% accuracy bound for 45% of the models we trained. We were unable to obtain nontrivial non-vacuous accuracy bounds using only explanations from previous works.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。