arXiv:2509.01874cs.LGcs.AI2025-09

提出SQS激活函数,让神经网络在保持性能的同时可解释权重特征。

Preserving Bilinear Weight Spectra with a Signed and Shrunk Quadratic Activation Function

  • 设计带符号的二次收缩激活函数,提升门控线性单元的可解释性。
  • 实验显示性能媲美顶尖激活函数,且支持权重层面的可解释分析。
  • 适合关注模型可解释性与计算效率的研究者或工程师。

理解机器学习模型的内部机制对确保其可靠性与鲁棒性至关重要。尽管许多机制可解释性方法聚焦于激活值分析,但能直接从神经网络权重中提取有意义特征的方法将提供更强的保证和更高的计算效率。现有通过权重分析模型特征的技术存在性能下降和数据效率低下的问题。本文提出签名二次收缩(Signed Quadratic Shrink, SQS)激活函数,使门控线性单元(GLUs)能够在不牺牲性能的前提下学习可解释特征。实验结果表明,SQS在性能上可与当前最先进的激活函数相媲美,同时实现了基于权重的可解释性。

原文摘要 · Abstract (English)

Understanding the inner workings of machine learning models is critical for ensuring their reliability and robustness. Whilst many techniques in mechanistic interpretability focus on activation driven analyses, being able to derive meaningful features directly from the weights of a neural network would provide greater guarantees and more computational efficiency. Existing techniques for analyzing model features through weights suffer from drawbacks such as reduced performance and data inefficiency. In this paper, we introduce Signed Quadratic Shrink (SQS), an activation function designed to allow Gated Linear Units (GLUs) to learn interpretable features without these drawbacks. Our experimental results show that SQS achieves performance competitive with state-of-the-art activation functions whilst enabling weight-based interpretability

可解释性激活函数权重分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。