提出新型激活函数HTAF,实现稳定二值化表示学习
A Composite Activation Function for Learning Stable Binary Representations

- 设计分段复合激活函数HTAF,近似海维赛德函数
- 在图像分类任务中达到与标准模型相当甚至更优性能
- 适合需要可解释性二值特征的视觉模型研究
激活函数在神经网络中起关键作用,决定内部表征形态。近年来,学习二值激活表示因其计算与存储效率高、可解释性强而备受关注。然而,由于海维赛德激活函数不可导,基于梯度的优化难以训练。本文提出重尾激活函数(HTAF),一种平滑逼近海维赛德函数的方法,支持基于梯度的稳定训练。HTAF由双曲正切与指数衰减的组合构成,理论表明其在零输入附近保持大梯度质量,并在尾部区域梯度衰减更慢。实验显示,脉冲神经网络、二值神经网络及深度海维赛德网络均可通过HTAF实现稳定训练。此外,我们引入隐式概念瓶颈模型(ICBMs),利用HTAF生成离散特征表示。在多种架构与图像数据集上的实验证明,ICBM在实现稳定离散化的同时,预测性能优于或相当于标准模型。
原文摘要 · Abstract (English)
Activation functions play a central role in neural networks by shaping internal representations. Recently, learning binary activation representations has attracted significant attention due to their advantages in computational and memory efficiency, as well as interpretability. However, training neural networks with Heaviside activations remains challenging, as their non-differentiability obstructs standard gradient-based optimization. In this paper, we propose Heavy Tailed Activation Function (HTAF), a smooth approximation to the Heaviside function that enables stable training with gradient-based optimization. We construct HTAF as a sigmoid hyperbolic tangent composite function and theoretically show that it maintains a large gradient mass around zero inputs while exhibiting slower gradient decay in the tail regions. We show that Spiking Neural Networks, Binary Neural Networks and Deep Heaviside neural Networks can be trained stably using HTAF with gradient-based optimization. Finally, we introduce Implicit Concept Bottleneck Models (ICBMs), an interpretable image model that leverages HTAF to induce discrete feature representations. Extensive experiments across various architectures and image datasets demonstrate that ICBM enables stable discretization while achieving prediction performance comparable to or better than standard models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。