研究正定矩阵神经网络的表达能力,发现常见约束会降低模型性能。
Expressivity of congruence-based architectures for DNNs on positive-definite matrices
- 采用左右乘权重矩阵的结构处理正定矩阵数据
- 半正交约束导致某些激活函数下模型退化为单层网络
- 分析了黎曼分类器与该架构的适配性,适合几何深度学习研究者
本文研究用于分类对称正定矩阵的神经架构,重点考察基于合同变换的层结构,即输入矩阵在左右分别与权重矩阵W及其转置相乘。这类结构是著名SPDNet的核心,并被独立用于正定数据的降维。研究发现,对W施加(半)正交性约束会限制其表达能力:对于某些激活函数,整个架构退化为等价于单隐层网络。这种表达能力的缺失源于半正交W下合同类层的谱多样性损失,由庞加莱分离定理直接决定。随后还探讨了最终分类器的选择,比较多种黎曼分类器,并讨论它们与合同类层生成特征图的兼容性。
原文摘要 · Abstract (English)
This work studies neural architectures for classifying symmetric positive-definite matrices, focusing on congruence-like layers, in which the input matrix is multiplied on the left and right by a (possibly rectangular) weight matrix $W$ and its transpose. Such layers lie at the core of the celebrated SPDNet and have also been employed independently for dimensionality reduction on positive-definite data. We show that the (semi)-orthogonality constraint commonly imposed on $W$ limits the expressivity of these layers: for certain activation functions, the resulting architecture collapses to a one-hidden-layer equivalent. This lack of expressivity follows from a loss of spectral diversity in congruence-like layers for semi-orthogonal $W$ and is a direct consequence of Poincaré's separation theorem. We then examine the choice of the final classifier, comparing several Riemannian classifiers and discussing their compatibility with the feature maps produced by congruence-like layers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。