揭示深度多项式神经网络可识别性的关键条件,解析结构与参数的关系。
Identifiability of Deep Polynomial Neural Networks
- 通过张量分解和Kruskal定理,建立深层多项式网络与可识别性的联系。
- 非递增层宽结构在弱条件下具通用可识别性,编码器-解码器结构需解码层增长不过快。
- 解决关于神经代数簇维数的开放猜想,给出激活度所需新界。
多项式神经网络(PNNs)具有丰富的代数与几何结构,但其可识别性——保障模型可解释性的关键属性——仍不明确。本文对带与不带偏置项的深层PNN架构进行了系统分析。结果揭示了激活度与层宽之间复杂的相互作用关系:在弱条件下,层宽非递增的架构具有通用可识别性;编码器-解码器结构在解码层宽度增长速度不超过激活度时具备可识别性。证明过程基于深层PNN与低秩张量分解的联系,并运用Kruskal型唯一性定理。此外,本文解决了关于PNN神经代数簇维数的开放猜想,给出了达到期望维数所需的激活度新上界。
原文摘要 · Abstract (English)
Polynomial Neural Networks (PNNs) possess a rich algebraic and geometric structure. However, their identifiability -- a key property for ensuring interpretability -- remains poorly understood. In this work, we present a comprehensive analysis of the identifiability of deep PNNs, including architectures with and without bias terms. Our results reveal an intricate interplay between activation degrees and layer widths in achieving identifiability. As special cases, we show that architectures with non-increasing layer widths are generically identifiable under mild conditions, while encoder-decoder networks are identifiable when the decoder widths do not grow too rapidly compared to the activation degrees. Our proofs are constructive and center on a connection between deep PNNs and low-rank tensor decompositions, and Kruskal-type uniqueness theorems. We also settle an open conjecture on the dimension of PNN's neurovarieties, and provide new bounds on the activation degrees required for it to reach the expected dimension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。