复杂任务下神经网络实际使用分段线性区域减少,揭示表达能力饱和现象。
Expressivity Saturation: Reduced Affine Region Usage Under Increasing Task Complexity

- 通过线段探测推导出分段线性区域数的上界,给出实现复杂信号的最小神经元需求
- 固定架构下任务越复杂,训练后实际激活的分段区域显著减少,出现表达力饱和
- 可视化显示训练过程区域划分持续优化,但高难度任务中决策边界会退化
分段线性神经网络(如使用ReLU或LeakyReLU激活)实现连续分段线性映射,其分段区域数量可作为表达能力的自然度量。然而,理论最大区域容量与训练后实际实现的区域数之间的差距仍不明确。本文从两个互补角度研究该问题:首先,针对多层感知机的分段线性激活,提出一个严格且依赖架构的定理,证明沿仿射线段探测时,实际实现的分段线性片段数被各层宽度项(及激活断点因子)的乘积所上界。这给出了表示特定一维复杂度目标函数所需的最小神经元阈值。其次,精确枚举在受控任务复杂度下有限2维及以上域内实际实现的分段区域。在固定架构和训练协议下,输入-标签复杂度增加时,评估域中实际使用的分段区域显著减少,尽管理论最大容量不变;我们称此为表达力饱和。此外,在最困难情形下,2维可视化显示区域使用崩溃常伴随决策边界退化。最后,可视化训练过程中分段区域划分与决策边界的动态演变,揭示优化过程具有稳定的细化特征。
原文摘要 · Abstract (English)
Piecewise-affine neural networks (e.g., with ReLU or LeakyReLU activations) implement continuous piecewise-affine maps, and the number of affine regions provides a natural proxy for expressive capacity. However, the gap between theoretical region capacity and the affine regions realized after training remains insufficiently understood. We study this gap from two complementary perspectives. First, we give a rigorous, architecture-dependent theorem for affine line-segment probes: for multilayer perceptrons with piecewise-affine activations, the number of affine pieces realized along an affine line-segment probe is upper bounded by an explicit product of layer-wise width terms (and activation breakpoint factors). This yields a neuron-threshold lower bound for representing target functions with prescribed one-dimensional piece complexity, formalizing the minimal region budget required for complex signals. Second, we exactly enumerate affine regions realized within bounded 2D and higher-dimensional domains under controlled task complexity. Under fixed architectures and training protocols, increasing input--label complexity yields trained solutions with markedly fewer realized regions in the evaluation domain, even though worst-case architectural capacity is unchanged; we call this reduced region usage expressivity saturation. Moreover, in the most challenging regimes, 2D visualizations show that region-usage collapse often coincides with degraded decision boundaries. Finally, we visualize the training dynamics of affine-region partitions and decision boundaries, revealing a consistent refinement process during optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。