揭示等变网络优化难题的根源:隐藏对称性阻碍学习,放松约束可跳出陷阱。
A Tale of Two Symmetries: Exploring the Loss Landscape of Equivariant Models
- 通过分析无约束模型的隐藏对称性,发现其影响等变子空间的损失曲面
- 在特定条件下,隐藏对称性会阻止达到全局最优解,导致学习失败
- 放松约束后网络跳转到不同群表示,提示需重新思考隐藏层的表示选择
等变神经网络在具有已知对称性的任务中表现优异,但其优化过程复杂,训练实践尚未成熟。近期研究发现,放松等变约束带来的收益微小,引发疑问:等变约束是否造成根本性优化障碍?还是仅需不同超参数调优?本文通过理论分析损失曲面几何,聚焦基于置换表示构建的网络,将其视为无约束MLP的子集。关键发现:无约束模型的参数对称性会对等变子空间的损失曲面产生非平凡影响,在特定条件下可证明阻碍全局最小值的学习。进一步实验表明,此时放松为无约束MLP有时可解决该问题。有趣的是,放松后找到的权重对应于隐藏层中不同的群表示选择。由此得出三点启示:(1) 观察无约束架构可揭示被约束强制打破的隐藏参数对称性;(2) 隐藏对称性对损失曲面有重要影响,可能生成临界点甚至极小值;(3) 隐藏对称性诱导的极小值可通过约束放松逃脱,网络会跳转至不同约束实现方式。有效的等变松弛可能需要重新考虑隐藏层固定的群表示选择。
原文摘要 · Abstract (English)
Equivariant neural networks have proven to be effective for tasks with known underlying symmetries. However, optimizing equivariant networks can be tricky and best training practices are less established than for standard networks. In particular, recent works have found small training benefits from relaxing equivariance constraints. This raises the question: do equivariance constraints introduce fundamental obstacles to optimization? Or do they simply require different hyperparameter tuning? In this work, we investigate this question through a theoretical analysis of the loss landscape geometry. We focus on networks built using permutation representations, which we can view as a subset of unconstrained MLPs. Importantly, we show that the parameter symmetries of the unconstrained model has nontrivial effects on the loss landscape of the equivariant subspace and under certain conditions can provably prevent learning of the global minima. Further, we empirically demonstrate in such cases, relaxing to an unconstrained MLP can sometimes solve the issue. Interestingly, the weights eventually found via relaxation corresponds to a different choice of group representation in the hidden layer. From this, we draw 3 key takeaways. (1) By viewing the unconstrained version of an architecture, we can uncover hidden parameter symmetries which were broken by choice of constraint enforcement (2) Hidden symmetries give important insights on loss landscapes and can induce critical points and even minima (3) Hidden symmetry induced minima can sometimes be escaped by constraint relaxation and we observe the network jumps to a different choice of constraint enforcement. Effective equivariance relaxation may require rethinking the fixed choice of group representation in the hidden layers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。