揭示有限ReLU网络中结构化混合选择特征的涌现机制
Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU Networks
- 通过等价变换分析ReLU网络学习动态
- 发现多任务情境下特征表示具高度可复用的结构化混合选择性
- 适合研究深度学习表征形成与网络归纳偏置的学者
尽管有限维度的ReLU神经网络是近期深度学习成功的关键因素,但其特征学习理论仍不明确。现有理论依赖于线性计算、无结构输入或无限宽度等假设。本文建立ReLU网络与门控深层线性网络的等价关系,利用后者更高的可解析性推导学习动态。研究了类似多任务学习或上下文控制的核心任务,发现此类任务中,ReLU网络具有向非严格模块化但高度结构化的潜在表示倾斜的归纳偏好,该特性随上下文数量和隐藏层加深而增强。本文为有限维度ReLU网络的特征学习理论迈出关键一步,并揭示了节点复用与学习速度共同导致结构化混合选择表示涌现的机制。
原文摘要 · Abstract (English)
In spite of finite dimension ReLU neural networks being a consistent factor behind recent deep learning successes, a theory of feature learning in these models remains elusive. Currently, insightful theories still rely on assumptions including the linearity of the network computations, unstructured input data and architectural constraints such as infinite width or a single hidden layer. To begin to address this gap we establish an equivalence between ReLU networks and Gated Deep Linear Networks, and use their greater tractability to derive dynamics of learning. We then consider multiple variants of a core task reminiscent of multi-task learning or contextual control which requires both feature learning and nonlinearity. We make explicit that, for these tasks, the ReLU networks possess an inductive bias towards latent representations which are not strictly modular or disentangled but are still highly structured and reusable between contexts. This effect is amplified with the addition of more contexts and hidden layers. Thus, we take a step towards a theory of feature learning in finite ReLU networks and shed light on how structured mixed-selective latent representations can emerge due to a bias for node-reuse and learning speed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。