揭示深度学习中数据与任务的局部性与组合性结构如何影响模型性能。
The Physics of Data and Tasks: Theories of Locality and Compositionality in Deep Learning
- 从局部性和组合性视角分析数据与任务的内在结构。
- 发现这些结构能显著提升泛化能力,训练样本越多效果越明显。
- 适合对深度学习原理感兴趣的科研人员与算法工程师。
深度神经网络取得了显著成功,但对其学习机制的理解仍有限。这些模型能够学习高维任务,而这类问题通常因维度灾难导致统计不可行。这一矛盾暗示可学习数据必然存在潜在的低维结构。这种结构的本质是什么?神经网络如何编码并利用它?它如何定量影响性能,例如泛化能力随训练样本数量增长的规律?本论文通过研究数据、任务与深度学习表征中的局部性与组合性,回答了这些问题。
原文摘要 · Abstract (English)
Deep neural networks have achieved remarkable success, yet our understanding of how they learn remains limited. These models can learn high-dimensional tasks, which is generally statistically intractable due to the curse of dimensionality. This apparent paradox suggests that learnable data must have an underlying latent structure. What is the nature of this structure? How do neural networks encode and exploit it, and how does it quantitatively impact performance - for instance, how does generalization improve with the number of training examples? This thesis addresses these questions by studying the roles of locality and compositionality in data, tasks, and deep learning representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。