arXiv:2411.01375stat.MLcs.AI2024-11被引 2

神经网络能利用隐藏因子结构,更高效地学习高维离散分布。

Learning with Hidden Factorial Structure

  • 通过分解复杂任务为简单子任务,揭示数据中的隐藏因子结构。
  • 实验表明神经网络可利用此类结构,提升学习效率。
  • 适合研究模型泛化能力与结构假设的学者参考。

高维空间中的统计学习若缺乏强数据结构将面临挑战。近期基础模型的发展表明,文本与图像数据包含此类隐藏结构,有助于缓解维度灾难。受非参数统计研究启发,我们假设这一现象可部分归因于将复杂任务分解为简单子任务。本文提出一个受控实验框架,检验神经网络是否真能利用这种“隐藏因子结构”。结果表明,神经网络确实能借助这些潜在模式,更高效地学习离散分布。我们还研究了结构假设与模型泛化能力之间的相互作用。

原文摘要 · Abstract (English)

Statistical learning in high-dimensional spaces is challenging without a strong underlying data structure. Recent advances with foundational models suggest that text and image data contain such hidden structures, which help mitigate the curse of dimensionality. Inspired by results from nonparametric statistics, we hypothesize that this phenomenon can be partially explained in terms of decomposition of complex tasks into simpler subtasks. In this paper, we present a controlled experimental framework to test whether neural networks can indeed exploit such "hidden factorial structures". We find that they do leverage these latent patterns to learn discrete distributions more efficiently. We also study the interplay between our structural assumptions and the models' capacity for generalization.

神经网络因子结构高维学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。