arXiv:2505.23960cs.LGcs.AI2025-05被引 4

提出量化映射结构的方法,解析大模型如何学习与泛化。

Information Structure in Mappings: An Approach to Learning, Representation, and Generalisation

  • 用信息论量化映射中的结构原语,分析表示空间。
  • 揭示语言结构与神经网络性能结构的相似性。
  • 适用于百万到十亿参数模型,支持跨模型比较。

尽管大规模神经网络取得显著成功,我们仍缺乏统一的符号体系来思考和描述其表示空间。目前尚无可靠方法揭示其表示结构如何形成、如何随训练演化,以及何种结构具有优势。本文提出定量方法识别映射中的系统性结构,并利用这些方法理解深度学习模型如何表示信息、哪些表征结构驱动泛化,以及设计选择如何影响结构生成。通过识别映射中的结构原语及其信息论度量,可分析多智能体强化学习模型、单任务序列到序列模型及大型语言模型的学习、结构与泛化过程。还提出一种新型高效向量空间熵估计方法,使该分析可应用于从100万到120亿参数的各类模型。实验揭示了大规模分布式认知模型的学习机制,并揭示其与人类认知系统的相似结构,表明语言结构及其约束在许多方面与当代神经网络的性能结构相呼应。

原文摘要 · Abstract (English)

Despite the remarkable success of large large-scale neural networks, we still lack unified notation for thinking about and describing their representational spaces. We lack methods to reliably describe how their representations are structured, how that structure emerges over training, and what kinds of structures are desirable. This thesis introduces quantitative methods for identifying systematic structure in a mapping between spaces, and leverages them to understand how deep-learning models learn to represent information, what representational structures drive generalisation, and how design decisions condition the structures that emerge. To do this I identify structural primitives present in a mapping, along with information theoretic quantifications of each. These allow us to analyse learning, structure, and generalisation across multi-agent reinforcement learning models, sequence-to-sequence models trained on a single task, and Large Language Models. I also introduce a novel, performant, approach to estimating the entropy of vector space, that allows this analysis to be applied to models ranging in size from 1 million to 12 billion parameters. The experiments here work to shed light on how large-scale distributed models of cognition learn, while allowing us to draw parallels between those systems and their human analogs. They show how the structures of language and the constraints that give rise to them in many ways parallel the kinds of structures that drive performance of contemporary neural networks.

表示学习信息论大模型泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。