arXiv:2504.19792cs.LGcs.AI2025-04

提出统一理论解释预训练为何有效,关键在输入与上下文的关联机制。

Contextures: The Mechanism of Representation Learning

  • 用输入与上下文变量的关联定义表征学习机制
  • 证明最优表征来自最大捕捉该关联的编码器
  • 强调设计更好上下文比单纯扩大模型更关键

本论文建立上下文结构(contexture)理论,数学化刻画表征学习(或预训练)的机制。尽管基础模型在实践中表现优异,但其学到的表征本质及为何适用于下游任务尚不清晰。科学理解表征学习对当前模型规模扩展趋于饱和的阶段尤为关键。以往研究对不同学习方法处理各异,而上下文结构理论提供统一分析框架。核心观点是:表征从输入X与上下文变量A的关联中学习。若编码器捕获该关联的最大信息量(即学习到上下文结构),则在与该上下文兼容的任务类别上达到最优。同时表明,当X与A的关联既不太强也不太弱时,上下文最有效。重要启示是仅增大模型规模将导致收益递减,进一步突破需更好上下文。我们证明多种预训练目标(如监督学习、自监督学习、生成模型等)均可学习上下文结构。进而提出两种通用目标SVME与KISE用于学习上下文结构,并展示如何混合多个上下文以生成更优上下文。最后,推导了表征学习的统计学习界,并讨论预训练与下游任务间数据分布偏移的影响。

原文摘要 · Abstract (English)

This dissertation establishes the contexture theory to mathematically characterize the mechanism of representation learning, or pretraining. Despite the remarkable empirical success of foundation models, it is not very clear what representations they learn, and why these representations are useful for various downstream tasks. A scientific understanding of representation learning is critical, especially at this point when scaling up the model size is producing diminishing returns, and designing new pretraining methods is imperative for further progress. Prior work treated different representation learning methods quite differently, whereas the contexture theory provides a unified framework for analyzing these methods. The central argument is that a representation is learned from the association between the input X and a context variable A. We prove that if an encoder captures the maximum information of this association, in which case we say that the encoder learns the contexture, then it will be optimal on the class of tasks that are compatible with the context. We also show that a context is the most useful when the association between X and A is neither too strong nor too weak. The important implication of the contexture theory is that increasing the model size alone will achieve diminishing returns, and further advancements require better contexts. We demonstrate that many pretraining objectives can learn the contexture, including supervised learning, self-supervised learning, generative models, etc. Then, we introduce two general objectives -- SVME and KISE, for learning the contexture. We also show how to mix multiple contexts together, an effortless way to create better contexts from existing ones. Then, we prove statistical learning bounds for representation learning. Finally, we discuss the effect of the data distribution shift from pretraining to the downstream task.

表征学习预训练理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。