arXiv:2505.01557cs.LGcs.AI2025-05ICML被引 2

提出上下文表征理论,解释大模型如何从输入与上下文关联中学习有效表示。

Contextures: Representations from Contexts

  • 将多种表示学习方法统一为从输入与上下文关联中学习
  • 证明学得的表征在适配任务上最优,且模型过大后提升有限
  • 提出无需下游任务即可评估上下文质量的新指标

尽管基础模型表现出色,但我们缺乏对其学习表征的系统性刻画。本文建立上下文表征理论(contexture theory),表明一大类表示学习方法可被理解为从输入与上下文变量的关联中学习。具体而言,许多主流方法旨在逼近由上下文诱导的期望算子的前d个奇异函数,此时称表征学习了上下文表征。我们通过证明,在监督、自监督和流形学习等多种范式下,表示学习均可从该视角分析。还证明,学习上下文表征的表示在与上下文兼容的任务上是最优的。一个重要启示是:一旦模型足够大以逼近前导奇异函数,继续扩大模型规模将带来边际收益递减。因此,仅靠扩展模型无法持续提升性能,需改进上下文。为此,我们研究了在不知下游任务时如何评估上下文有效性,提出了一个新度量,并通过实验验证其与编码器在多个真实数据集上的实际表现高度相关。

原文摘要 · Abstract (English)

Despite the empirical success of foundation models, we do not have a systematic characterization of the representations that these models learn. In this paper, we establish the contexture theory. It shows that a large class of representation learning methods can be characterized as learning from the association between the input and a context variable. Specifically, we show that many popular methods aim to approximate the top-d singular functions of the expectation operator induced by the context, in which case we say that the representation learns the contexture. We demonstrate the generality of the contexture theory by proving that representation learning within various learning paradigms -- supervised, self-supervised, and manifold learning -- can all be studied from such a perspective. We also prove that the representations that learn the contexture are optimal on those tasks that are compatible with the context. One important implication of the contexture theory is that once the model is large enough to approximate the top singular functions, further scaling up the model size yields diminishing returns. Therefore, scaling is not all we need, and further improvement requires better contexts. To this end, we study how to evaluate the usefulness of a context without knowing the downstream tasks. We propose a metric and show by experiments that it correlates well with the actual performance of the encoder on many real datasets.

表示学习基础模型上下文表征理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。