arXiv:2505.17838stat.MLcs.LG2025-05被引 3

连续变换器通过算子梯度下降实现上下文学习,理论揭示其等价于贝叶斯最优预测。

Continuum Transformers Perform In-Context Learning by Operator Gradient Descent

  • 将标准Transformer推广到函数空间,构建连续变换器处理无限维输入。
  • 证明连续变换器在上下文中执行算子梯度下降,收敛于贝叶斯最优预测。
  • 适用于科学计算中的微分方程代理建模,适合关注理论机制的研究者。

Transformer在不更新参数的情况下,仅通过在上下文窗口中放置训练样本即可提升任务预测准确率,这种现象称为上下文学习。近期研究发现,标准Transformer通过前向传播实现梯度下降来达成此目标。然而,该结论局限于处理有限维输入的架构。在偏微分方程(PDE)代理建模领域,一种可处理无限维函数输入的广义Transformer——连续变换器已被提出,并展现出类似的上下文学习能力。尽管实证表现优异,但其理论机制尚不明确。本文首次证明:连续变换器通过在算子再生核希尔伯特空间(operator RKHS)中执行梯度下降,实现上下文算子学习。我们采用新型证明策略,基于希尔伯特空间的广义表示定理及泛函空间上的梯度流分析,严格建立该结论。进一步证明,在变换器深度趋于无穷时,上下文中学习到的算子即为贝叶斯最优预测器。最后,我们通过实验验证了该最优性结果,并展示了连续变换器训练过程中可恢复出执行梯度下降的参数配置。

原文摘要 · Abstract (English)

Transformers robustly exhibit the ability to perform in-context learning, whereby their predictive accuracy on a task can increase not by parameter updates but merely with the placement of training samples in their context windows. Recent works have shown that transformers achieve this by implementing gradient descent in their forward passes. Such results, however, are restricted to standard transformer architectures, which handle finite-dimensional inputs. In the space of PDE surrogate modeling, a generalization of transformers to handle infinite-dimensional function inputs, known as "continuum transformers," has been proposed and similarly observed to exhibit in-context learning. Despite impressive empirical performance, such in-context learning has yet to be theoretically characterized. We herein demonstrate that continuum transformers perform in-context operator learning by performing gradient descent in an operator RKHS. We demonstrate this using novel proof strategies that leverage a generalized representer theorem for Hilbert spaces and gradient flows over the space of functionals of a Hilbert space. We additionally show the operator learned in context is the Bayes Optimal Predictor in the infinite depth limit of the transformer. We then provide empirical validations of this optimality result and demonstrate that the parameters under which such gradient descent is performed are recovered through the continuum transformer training.

连续变换器算子学习贝叶斯最优上下文学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。