arXiv:2409.12293cs.LGcs.NA2024-09被引 5

用线性Transformer解线性方程组,给出泛化误差的理论边界。

In-Context Learning of Linear Systems: Generalization Theory and Applications to Operator Learning

  • 基于线性Transformer设计上下文学习方法求解线性系统。
  • 发现任务多样性是模型跨分布泛化的关键条件。
  • 适用于微分方程算子学习等实际场景,理论可验证。

我们研究了使用线性Transformer架构在上下文学习中求解线性系统的理论保证。对于域内泛化,给出了神经尺度定律,将泛化误差与训练和推理中使用的任务数量及样本规模关联起来。对于域外泛化,发现训练后Transformer在任务分布变化下的表现关键取决于训练时遇到的任务分布。我们引入新的任务多样性概念,并证明其是预训练Transformer在任务分布偏移下实现泛化的充要条件。此外,还探索了在线性系统上下文学习中的应用,如偏微分方程(PDE)的算子学习。最后,通过数值实验验证了所提出的理论。

原文摘要 · Abstract (English)

We study theoretical guarantees for solving linear systems in-context using a linear transformer architecture. For in-domain generalization, we provide neural scaling laws that bound the generalization error in terms of the number of tasks and sizes of samples used in training and inference. For out-of-domain generalization, we find that the behavior of trained transformers under task distribution shifts depends crucially on the distribution of the tasks seen during training. We introduce a novel notion of task diversity and show that it defines a necessary and sufficient condition for pre-trained transformers generalize under task distribution shifts. We also explore applications of learning linear systems in-context, such as to in-context operator learning for PDEs. Finally, we provide some numerical experiments to validate the established theory.

线性系统上下文学习泛化理论算子学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。