arXiv:2506.10959cs.LGcs.AI2025-06被引 11

揭示了Transformer在流形上做上下文学习的几何机制。

Understanding In-Context Learning on Structured Manifolds: Bridging Attention to Kernel Methods

  • 将注意力机制与核方法关联,解释Transformer如何通过提示进行预测。
  • 证明了提示长度越长,误差下降越快,且依赖于流形内在维数。
  • 适合研究大模型理论、几何机器学习的研究者阅读。

尽管上下文学习(ICL)在自然语言和视觉领域取得显著成功,但其在结构化几何数据上的理论理解仍属空白。本文首次对流形上霍尔德函数的回归任务中的ICL进行理论研究。我们建立注意力机制与经典核方法的新联系,表明Transformer通过与提示的交互,实质上执行基于核的预测。数值实验验证了该观点:对于霍尔德函数,学习到的查询-提示得分与高斯核高度相关。基于此,我们推导出泛化误差界,其依赖于提示长度和训练任务数量。当训练任务足够多时,Transformer实现了流形上霍尔德函数的极小极大回归率,该速率随提示长度呈指数级下降,指数依赖于流形的内在维度而非嵌入空间维度。结果还刻画了泛化误差随训练任务数的变化规律,揭示了Transformer作为上下文核算法学习者的复杂性。研究为几何在ICL中的作用提供了基础洞见,并提出了分析非线性模型ICL的新工具。

原文摘要 · Abstract (English)

While in-context learning (ICL) has achieved remarkable success in natural language and vision domains, its theoretical understanding-particularly in the context of structured geometric data-remains unexplored. This paper initiates a theoretical study of ICL for regression of Hölder functions on manifolds. We establish a novel connection between the attention mechanism and classical kernel methods, demonstrating that transformers effectively perform kernel-based prediction at a new query through its interaction with the prompt. This connection is validated by numerical experiments, revealing that the learned query-prompt scores for Hölder functions are highly correlated with the Gaussian kernel. Building on this insight, we derive generalization error bounds in terms of the prompt length and the number of training tasks. When a sufficient number of training tasks are observed, transformers give rise to the minimax regression rate of Hölder functions on manifolds, which scales exponentially with respect to the prompt length with the exponent depending on the intrinsic dimension of the manifold, rather than the ambient space dimension. Our result also characterizes how the generalization error scales with the number of training tasks, shedding light on the complexity of transformers as in-context kernel algorithm learners. Our findings provide foundational insights into the role of geometry in ICL and novels tools to study ICL of nonlinear models.

上下文学习核方法流形学习Transformer理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。