arXiv:2602.04872stat.MLcs.AI2026-02被引 3

证明多层交叉注意力可实现多模态上下文学习的最优性能

Multi-layer Cross-attention is Provably Optimal for Multi-modal In-context Learning

  • 构建多模态潜在因子模型,分析Transformer在上下文学习中的表达能力
  • 单层线性自注意力无法统一达到贝叶斯最优,而多层交叉注意力在梯度流优化下可实现最优
  • 揭示深度与交叉注意力对多模态学习的关键作用,适合研究模型机制的研究者

近期进展加速了我们对基于注意力神经网络中上下文学习机制的理解。然而,现有成果仅聚焦于单模态数据;相比之下,多模态上下文学习的理论基础仍不清晰。本文提出一个数学上可处理的框架来研究多模态学习,并探讨Transformer类架构何时能实现上下文中的贝叶斯最优性能。为建模多模态问题,假设观测数据来自潜在因子模型。首个结果表明:在任务分布上,单层线性自注意力无法统一恢复贝叶斯最优预测器。为克服此限制,我们引入一种新型线性化交叉注意力机制,并在交叉注意力层数和上下文长度均较大的情形下进行研究。结果表明,该机制在使用梯度流优化时可被严格证明达到贝叶斯最优。研究凸显了深度在上下文学习中的优势,并确立了交叉注意力对多模态分布的可证明效用。

原文摘要 · Abstract (English)

Recent progress has rapidly advanced our understanding of the mechanisms underlying in-context learning in modern attention-based neural networks. However, existing results focus exclusively on unimodal data; in contrast, the theoretical underpinnings of in-context learning for multi-modal data remain poorly understood. We introduce a mathematically tractable framework for studying multi-modal learning and explore when transformer-like architectures can recover Bayes-optimal performance in-context. To model multi-modal problems, we assume the observed data arises from a latent factor model. Our first result comprises a negative take on expressibility: we prove that single-layer, linear self-attention fails to recover the Bayes-optimal predictor uniformly over the task distribution. To address this limitation, we introduce a novel, linearized cross-attention mechanism, which we study in the regime where both the number of cross-attention layers and the context length are large. We show that this cross-attention mechanism is provably Bayes optimal when optimized using gradient flow. Our results underscore the benefits of depth for in-context learning and establish the provable utility of cross-attention for multi-modal distributions.

多模态注意力机制上下文学习理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。