arXiv:2411.02199cs.LGstat.ML2024-11NeurIPS被引 9

揭示了大模型如何利用多概念语义实现高效上下文学习。

Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning

  • 基于概念编码的稀疏提示模型,分析变压器如何利用多概念语义。
  • 证明在非凸训练下损失呈指数级收敛,支持强泛化能力。
  • 为理解大模型创新解题机制提供理论依据,适合研究者参考。

基于Transformer的大语言模型展现出卓越的创造力和涌现能力。现有研究表明,这些模型的强大涌现能力与其上下文学习(ICL)能力密切相关,即仅通过任务相关的提示即可解决新任务而无需微调。同时,已有实证与理论研究发现,这类模型中多概念语义表示具有线性规律。然而,现有理论未能建立该规律与ICL创新能力之间的联系。此外,以往工作常聚焦于线性变压器或不现实的损失函数,且仅能达到线性或次线性收敛速度。本文通过精细的数学分析,揭示了变压器如何利用词语的多概念语义实现强大的上下文学习及出色的分布外泛化能力,深入解析了其对包含多跨概念语义的未见任务的创新求解机制。受大模型线性潜在几何的实证启发,分析基于一种概念基础的低噪声稀疏编码提示模型,借助先进技术,首次在包含softmax自注意力、ReLU激活的MLP和交叉熵损失的复杂设置下,证明了0-1损失的指数级收敛,验证了理论结果的正确性。

原文摘要 · Abstract (English)

Transformer-based large language models (LLMs) have displayed remarkable creative prowess and emergence capabilities. Existing empirical studies have revealed a strong connection between these LLMs' impressive emergence abilities and their in-context learning (ICL) capacity, allowing them to solve new tasks using only task-specific prompts without further fine-tuning. On the other hand, existing empirical and theoretical studies also show that there is a linear regularity of the multi-concept encoded semantic representation behind transformer-based LLMs. However, existing theoretical work fail to build up an understanding of the connection between this regularity and the innovative power of ICL. Additionally, prior work often focuses on simplified, unrealistic scenarios involving linear transformers or unrealistic loss functions, and they achieve only linear or sub-linear convergence rates. In contrast, this work provides a fine-grained mathematical analysis to show how transformers leverage the multi-concept semantics of words to enable powerful ICL and excellent out-of-distribution ICL abilities, offering insights into how transformers innovate solutions for certain unseen tasks encoded with multiple cross-concept semantics. Inspired by empirical studies on the linear latent geometry of LLMs, the analysis is based on a concept-based low-noise sparse coding prompt model. Leveraging advanced techniques, this work showcases the exponential 0-1 loss convergence over the highly non-convex training dynamics, which pioneeringly incorporates the challenges of softmax self-attention, ReLU-activated MLPs, and cross-entropy loss. Empirical simulations corroborate the theoretical findings.

大模型上下文学习语义表征理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。