揭示了用小教师模型训练Transformer的样本复杂度规律
From Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASP
- 基于C-RASP构造,研究Transformer学习小教师模型的理论边界
- 首次给出Transformer学习此类构造所需的样本量下限
- 为理解大模型泛化能力提供新视角,适合关注理论机制的研究者
对Transformer的理论理解对于把握大语言模型的能力与局限至关重要。已有大量工作分析注意力模型的表达能力,通过设计特定权重或使用计算复杂性论证,试图界定哪些任务属于Transformer的假设类。然而,很少有研究探讨这些解的可学习性。本文在此方向取得进展:受近期损失景观分析启发,我们提出了学习C-RASP构造的初步样本复杂度界。
原文摘要 · Abstract (English)
A theoretical understanding of Transformers is crucial to better understand the capacities and limitations of large language models (LLMs). There is much work analyzing the expressivity of attention-based models. By proposing handcrafted weights or using computational complexity arguments, a large amount of past theoretical works have sought to characterize which tasks are and which are not in the hypothesis class of Transformer models. However, little work investigates the learnability of such solutions. In this work, we make progress towards this goal. Inspired by recent loss landscape analysis work, we propose preliminary sample complexity bounds for learning C-RASP constructions with Transformers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。