arXiv:2502.03503stat.MLcs.AI2025-02被引 5
揭示大模型上下文学习的理论极限,挑战其类学习算法的假说。
Analyzing limits for in-context learning
- 通过数学分析证明变压器架构固有缺陷
- 实验证明其无法实现通用预测准确率
- 适合关注AI底层机制的研究者阅读
本文质疑先前研究中关于基于Transformer的模型在上下文学习时隐式实现标准学习算法的说法。我们提供了与该观点不一致的实证证据,并进行了数学分析,表明由于固有的架构限制,Transformer无法实现通用预测准确性。研究结果揭示了当前大模型在上下文学习中的根本性局限。
原文摘要 · Abstract (English)
Our paper challenges claims from prior research that transformer-based models, when learning in context, implicitly implement standard learning algorithms. We present empirical evidence inconsistent with this view and provide a mathematical analysis demonstrating that transformers cannot achieve general predictive accuracy due to inherent architectural limitations.
上下文学习Transformer理论分析
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。