大模型预测不确定性难降,根源在于其生成非高斯输出的机制。
The wall confronting large language models
- 利用高斯输入生成非高斯输出的机制推动学习,却导致误差堆积。
- 数据规模增大时,虚假相关性激增,加剧模型不可靠性。
- 需更强问题理解力来避免大模型走向退化,适合关注可信AI的研究者。
我们发现,决定大语言模型性能的缩放定律严重限制了其降低预测不确定性的能力。因此,通过合理手段提升其可靠性以满足科学探究标准几乎不可能。我们认为,驱动大模型强大学习能力的核心机制——即从高斯输入生成非高斯输出的能力——很可能正是其产生误差堆积、引发信息灾难和退化人工智能行为的根源。这种学习与准确性的矛盾,可能是观测到的缩放系数偏低的潜在机制。该矛盾因Calude和Longo指出的虚假相关性急剧增加而进一步加剧,这类相关性随数据集规模增大而迅速出现,无论数据性质如何。尽管退化人工智能路径在大模型景观中极为可能,但并不意味着必然发生。本文还探讨了规避此路径的方法,强调必须更高程度重视对问题结构特征的洞察与理解。
原文摘要 · Abstract (English)
We show that the scaling laws which determine the performance of large language models (LLMs) severely limit their ability to improve the uncertainty of their predictions. As a result, raising their reliability to meet the standards of scientific inquiry is intractable by any reasonable measure. We argue that the very mechanism which fuels much of the learning power of LLMs, namely the ability to generate non-Gaussian output distributions from Gaussian input ones, might well be at the roots of their propensity to produce error pileup, ensuing information catastrophes and degenerative AI behaviour. This tension between learning and accuracy is a likely candidate mechanism underlying the observed low values of the scaling components. It is substantially compounded by the deluge of spurious correlations pointed out by Calude and Longo which rapidly increase in any data set merely as a function of its size, regardless of its nature. The fact that a degenerative AI pathway is a very probable feature of the LLM landscape does not mean that it must inevitably arise in all future AI research. Its avoidance, which we also discuss in this paper, necessitates putting a much higher premium on insight and understanding of the structural characteristics of the problems being investigated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。