随机语言任务暴露大模型在上下文学习中的根本缺陷
Randomly Sampled Language Reasoning Problems Elucidate Limitations of In-Context Learning
- 用随机生成的简单语言任务测试上下文学习能力
- 大模型表现差于n-gram模型,无论是否使用思维链
- 揭示了大模型依赖参数知识而非真正推理能力
尽管大语言模型在众多任务中表现出色,但其仍存在幻觉和对非标准任务表现不佳的问题。现有研究多关注复杂任务的微小扰动,而本文旨在通过极简语言任务分离出新颖性这一变量。我们设计了一种基于简单语法规则的随机语言集合,用于测试大模型在下一词预测任务上的表现。这些任务完全未见过,且不依赖模型的预训练知识。实验表明,无论作为直接预测器还是结合思维链使用,大模型的表现均劣于n-gram模型,说明其在上下文学习中的局限性并非源于复杂性,而是对新颖任务的适应能力不足。
原文摘要 · Abstract (English)
While LLMs have revolutionized the field of machine learning due to their high performance on a strikingly wide range of problems, they are also known to hallucinate false answers and underperform on less canonical versions of the same tasks. There are several emerging theories of LLM performance, among them that LLMs lack world modeling ability, that they have an undesirable bias towards an autoregressive prior, and that they struggle on more novel problems. The existing literature on LLM input novelty has focused on tasks of relatively high complexity, studying perturbations of canonical but complex problems. In this paper, we attempt to minimize complexity in order to isolate novelty as a factor in LLM underperformance and investigate the power of in-context-learning. To this end, we consider an extremely simple domain: next token prediction on simple language tasks. The twist is that these language tasks are wholly unseen, as they are randomly drawn from a large, parsimoniously defined set of languages arising from simple grammar rules. This experimental setup allows us to evaluate ICL independently of models' parametric knowledge. We find that LLMs uniformly underperform n-gram models on this task, both when used as next token predictors and in chain-of-thought.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。