arXiv:2501.10361cs.CYcs.CL2025-01

把大模型看作外推机器,解释其生成与幻觉的本质

How Large Language Models (LLMs) Extrapolate: From Guided Missiles to Guided Prompts

  • 将大模型视为外推工具,而非单纯记忆或生成
  • 外推能力是模型高效响应的核心,过度外推导致幻觉
  • 追溯外推思想从导弹控制到现代AI的科学脉络

本文主张应将大语言模型视为外推机器。外推是一种用于预测序列中下一个值的统计函数,它既推动了GPT的成功,也引发了关于其幻觉现象的争议。所谓‘幻觉’暗示故障,但本文认为这实则反映了聊天机器人在信息外推上的高效性,只是存在过度现象。文章具有历史维度:追溯外推思想至控制论早期。1941年,诺伯特·维纳从导弹科学转向通信工程时,采纳的核心概念正是外推。苏联数学家安德烈·柯尔莫戈洛夫于1939年提出的另一项外推研究,后来被维纳发现与其构思高度相似。本文揭示了热战科学、冷战控制论与当代大模型性能争论之间的深层关联。

原文摘要 · Abstract (English)

This paper argues that we should perceive LLMs as machines of extrapolation. Extrapolation is a statistical function for predicting the next value in a series. Extrapolation contributes to both GPT successes and controversies surrounding its hallucination. The term hallucination implies a malfunction, yet this paper contends that it in fact indicates the chatbot efficiency in extrapolation, albeit an excess of it. This article bears a historical dimension: it traces extrapolation to the nascent years of cybernetics. In 1941, when Norbert Wiener transitioned from missile science to communication engineering, the pivotal concept he adopted was none other than extrapolation. Soviet mathematician Andrey Kolmogorov, renowned for his compression logic that inspired OpenAI, had developed in 1939 another extrapolation project that Wiener later found rather like his own. This paper uncovers the connections between hot war science, Cold War cybernetics, and the contemporary debates on LLM performances.

大模型外推幻觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。