arXiv:2603.23543physics.soc-phcs.AI2026-03被引 3

LLMs lack human-like scientific intelligence due缺失隐性知识与社会对话

Large Language Models and Scientific Discourse: Where's the Intelligence?

  • 对比人类科学认知形成机制,指出LLMs依赖书面文献而无法获取专家间隐性交流
  • 在蒙提霍尔问题上,早期LLMs失败,但后续进步源于人类书写内容的更新而非模型自身推理增强
  • 当主流话语主导时,LLMs会忽略微小提示变化,暴露其缺乏真正理解力

我们通过对比人类构建知识的方式与大语言模型(LLMs)的数据获取方式,探讨其能力边界。以引力波物理领域2014年研究为例,科学家忽略‘边缘科学’论文主要依赖专家圈内非正式交流中积累的默会知识。而当前LLMs无法接触此类话语,导致其在科学知识初始阶段的理解不稳。2023年ChatGPT未能解决科林·弗雷泽提出的‘愚笨蒙提霍尔问题’,但一年后部分模型成功——这并非模型推理能力提升,而是人类书写内容的变化或人工干预所致。我们设计新蒙提霍尔提示,比较了多组LLMs与人类的回答,结果差异显著;但未来随着人类话语演变,LLMs将重新对齐人类认知。最后,我们分析‘遮蔽’现象:当既定话语体系占据主导,即使提示微调使原答案失效,LLMs仍机械沿用旧回答。因此,真正的智能仍在人类身上。

原文摘要 · Abstract (English)

We explore the capabilities of Large Language Models (LLMs) by comparing the way they gather data with the way humans build knowledge. Here we examine how scientific knowledge is made and compare it with LLMs. The argument is structured by reference to two figures, one representing scientific knowledge and the other LLMs. In a 2014 study, scientists explain how they choose to ignore a 'fringe science' paper in the domain of gravitational wave physics: the decisions are made largely as a result of tacit knowledge built up in social discourse, most spoken discourse, within closed groups of experts. It is argued that LLMs cannot or do not currently access such discourse, but it is typical of the early formation of scientific knowledge. LLMs 'understanding' builds on written literatures and is therefore insecure in the case of the initial stages of knowledge building. We refer to Colin Fraser's 'Dumb Monty Hall problem' where in 2023 ChatGPT failed though a year later or so later LLMs were succeeding. We argue that this is not a matter of improvement in LLMs ability to reason but in the change in the body of human written discourse on which they can draw (or changes being put in by humans 'by hand'). We then invent a new Monty Hall prompt and compare the responses of a panel of LLMs and a panel of humans: they are starkly different but we explain that the previous mechanisms will soon allow the LLMs to align themselves to humans once more. Finally, we look at 'overshadowing' where a settled body of discourse becomes so dominant that LLMs fail to respond to small variations in prompts which render the old answers nonsensical. The 'intelligence' we argue is in the humans not the LLMs

大模型局限科学认知隐性知识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。