用网格任务测试大模型对物理概念的理解能力,发现其远不如人类。
The Stochastic Parrot on LLM's Shoulder: A Summative Assessment of Physical Concept Understanding
- 设计网格格式任务,避免记忆干扰,考察深层理解。
- 顶尖模型在任务中表现比人类差约40%。
- 适合关注大模型认知局限的研究者和开发者。
我们系统性地探讨了一个常见问题:大语言模型(LLMs)是否真正理解它们所说的内容?为此,我们提出了一项针对物理概念理解的综合性评估任务 PhysiCo。该任务通过网格格式输入抽象描述物理现象,涵盖核心现象、应用实例及与其他抽象模式的类比。实验表明:(1)包括 GPT-4o、o1 和 Gemini 2.0 flash thinking 在内的顶尖模型,在任务中表现比人类落后约 40%;(2)模型在自然语言中能描述和识别相同概念,但在网格任务中失败,印证了‘随机鹦鹉’现象;(3)任务挑战源于内在难度,而非格式陌生,因为在相同格式数据上进行上下文学习或微调,性能提升有限。
原文摘要 · Abstract (English)
In a systematic way, we investigate a widely asked question: Do LLMs really understand what they say?, which relates to the more familiar term Stochastic Parrot. To this end, we propose a summative assessment over a carefully designed physical concept understanding task, PhysiCo. Our task alleviates the memorization issue via the usage of grid-format inputs that abstractly describe physical phenomena. The grids represents varying levels of understanding, from the core phenomenon, application examples to analogies to other abstract patterns in the grid world. A comprehensive study on our task demonstrates: (1) state-of-the-art LLMs, including GPT-4o, o1 and Gemini 2.0 flash thinking, lag behind humans by ~40%; (2) the stochastic parrot phenomenon is present in LLMs, as they fail on our grid task but can describe and recognize the same concepts well in natural language; (3) our task challenges the LLMs due to intrinsic difficulties rather than the unfamiliar grid format, as in-context learning and fine-tuning on same formatted data added little to their performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。