探究大模型能否通过结构对应关系表征真实世界内容
Can structural correspondences ground real world representational content in Large Language Models?
- 基于结构对应理论分析大模型的表征能力
- 仅存在结构对应不足以支撑真实表征
- 需任务中有效利用对应关系才可能实现表征
大型语言模型(如GPT-4)对多种提示生成令人信服的响应,但其表征能力尚不明确。许多大模型与语言外现实无直接接触:输入、输出及训练数据均仅由文本构成,这引发两个问题:(1) 大模型能否表征任何事物?(2) 若能,表征的是什么?本文根据基于结构对应的关系表征理论,探讨回答这些问题所需条件,并初步调查相关证据。论证指出,仅存在大模型与现实实体之间的结构对应关系,不足以奠基这些实体的真实表征。然而,若这些结构对应关系在任务中发挥适当作用——即被有效利用以解释成功的表现——则可能奠基真实世界的表征内容。这需要克服一个挑战:大模型的文本封闭性表面上似乎阻止其参与恰当的任务。
原文摘要 · Abstract (English)
Large Language Models (LLMs) such as GPT-4 produce compelling responses to a wide range of prompts. But their representational capacities are uncertain. Many LLMs have no direct contact with extra-linguistic reality: their inputs, outputs and training data consist solely of text, raising the questions (1) can LLMs represent anything and (2) if so, what? In this paper, I explore what it would take to answer these questions according to a structural-correspondence based account of representation, and make an initial survey of this evidence. I argue that the mere existence of structural correspondences between LLMs and worldly entities is insufficient to ground representation of those entities. However, if these structural correspondences play an appropriate role - they are exploited in a way that explains successful task performance - then they could ground real world contents. This requires overcoming a challenge: the text-boundedness of LLMs appears, on the face of it, to prevent them engaging in the right sorts of tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。