大模型能仅靠语言生成比人类更可靠的视觉想象。
Artificial Phantasia: Emergent Mental Imagery in Large Language Models
- 用语言描述字母形状变换,测试模型想象能力。
- 最佳模型表现显著优于100名人类参与者(p<0.0001)。
- 适合对认知科学与AI想象力交叉研究感兴趣者。
视觉意象能否仅由语言驱动?传统认知科学认为视觉心理意象必须依赖图像表征。大型语言模型(LLMs)提供了初步证据,表明仅通过命题表征即可实现视觉心理意象,且可能比人类想象更稳定。我们设计了数十个新任务项,扩展了一项经典任务——该任务传统上被认为只能通过图像表征解决(仅语言不足以完成)。受试者需想象一系列组合式字母与形状变换,并识别结果“图像”。结果显示,最优的LLM在该任务上的表现显著优于100名人类参与者(p < .0001),表明存在一种人工幻象(artificial phantasia),即非图像形式的新兴“视觉”心理意象。此外,我们测试了不同推理令牌分配下的推理模型,发现较长的推理链表现最佳,说明语言本身对任务有显著影响——语言单独可能已足够。我们检验了三种新兴意象假设:纯命题意象、含视觉-语言先验的命题意象,或经典图像式视觉意象。本研究不仅揭示了LLM一种此前未报告的涌现认知能力,也重新引发了关于心理意象是否必须以图像格式存在的学术争论。
原文摘要 · Abstract (English)
Can visual imagery be driven solely by language? This idea goes against cognitive science's traditional view that visual mental imagery is only possible through pictorial representations. Large Language Models (LLMs) provide nascent evidence not only that visual mental imagery via propositional-representations is possible, but that it can be more robust than human imagination. We created dozens of novel items for an extension to a classic task which is argued to be solvable exclusively via pictorial representations (i.e., language alone would be insufficient). Subjects were asked to imagine a series of compositional letter and shape transformations and identify the resultant "image". We found that the best LLMs performed significantly better than humans ($n = 100$ human participants, $p < .0001$), indicating the existence of an artificial phantasia, or emergent "visual" mental imagery that may not be pictorial. Furthermore, we tested reasoning models with variable reasoning-token allocation and found that models perform best with longer reasoning chains, demonstrating a linguistic impact on the task -- language alone may be sufficient. We examined three emergent imagery hypotheses: pure propositional imagery, propositional imagery with visio-linguistic priors, or pictorial visual imagery (classical visual imagery). Our study not only presents evidence for a previously unreported emergent cognitive capacity of LLMs, but also reignites debate on the requirement for a pictorial format in mental imagery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。