arXiv:2409.13929cs.AI2024-09被引 6

用认知科学方法测试GPT-4o的视角转换能力,发现其空间理解与人类有本质差异。

Failures in Perspective-taking of Multimodal AI Systems

  • 引入认知科学方法评估GPT-4o的视角转换能力
  • 发现模型依赖命题表征,缺乏类人模拟表征
  • 为下一代多模态模型发展提供人类认知参照

本研究拓展了对多模态人工智能系统中空间表征的理解。尽管当前模型能从图像中提取丰富的空间信息,但这些信息基于命题表征,与人类及动物空间认知中的模拟表征存在本质差异。为深入探究此局限性,我们采用认知科学与发育科学中的方法,评估GPT-4o的视角转换能力。该分析实现了对人类大脑认知发展与多模态人工智能认知发展之间的对比,为未来研究和模型设计提供了指导。

原文摘要 · Abstract (English)

This study extends previous research on spatial representations in multimodal AI systems. Although current models demonstrate a rich understanding of spatial information from images, this information is rooted in propositional representations, which differ from the analog representations employed in human and animal spatial cognition. To further explore these limitations, we apply techniques from cognitive and developmental science to assess the perspective-taking abilities of GPT-4o. Our analysis enables a comparison between the cognitive development of the human brain and that of multimodal AI, offering guidance for future research and model development.

多模态AI认知建模视角转换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。