arXiv:2505.22112cs.AIq-bio.NC2025-05被引 4

VLLM在卡片分类测试中表现堪比人类,展现高级认知灵活性。

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test

  • 用思维链提示让大模型完成卡片分类任务,模拟人类思维转换。
  • GPT-4o等模型在文本输入下达到或超过人类的换牌能力。
  • 通过角色扮演可模拟脑损伤症状,揭示其类脑认知结构。

认知灵活性是人类认知的重要组成部分,但在视觉大语言模型(VLLMs)中的研究仍不充分。本研究采用经典的威斯康星卡片分类测试(WCST)评估前沿VLLM(GPT-4o、Gemini-1.5 Pro、Claude-3.5 Sonnet)的认知灵活性。结果表明,在使用思维链(chain-of-thought)提示且输入为文本时,VLLMs的换牌能力达到甚至超过人类水平。然而,其表现高度依赖输入模态与提示策略。此外,通过角色扮演,VLLMs能模拟多种与认知灵活性受损相关的功能缺陷,暗示其在换牌能力方面可能具备类似大脑的认知架构。本研究揭示了当前VLLM已在人类高阶认知的关键成分上接近人类水平,并展示了其作为复杂脑过程模拟工具的潜力。

原文摘要 · Abstract (English)

Cognitive flexibility has been extensively studied in human cognition but remains relatively unexplored in the context of Visual Large Language Models (VLLMs). This study assesses the cognitive flexibility of state-of-the-art VLLMs (GPT-4o, Gemini-1.5 Pro, and Claude-3.5 Sonnet) using the Wisconsin Card Sorting Test (WCST), a classic measure of set-shifting ability. Our results reveal that VLLMs achieve or surpass human-level set-shifting capabilities under chain-of-thought prompting with text-based inputs. However, their abilities are highly influenced by both input modality and prompting strategy. In addition, we find that through role-playing, VLLMs can simulate various functional deficits aligned with patients having impairments in cognitive flexibility, suggesting that VLLMs may possess a cognitive architecture, at least regarding the ability of set-shifting, similar to the brain. This study reveals the fact that VLLMs have already approached the human level on a key component underlying our higher cognition, and highlights the potential to use them to emulate complex brain processes.

视觉大模型认知灵活性类脑智能思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。