arXiv:2410.00324cs.AI2024-10被引 17

测试发现视觉语言模型懂意图却难换位思考。

Vision Language Models See What You Want but not What You See

  • 构建两个认知基准测试,模拟真实场景中的心智理论任务。
  • 模型在理解意图上表现良好,但第二层换位思考能力很差。
  • 提示现有模型缺乏基于模型的心理状态推理能力,适合研究者参考。

理解他人意图和换位思考是人类智能的核心,被视为心智理论的具体体现。为探究视觉语言模型(VLMs)在意图理解和二级视角推理方面的能力,我们构建了IntentBench和PerspectBench两个基准,共包含300多个基于现实场景和经典认知任务的实验。结果显示,VLMs在意图理解任务上表现优异,但在二级视角推理任务中表现较差。这表明VLMs可能存在基于模拟与基于理论的心智理论能力分离现象,提示其难以运用基于模型的推理来推断他人心理状态。

原文摘要 · Abstract (English)

Knowing others' intentions and taking others' perspectives are two core components of human intelligence that are considered to be instantiations of theory-of-mind. Infiltrating machines with these abilities is an important step towards building human-level artificial intelligence. Here, to investigate intentionality understanding and level-2 perspective-taking in Vision Language Models (VLMs), we constructed the IntentBench and PerspectBench, which together contains over 300 cognitive experiments grounded in real-world scenarios and classic cognitive tasks. We found VLMs achieving high performance on intentionality understanding but low performance on level-2 perspective-taking. This suggests a potential dissociation between simulation-based and theory-based theory-of-mind abilities in VLMs, highlighting the concern that they are not capable of using model-based reasoning to infer others' mental states.

心智理论视觉语言模型认知测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。