视觉语言模型懂逆转但不懂数量守恒,认知能力有缺陷
Vision Language Models Know Law of Conservation without Understanding More-or-Less
- 设计4维度365题认知测试集,检验模型对守恒的理解
- 模型在可逆操作任务中表现良好,但在数量判断上普遍失败
- 揭示模型缺乏对数量本质理解,适合研究模型认知局限的人看
理解守恒定律是人类认知发展的关键里程碑,被认为依赖于数量概念和操作可逆性的掌握。为评估该核心智能是否在视觉语言模型中出现,我们构建了ConserveBench,一个涵盖体积、固体量、长度和数量四个物理维度的365个认知实验测试集。前两个维度涉及需要可逆性理解的变换任务,后两个涉及数量理解的非变换任务。令人惊讶的是,尽管视觉语言模型在变换任务上总体表现良好,却往往在非变换任务中失败。这表明模型在操作可逆性与数量概念理解之间存在分离,而这两者均被认为是人类理解守恒定律的基础。
原文摘要 · Abstract (English)
Understanding law of conservation is a critical milestone in human cognitive development considered to be supported by the apprehension of quantitative concepts and the reversibility of operations. To assess whether this critical component of human intelligence has emerged in Vision Language Models, we have curated the ConserveBench, a battery of 365 cognitive experiments across four dimensions of physical quantities: volume, solid quantity, length, and number. The former two involve transformational tasks which require reversibility understanding. The latter two involve non-transformational tasks which assess quantity understanding. Surprisingly, we find that while Vision Language Models are generally good at transformational tasks, they tend to fail at non-transformational tasks. There is a dissociation between understanding the reversibility of operations and understanding the concept of quantity, which both are believed to be the cornerstones of understanding law of conservation in humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。