arXiv:2410.00318cs.AIq-bio.NC2024-10被引 13

测试26个视觉语言模型在机械推理上的表现,发现其远逊于人类。

Probing Mechanical Reasoning in Large Vision Language Models

  • 用155个认知实验评测模型对机械系统的理解能力
  • 齿轮和流体力学任务表现最差,且参数量越大越不改善
  • 提示当前模型缺乏心理模拟能力,难以掌握机械机制

机械推理是人类智能的标志,广泛应用于日常活动到土木工程。将机械推理能力赋予机器是迈向人类水平人工智能的重要一步。本文利用155项认知实验,测试26个视觉语言模型(VLMs)在系统稳定性、齿轮与滑轮系统、杠杆原理、惯性与运动、流体力学等方面的理解能力。结果表明,所有任务中模型表现均显著低于人类,尤其在齿轮系统和流体力学任务上困难重重。值得注意的是,模型性能并未随参数量增加而提升,暗示当前基于注意力的架构可能无法捕捉机械推理所需的深层机制,尤其是心理模拟能力。

原文摘要 · Abstract (English)

Mechanical reasoning is a hallmark of human intelligence, defined by its ubiquitous yet irreplaceable role in human activities ranging from routine tasks to civil engineering. Embedding machines with mechanical reasoning is therefore an important step towards building human-level artificial intelligence. Here, we leveraged 155 cognitive experiments to test the understanding of system stability, gears and pulley systems, leverage principle, inertia and motion, and fluid mechanics in 26 Vision Language Models (VLMs). Results indicate that VLMs consistently perform worse than humans on all domains, while demonstrate significant difficulty in reasoning about gear systems and fluid mechanics. Notably, their performance on these tasks do not improve as number of parameters increase, suggesting that current attention-based architecture may fail to grasp certain underlying mechanisms required for mechanical reasoning, particularly those pertaining to mental simulations.

机械推理视觉语言模型心智模拟认知实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。