让大模型在3D环境中动手操作,测试其物理常识理解能力。
A little less conversation, a little more action, please: Investigating the physical common-sense of LLMs in a 3D embodied environment
- 将大模型置于3D虚拟环境,通过控制代理直接执行任务。
- 模型在物体追踪、距离判断等任务上表现不及人类儿童。
- 方法基于认知科学实验,适合评估真实物理推理能力。
大型语言模型(LLMs)需在日常物理环境中进行推理。现有研究多依赖静态文本或图像基准,但难以捕捉真实物理过程的复杂性。本文提出将LLM‘具身化’,赋予其在3D环境中的代理控制权,首次实现对LLM物理常识推理的具身化评估。采用Animal-AI(AAI)环境及AAI Testbed实验套件,复现非人动物实验室研究,考察距离估计、物体追踪和工具使用等能力。结果表明,未经微调的多模态模型可完成此类任务,表现接近2019年动物AI奥林匹克竞赛参赛者,但仍逊于人类儿童。该方法借鉴认知科学生态有效实验,提升对LLM推理能力预测与评估的可靠性。
原文摘要 · Abstract (English)
As general-purpose tools, Large Language Models (LLMs) must often reason about everyday physical environments. In a question-and-answer capacity, understanding the interactions of physical objects may be necessary to give appropriate responses. Moreover, LLMs are increasingly used as reasoning engines in agentic systems, designing and controlling their action sequences. The vast majority of research has tackled this issue using static benchmarks, comprised of text or image-based questions about the physical world. However, these benchmarks do not capture the complexity and nuance of real-life physical processes. Here we advocate for a second, relatively unexplored, approach: 'embodying' the LLMs by granting them control of an agent within a 3D environment. We present the first embodied and cognitively meaningful evaluation of physical common-sense reasoning in LLMs. Our framework allows direct comparison of LLMs with other embodied agents, such as those based on Deep Reinforcement Learning, and human and non-human animals. We employ the Animal-AI (AAI) environment, a simulated 3D virtual laboratory, to study physical common-sense reasoning in LLMs. For this, we use the AAI Testbed, a suite of experiments that replicate laboratory studies with non-human animals, to study physical reasoning capabilities including distance estimation, tracking out-of-sight objects, and tool use. We demonstrate that state-of-the-art multi-modal models with no finetuning can complete this style of task, allowing meaningful comparison to the entrants of the 2019 Animal-AI Olympics competition and to human children. Our results show that LLMs are currently outperformed by human children on these tasks. We argue that this approach allows the study of physical reasoning using ecologically valid experiments drawn directly from cognitive science, improving the predictability and reliability of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。