arXiv:2502.16606cs.CLcs.AI2025-02被引 1

测试大模型能否像人一样用奇怪物品完成任务,发现新版模型表现接近人类。

Reasoning about Affordances: Causal and Compositional Reasoning in LLMs

  • 设计新任务让模型选非常规工具完成目标,避免数据泄露
  • GPT-4o和Claude 3.5在复杂条件下仍能正确推理,准确率远超随机
  • 视觉输入对部分模型影响大,说明其理解方式仍有差异

随着大语言模型(LLMs)的快速发展,理解其能力与局限变得愈发重要。本文通过两项实验,研究了大模型与人类在物体功能属性(affordances)领域中的因果与组合推理能力,该领域传统上与具身认知相关。任务从零设计以避免数据污染,要求决策者选择非典型物品替代常规工具完成特定目的,例如用乒乓球拍挖洞。实验1评估GPT-3.5与GPT-4o,发现经思维链提示后,GPT-4o表现与人类相当,而GPT-3.5明显落后。实验2引入干扰项(更多选项,增加难度)和图像呈现(视觉展示物品)两种新条件,并额外测试Claude 3 Sonnet与Claude 3.5 Sonnet。干扰项显著降低人类及所有模型的表现,但GPT-4o与Claude 3.5仍显著高于随机水平。令人意外的是,图像条件对人类和GPT-4o影响甚微,却显著降低Claude 3.5的准确率。定性分析表明,GPT-4o与Claude 3.5相比前代模型,在识别并灵活运用因果相关属性方面更具优势。从GPT-3.5到GPT-4o、Claude 3到Claude 3.5的提升,表明这些模型在某些领域已具备更强的因果与组合推理能力,但其内在机制仍需进一步研究。

原文摘要 · Abstract (English)

With the rapid progress of Large Language Models (LLMs), it becomes increasingly important to understand their abilities and limitations. In two experiments, we investigate the causal and compositional reasoning abilities of LLMs and humans in the domain of object affordances, an area traditionally linked to embodied cognition. The tasks, designed from scratch to avoid data contamination, require decision-makers to select unconventional objects to replace a typical tool for a particular purpose, such as using a table tennis racket to dig a hole. In Experiment 1, we evaluated GPT-3.5 and GPT-4o, finding that GPT-4o, when given chain-of-thought prompting, performed on par with human participants, while GPT-3.5 lagged significantly. In Experiment 2, we introduced two new conditions, Distractor (more object choices, increasing difficulty) and Image (object options presented visually), and evaluated Claude 3 Sonnet and Claude 3.5 Sonnet in addition to the GPT models. The Distractor condition significantly impaired performance across humans and models, although GPT-4o and Claude 3.5 still performed well above chance. Surprisingly, the Image condition had little impact on humans or GPT-4o, but significantly lowered Claude 3.5's accuracy. Qualitative analysis showed that GPT-4o and Claude 3.5 have a stronger ability than their predecessors to identify and flexibly apply causally relevant object properties. The improvement from GPT-3.5 and Claude 3 to GPT-4o and Claude 3.5 suggests that models are increasingly capable of causal and compositional reasoning in some domains, although further mechanistic research is necessary to understand how LLMs reason.

因果推理大模型组合推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。