arXiv:2505.00603cs.AIcs.HC2025-05被引 3

GPT-4能生成大量类比,但人类更擅长判断类比的深层逻辑。

Can LLMs Help Improve Analogical Reasoning For Strategic Decisions? Experimental Evidence from Humans and GPT-4

  • 用源域到目标域匹配法测试类比推理能力
  • GPT-4召回率高但误用率高,人类召回低但准确率高
  • 适合人机协作决策:AI出主意,人把关

本研究探究大型语言模型(特别是GPT-4)在战略决策情境下的类比推理能力是否可媲美人类。通过新颖的源域到目标域匹配实验设计,发现GPT-4在检索所有可能类比方面表现出高召回率,但精度低,常基于表面相似性错误应用类比。而人类参与者则表现出高精度但低召回率,所选类比虽少却具有更强的因果一致性。该发现推进了理论,揭示类比推理中的‘匹配’阶段是一个需要精确因果映射的独立步骤,超越简单检索。尽管当前大模型在生成候选类比方面表现优异,人类在识别跨领域深层结构相似性上仍具优势。错误分析显示,AI错误源于表面匹配,人类错误则源于因果结构误读。综合结果表明,在人工智能辅助的组织决策中,存在一种高效分工:大模型作为广泛类比生成器,人类作为关键评估者,筛选最契合情境的类比用于战略问题解决。

原文摘要 · Abstract (English)

This study investigates whether large language models, specifically GPT4, can match human capabilities in analogical reasoning within strategic decision making contexts. Using a novel experimental design involving source to target matching, we find that GPT4 achieves high recall by retrieving all plausible analogies but suffers from low precision, frequently applying incorrect analogies based on superficial similarities. In contrast, human participants exhibit high precision but low recall, selecting fewer analogies yet with stronger causal alignment. These findings advance theory by identifying matching, the evaluative phase of analogical reasoning, as a distinct step that requires accurate causal mapping beyond simple retrieval. While current LLMs are proficient in generating candidate analogies, humans maintain a comparative advantage in recognizing deep structural similarities across domains. Error analysis reveals that AI errors arise from surface level matching, whereas human errors stem from misinterpretations of causal structure. Taken together, the results suggest a productive division of labor in AI assisted organizational decision making where LLMs may serve as broad analogy generators, while humans act as critical evaluators, applying the most contextually appropriate analogies to strategic problems.

类比推理大模型人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。