用故事类类比任务测试大模型推理能力,发现其表现与人类相似但仍有差距。
Modeling Understanding of Story-Based Analogies Using Large Language Models
- 通过句子嵌入分析模型对类比语义的捕捉能力
- 70B参数模型在类比推理上优于8B模型,接近人类水平
- 首次逐个分析类比题表现,揭示模型与人类的推理模式差异
大语言模型(LLMs)在多项任务中已接近人类认知水平。它们在识别和映射类比关系方面表现如何?以往研究显示,尽管模型能提取类比中的相似性,但缺乏稳健的人类式推理能力。本研究基于Webb、Holyoak和Lu(2023)的工作,聚焦故事类类比映射任务,开展细粒度评估,比较模型与人类的推理表现。首先,利用句子嵌入评估模型是否能准确捕捉类比源文本与目标文本间的相似性,以及源文本与干扰项之间的差异性。其次,考察显式提示模型解释类比的效果。研究不只关注整体准确率,而是以单个类比为单位评估推理过程。实验涵盖不同模型规模(80亿与700亿参数)及主流架构如GPT-4和LLaMA3的表现差异。结果深化了我们对大模型类比推理能力的理解,也为将其作为人类推理模型提供了依据。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) have brought them closer to matching human cognition across a variety of tasks. How well do these models align with human performance in detecting and mapping analogies? Prior research has shown that LLMs can extract similarities from analogy problems but lack robust human-like reasoning. Building on Webb, Holyoak, and Lu (2023), the current study focused on a story-based analogical mapping task and conducted a fine-grained evaluation of LLM reasoning abilities compared to human performance. First, it explored the semantic representation of analogies in LLMs, using sentence embeddings to assess whether they capture the similarity between the source and target texts of an analogy, and the dissimilarity between the source and distractor texts. Second, it investigated the effectiveness of explicitly prompting LLMs to explain analogies. Throughout, we examine whether LLMs exhibit similar performance profiles to those observed in humans by evaluating their reasoning at the level of individual analogies, and not just at the level of overall accuracy (as prior studies have done). Our experiments include evaluating the impact of model size (8B vs. 70B parameters) and performance variation across state-of-the-art model architectures such as GPT-4 and LLaMA3. This work advances our understanding of the analogical reasoning abilities of LLMs and their potential as models of human reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。