arXiv:2501.11833cs.CLcs.AI2025-01被引 3

研究大模型在固有思维定式下的推理表现,揭示其适应新情境的能力短板。

Is your LLM trapped in a Mental Set? Investigative study on how mental sets affect the reasoning capabilities of LLMs

  • 引入认知心理学中的思维定式概念,评估大模型在陌生场景中的适应性。
  • 对比Llama-3.1-8B/70B与GPT-4o在思维定式干扰下的推理表现差异。
  • 为模型选型提供新视角,适合关注推理鲁棒性与可解释性的研究者。

本文首次将认知心理学中的思维定式(Mental Set)概念引入大模型复杂推理能力的评估中。尽管大模型在自然语言处理任务上表现优异,依赖参数高效微调(PEFT)和上下文学习(ICL)等技术,但现有评估方法如MMLU、MATH、GSM8K等指标,仅关注性能得分或由大模型生成的推理链,忽视了模型在面对新情境时克服固有策略的能力。思维定式指个体坚持以往有效策略,即使其已失效,这严重影响问题解决效率。本研究对比了Llama-3.1-8B-Instruct、Llama-3.1-70B-Instruct与GPT-4o在思维定式干扰下的表现,揭示了不同模型在应对非标准问题时的适应性差异。该研究为理解大模型的真实推理能力提供了更深层洞见。

原文摘要 · Abstract (English)

In this paper, we present an investigative study on how Mental Sets influence the reasoning capabilities of LLMs. LLMs have excelled in diverse natural language processing (NLP) tasks, driven by advancements in parameter-efficient fine-tuning (PEFT) and emergent capabilities like in-context learning (ICL). For complex reasoning tasks, selecting the right model for PEFT or ICL is critical, often relying on scores on benchmarks such as MMLU, MATH, and GSM8K. However, current evaluation methods, based on metrics like F1 Score or reasoning chain assessments by larger models, overlook a key dimension: adaptability to unfamiliar situations and overcoming entrenched thinking patterns. In cognitive psychology, Mental Set refers to the tendency to persist with previously successful strategies, even when they become inefficient - a challenge for problem solving and reasoning. We compare the performance of LLM models like Llama-3.1-8B-Instruct, Llama-3.1-70B-Instruct and GPT-4o in the presence of mental sets. To the best of our knowledge, this is the first study to integrate cognitive psychology concepts into the evaluation of LLMs for complex reasoning tasks, providing deeper insights into their adaptability and problem-solving efficacy.

思维定式推理能力大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。