测试大模型在几何推理任务中的表现与提示依赖性。
Reasoning Capabilities and Invariability of Large Language Models
- 构建几何图形推理新基准,避免常识干扰。
- 700亿参数以上模型零样本表现更好但仍有提升空间。
- 思维链提示效果因顺序而异,可能增益或损害性能。
大型语言模型(LLMs)在自然语言处理中表现出色,但其简单推理能力常受质疑。本文针对模型的提示依赖性,开展全面分析,提出一个包含浅层逻辑推理题的新基准数据集。题目围绕几何图形展开,符合认知心理学标准,确保答案仅依赖演绎推理而非外部常识。对24个不同规模的LLM进行零样本和少样本提示实验,结果表明:参数量超过700亿的模型在零样本下表现更优,但仍存在显著提升空间。进一步对22个LLM采用思维链(chain-of-thought)提示发现,若要求在答案前提供推理过程,性能可能提升;若在答案后补充推理,则可能产生负面影响。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown remarkable capabilities in manipulating natural language across multiple applications, but their ability to handle simple reasoning tasks is often questioned. In this work, we aim to provide a comprehensive analysis of LLMs' reasoning competence, specifically focusing on their prompt dependency. In particular, we introduce a new benchmark dataset with a series of simple reasoning questions demanding shallow logical reasoning. Aligned with cognitive psychology standards, the questions are confined to a basic domain revolving around geometric figures, ensuring that responses are independent of any pre-existing intuition about the world and rely solely on deduction. An empirical analysis involving zero-shot and few-shot prompting across 24 LLMs of different sizes reveals that, while LLMs with over 70 billion parameters perform better in the zero-shot setting, there is still a large room for improvement. An additional test with chain-of-thought prompting over 22 LLMs shows that this additional prompt can aid or damage the performance of models, depending on whether the rationale is required before or after the answer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。