对比大模型与人类在逻辑抽象推理上的表现差异。
Human-Level Reasoning: A Comparative Study of Large Language Models on Logical and Abstract Reasoning
- 设计8个定制推理题,评估大模型的逻辑与抽象能力。
- 多款大模型在推理任务中表现显著低于人类水平。
- 揭示大模型在演绎推理中的短板,适合研究AI认知局限者参考。
评估大型语言模型(LLMs)的推理能力对推动人工智能发展至关重要,因其超越了单纯的语言任务表现,涉及模型是否真正理解信息、进行推理并以逻辑有效方式得出结论。本研究通过一组八个定制设计的推理问题,比较了包括GPT、Claude、DeepSeek、Gemini、Grok、Llama、Mistral、Perplexity和Sabiá在内的多个大模型在逻辑与抽象推理方面的能力,并将其结果与人类在相同任务上的表现进行基准对比,揭示了显著差异,表明大模型在演绎推理方面存在明显不足。
原文摘要 · Abstract (English)
Evaluating reasoning ability in Large Language Models (LLMs) is important for advancing artificial intelligence, as it transcends mere linguistic task performance. It involves understanding whether these models truly understand information, perform inferences, and are able to draw conclusions in a logical and valid way. This study compare logical and abstract reasoning skills of several LLMs - including GPT, Claude, DeepSeek, Gemini, Grok, Llama, Mistral, Perplexity, and Sabiá - using a set of eight custom-designed reasoning questions. The LLM results are benchmarked against human performance on the same tasks, revealing significant differences and indicating areas where LLMs struggle with deduction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。