arXiv:2504.02671cs.CL2025-04

探究大模型解决费米问题的能力与局限,发现其表现远未达人类水平。

LLM for Complex Reasoning Task: An Exploratory Study in Fermi Problems

  • 按TELeR分类框架设计提示词,评估三款大模型在费米问题上的表现。
  • 所有模型得分均低于0.5,标准类问题比具体类问题更易解决。
  • 揭示大模型在模糊现实推理中的短板,适合对推理能力有研究兴趣者阅读。

费米问题(Fermi Problems, FPs)是需要类人逻辑与数值推理的数学推理任务。与其它推理题不同,费米问题常涉及现实不可行性或模糊概念,即使对人类也极具挑战。尽管人工智能取得进展,特别是大语言模型(LLMs)在各类推理任务中表现优异,但费米问题仍相对未被充分探索。本研究开展一项探索性研究,评估大模型在解决费米问题上的能力与局限。我们首先使用公开的费米问题数据集,评估三款先进大模型的整体表现,并依据近期提出的TELeR分类框架设计提示词,包括零样本场景。结果表明,所有模型的fp_score(0-1区间)均低于0.5,凸显此类推理任务的内在难度。为进一步分析,我们将费米问题分为标准类与特定类,假设大模型在表述清晰简洁的标准类问题上表现更好。对比实验验证了该假设,显示大模型在准确率和效率上均更优地应对标准类费米问题。

原文摘要 · Abstract (English)

Fermi Problems (FPs) are mathematical reasoning tasks that require human-like logic and numerical reasoning. Unlike other reasoning questions, FPs often involve real-world impracticalities or ambiguous concepts, making them challenging even for humans to solve. Despite advancements in AI, particularly with large language models (LLMs) in various reasoning tasks, FPs remain relatively under-explored. This work conducted an exploratory study to examine the capabilities and limitations of LLMs in solving FPs. We first evaluated the overall performance of three advanced LLMs using a publicly available FP dataset. We designed prompts according to the recently proposed TELeR taxonomy, including a zero-shot scenario. Results indicated that all three LLMs achieved a fp_score (range between 0 - 1) below 0.5, underscoring the inherent difficulty of these reasoning tasks. To further investigate, we categorized FPs into standard and specific questions, hypothesizing that LLMs would perform better on standard questions, which are characterized by clarity and conciseness, than on specific ones. Comparative experiments confirmed this hypothesis, demonstrating that LLMs performed better on standard FPs in terms of both accuracy and efficiency.

费米问题大模型推理逻辑推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。