arXiv:2510.11652cs.CL2025-10被引 1

构建学术推理新基准,揭示大模型在高阶研究任务中的能力短板

ACADREASON: Exploring the Limits of Reasoning Models with Academic Research Problems

  • 设计跨五领域的50道专家标注学术题,覆盖计算机、经济、法律等
  • 主流大模型平均分不足20,顶级GPT-5仅得16分,代理模型最高40分
  • 适合评估AI在科研级推理中的真实能力,推动智能研究系统发展

近年来,大语言模型与智能体的研究重点从展示新能力转向复杂推理与挑战性任务。然而,现有评估多集中于数学/编程竞赛或通用任务,多领域学术基准又缺乏足够的推理深度,导致高阶推理缺乏严谨评测标准。为此,我们提出Acadreason基准,用于评估大模型与智能体获取并运用学术知识的能力。该基准包含50道来自近年顶级期刊的专家标注学术问题,涵盖计算机科学、经济学、法学、数学和哲学五个高推理强度领域。所有题目均经过严格标注与质量控制,确保兼具挑战性与可答性。我们对超过10个主流大模型与智能体进行了系统评估。结果显示,大多数大模型得分低于20分,即使是前沿的GPT-5也仅获16分;智能体表现更优,但最高分未超40分。这揭示了当前大模型与智能体在超智能科研任务中的显著能力差距,并凸显Acadreason的挑战性。

原文摘要 · Abstract (English)

In recent years, the research focus of large language models (LLMs) and agents has shifted increasingly from demonstrating novel capabilities to complex reasoning and tackling challenging tasks. However, existing evaluations focus mainly on math/code contests or general tasks, while existing multi-domain academic benchmarks lack sufficient reasoning depth, leaving the field without a rigorous benchmark for high-level reasoning. To fill this gap, we introduce the Acadreason benchmark, designed to evaluate the ability of LLMs and agents to acquire and reason over academic knowledge. It consists of 50 expert-annotated academic problems across five high-reasoning domains, including computer science, economics, law, mathematics, and philosophy. All questions are sourced from top-tier publications in recent years and undergo rigorous annotation and quality control to ensure they are both challenging and answerable. We conduct systematic evaluations of over 10 mainstream LLMs and agents. The results show that most LLMs scored below 20 points, with even the cutting-edge GPT-5 achieving only 16 points. While agents achieved higher scores, none exceeded 40 points. This demonstrates the current capability gap between LLMs and agents in super-intelligent academic research tasks and highlights the challenges of Acadreason.

学术推理大模型评估智能体科研能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。