arXiv:2604.14140cs.LGcs.AI2026-04被引 2

测试大模型长链条推理能力,发现顶尖模型准确率不足10%。

LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning

论文配图:LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
图 1 · 摘自论文原文
  • 设计2500道跨学科复杂题,检验模型长序列思维链能力。
  • 当前最优模型在长推理任务上准确率低于10%(如GPT-5.2仅9.8%)。
  • 适合评估大模型长期规划与复杂任务处理能力的研究者使用。

随着语言模型被用于复杂的自主任务,其在长时程范围内进行准确推理的能力变得至关重要。其中关键在于规划和管理复杂的长链条思维(Chain-of-Thought, CoT)。我们提出LongCoT,一个可扩展的基准测试集,包含2500道由专家设计的问题,涵盖化学、数学、计算机科学、国际象棋和逻辑等领域,旨在直接衡量前沿模型的长时程思维链推理能力。每道题包含简短输入和可验证答案,解题需经过数十万至数十万推理标记的相互依赖步骤构成的图结构。每个局部步骤对前沿模型而言均具备可处理性,因此失败原因反映的是长时程推理局限。发布时,最佳模型在LongCoT上的准确率低于10%(GPT-5.2:9.8%;Gemini 3 Pro:6.1%),揭示了当前能力的重大差距。总体而言,LongCoT为长时程推理提供了严谨的度量标准,可追踪前沿模型在长时间跨度下可靠推理的能力。

原文摘要 · Abstract (English)

As language models are increasingly deployed for complex autonomous tasks, their ability to reason accurately over longer horizons becomes critical. An essential component of this ability is planning and managing a long, complex chain-of-thought (CoT). We introduce LongCoT, a scalable benchmark of 2,500 expert-designed problems spanning chemistry, mathematics, computer science, chess, and logic to isolate and directly measure the long-horizon CoT reasoning capabilities of frontier models. Problems consist of a short input with a verifiable answer; solving them requires navigating a graph of interdependent steps that span tens to hundreds of thousands of reasoning tokens. Each local step is individually tractable for frontier models, so failures reflect long-horizon reasoning limitations. At release, the best models achieve <10% accuracy (GPT 5.2: 9.8%; Gemini 3 Pro: 6.1%) on LongCoT, revealing a substantial gap in current capabilities. Overall, LongCoT provides a rigorous measure of long-horizon reasoning, tracking the ability of frontier models to reason reliably over extended periods.

长链条推理评测基准大模型能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。