arXiv:2602.21061cs.AI2026-02

通过测试时搜索让大模型实现超智能,关键在工具调用的精准性。

Tool Building as a Path to "Superintelligence"

  • 设计基于GF(2)电路重构的任务,检验模型逐步推理能力。
  • 小模型推理成功率随深度下降超线性,前沿模型具部分鲁棒性。
  • 精准工具调用是实现通用超智能的关键,工具构建至关重要。

Diligent Learner 框架认为,只要每一步的成功概率 $γ$ 足够高,大语言模型可通过测试时搜索实现超智能。本文设计了一个基准,用于衡量逻辑分布外推理中的 $γ$ 值。我们构造了一类随推理步骤增加而变难的任务,涉及 GF(2) 电路重构,从信息论角度看,除非模型能精确整合所有输入信息,否则无法可靠求解。分析表明,小模型的 $γ$ 随深度增长呈超线性下降,而前沿模型在此任务上表现出部分鲁棒性。此外,大规模成功推理依赖于精确的工具调用,凸显工具设计对实现通用超智能的重要性。

原文摘要 · Abstract (English)

The Diligent Learner framework suggests LLMs can achieve superintelligence via test-time search, provided a sufficient step-success probability $γ$. In this work, we design a benchmark to measure $γ$ on logical out-of-distribution inference. We construct a class of tasks involving GF(2) circuit reconstruction that grow more difficult with each reasoning step, and that are, from an information-theoretic standpoint, impossible to reliably solve unless the LLM carefully integrates all of the information provided. Our analysis demonstrates that while the $γ$ value for small LLMs declines superlinearly as depth increases, frontier models exhibit partial robustness on this task. Furthermore, we find that successful reasoning at scale is contingent upon precise tool calls, identifying tool design as a critical capability for LLMs to achieve general superintelligence through the Diligent Learner framework.

超智能工具调用推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。