arXiv:2608.31075cs.AI2026-08

探索大模型在无人监督下持续进化路径,迈向超级智能。

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

论文配图:Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
图 1 · 摘自论文原文
  • 构建从人工标注到自动生成奖励与经验的五级演化阶梯。
  • 发现自主奖励机制可能引发奖励劫持、反馈漂移等风险。
  • 适合关注通用人工智能演进与自主学习系统的研究者。

近期大型推理模型(LRMs)的研究表明,基于可验证奖励的强化学习(RLVR)能显著提升数学与代码任务中的推理能力,因为这些任务的结果可自动验证。然而,将这一进展扩展到开放性与代理式任务仍面临挑战,因可靠奖励难以获取,且人工监督无法跟上模型生成经验的规模与复杂度。本文研究当人类监督逐渐退出学习循环时,LRMs如何继续提升。我们分析两个关键维度:奖励轴从单实例人工判断,发展为可复用的验证器与无须人类反馈的奖励机制;经验轴则从人工构建的任务与环境,演变为自生成课程、构造环境及自主协同演化。通过一个从L0到L4的五级阶梯,我们界定学习过程中仍受人类控制的部分。分析还揭示了日益自主的奖励与经验生成带来的风险,包括奖励劫持、反馈漂移、课程坍塌和环境错误。为此,我们提出对策略能力、反馈保真度与经验质量三类互补对象的评估框架。本研究为超越人类监督的LRM扩展提供了结构化视角,并指出了通向超级智能的自维持学习系统所面临的开放问题。此外,我们维护一个持续更新的GitHub仓库以追踪最新进展。

原文摘要 · Abstract (English)

Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per-instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human-curated tasks and environments toward self-generated curricula, constructed environments, and autonomous co-evolution. We connect these dimensions through a five-level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated \href{https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision}{GitHub repository} to track the latest advances.

大模型推理自主学习超智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。